NMT-Clinical: Investigation into Multilingual Pre-Trained Language Models and Transfer-Learning

Опубликовано: 21 Март 2026
на канале: Aaron Poet
35
0

Check out the detailed presentation on our accepted manuscript “Neural Machine Translation of Clinical Text: An Empirical Investigation into Multilingual Pre-Trained Language Models and Transfer-Learning” by #Frontiers in #Digital #Health
PPT https://docs.google.com/presentation/...

(https://lnkd.in/denbqrAJ)

#Highlights

#Motivation: Clinical texts and documents contain a wealth of information and knowledge in the field of healthcare, and their processing, using state-of-the-art language technology, has become very important for building intelligent systems capable of supporting healthcare and providing greater social good.

This processing includes creating language understanding models and translating resources into other natural languages to share domain-specific cross-lingual knowledge.

#Methods: we conduct investigations on clinical text machine translation by examining multilingual neural network models using deep learning methods such as Transformer-based structures.

Furthermore, to address the issue of language resource imbalance, we also carry out experiments using a transfer learning methodology based on massive multilingual pre-trained language models (MMPLMs).

#Outcomes: The experimental results on three sub-tasks including 1) clinical case (CC), 2) clinical terminology (CT), and 3) ontological concept (OC) show that our models achieved top-level performances in the ClinSpEn-2022 shared task on English-Spanish clinical domain data.

Furthermore, our expert-based human evaluations demonstrate that the small-sized pre-trained language model (PLM) wins in the clinical domain fine-tuning over the other two extra-large language models by a large margin. This finding has never been previously reported in the field.

Finally, the transfer learning method works well in our experimental setting using the WMT21fb model to accommodate a new Spanish language space that was not seen at the pretraining stage within WMT21fb itself – and this deserves further exploration for clinical knowledge transformation, e.g. investigation into more languages.

#Perspectives: These research findings can shed some light on domain-specific machine translation development, especially in clinical and healthcare fields.

Further research projects can be carried out based on our work to improve healthcare text analytics and knowledge transformations. Our data will be openly available for research purposes at https://github.com/HECTA-UoM/ClinicalNMT


#Keywords:

Neural Machine Translation; Clinical Text Translation; Multilingual Pre-trained Language Model; Large Language Model; Transfer Learning; Clinical Knowledge Transformation; Spanish-English Translation #LLM #MT #Spanish #ClinicalNLP Logrus Global The University of Manchester Frontiers