[Connected 2021] Extending NLP Support for Less-Resourced Language Using Word Vectors

Опубликовано: 04 Апрель 2026
на канале: Globalization and Localization Association (GALA)
20
0

Natural language processing tools are now omnipresent in the localization industry. Automatic creation and querying of electronic dictionaries, enhanced translation memory lookup, automatic post editing and machine translation are a must-have in today’s world. Leveraging their potential both speeds up and improves the quality of translation. However, the presence of these natural language processing tools for a given language relies heavily on the availability of electronic language resources in this language. Resources of the highest value include: mono and bilingual dictionaries and parallel or monolingual text corpora. Lack of these resources constitutes a significant impediment in the development of natural language processing for some languages. These languages are often referred to as less-resourced languages.

Importantly, the status of a less-resourced language does not necessarily correlate with the number of people speaking the language and, in consequence, the market demand for translation services involving this language. In our scenario, we focused on several languages of India: Assamese, Bengali, Gujarati, Hindi, Kannada, Malayalam, Marathi, Nepali, Panjabi and Sindhi. We aimed at providing bilingual word-level alignment - a fundamental functionality that can serve to develop several other natural language processing mechanisms.

For more information about GALA: https://www.gala-global.org/

For more on globalization and localization news, subscribe to our newsletter: https://bit.ly/3I2FRmq

For more on GALA events: https://bit.ly/3FDKZNY

#localization #SmallLanguages #machinetranslation