Logo Logo
FAQ
Contact
Switch language to German
Towards multilinguality: overcoming language and script barriers in natural language processing
Towards multilinguality: overcoming language and script barriers in natural language processing
With over 7,000 languages written in at least 293 scripts, the majority of the world's languages remain underrepresented or even unsupported in current pretrained language models (PLMs) and large language models (LLMs), resulting in inequitable progress in the field of natural language processing (NLP). Moreover, many related languages use different scripts, complicating crosslingual knowledge transfer, which often relies on lexical overlap. This thesis tackles these challenges by addressing language and script barriers through a series of studies in highly multilingual contexts. The first series of studies investigates how highly multilingual resources can be leveraged to build valuable NLP models for as many languages as possible. We introduce Conceptualizer, a method for automatically extracting colexification patterns for a selected set of concepts from multilingual corpora encompassing over 1,300 languages. Our findings reveal that the concreteness of a concept predicts its crosslingual stability. We introduce a measure of conceptual similarity between languages, complementing standard genealogical, typological, and surface similarity measures. We further expand our coverage to over 2,000 concepts and build multilingual graphs, ColexNet and ColexNet+, from which high-quality multilingual word vectors are learned to support crosslingual transfer. The next series of studies focuses on enhancing large-scale multilingual pretraining with the help of available multilingual resources and small models derived from such resources, with the aim of improving the multilinguality of PLMs. We propose two strategies to achieve this objective: OFA and LangSAMP. The former leverages well-aligned external multilingual word vectors derived from ColexNet+ and matrix factorization to initialize subword embeddings, achieving superior performance with lower computational and environmental costs than existing pretraining methods. The latter introduces language and script embeddings during pretraining, fostering more language-neutral contextualized representations and boosting crosslingual transfer. The final series of studies provides viable methods for breaking script barriers in PLMs in order to improve crosslingual transfer across languages written in different scripts. We introduce transliteration-based learning objectives: transliteration contrastive modeling (TCM) and transliteration language modeling (TLM). These approaches enhance crosslingual alignment, especially for languages with underrepresented scripts, without requiring parallel data. Our methods significantly improve crosslingual transfer performance, revealing the potential of transliteration for improving alignment. In addition, we systematically investigate how and why transliteration-based approaches achieve better crosslingual alignment. This thesis contributes to multilingual NLP by expanding the coverage of languages supported by NLP models, improving the efficiency and effectiveness of multilingual pretraining, enhancing the crosslingual alignment of existing models, and advancing transfer performance across languages and scripts.
Multilingual NLP, Crosslingual Transfer, Script Barriers, Large Language Models, Crosslingual Alignment
Liu, Yihong
2026
English
Universitätsbibliothek der Ludwig-Maximilians-Universität München
Liu, Yihong (2026): Towards multilinguality: overcoming language and script barriers in natural language processing. Dissertation, LMU München: Faculty of Mathematics, Computer Science and Statistics
[thumbnail of Liu_Yihong.pdf]
Preview
PDF
Liu_Yihong.pdf

17MB

Abstract

With over 7,000 languages written in at least 293 scripts, the majority of the world's languages remain underrepresented or even unsupported in current pretrained language models (PLMs) and large language models (LLMs), resulting in inequitable progress in the field of natural language processing (NLP). Moreover, many related languages use different scripts, complicating crosslingual knowledge transfer, which often relies on lexical overlap. This thesis tackles these challenges by addressing language and script barriers through a series of studies in highly multilingual contexts. The first series of studies investigates how highly multilingual resources can be leveraged to build valuable NLP models for as many languages as possible. We introduce Conceptualizer, a method for automatically extracting colexification patterns for a selected set of concepts from multilingual corpora encompassing over 1,300 languages. Our findings reveal that the concreteness of a concept predicts its crosslingual stability. We introduce a measure of conceptual similarity between languages, complementing standard genealogical, typological, and surface similarity measures. We further expand our coverage to over 2,000 concepts and build multilingual graphs, ColexNet and ColexNet+, from which high-quality multilingual word vectors are learned to support crosslingual transfer. The next series of studies focuses on enhancing large-scale multilingual pretraining with the help of available multilingual resources and small models derived from such resources, with the aim of improving the multilinguality of PLMs. We propose two strategies to achieve this objective: OFA and LangSAMP. The former leverages well-aligned external multilingual word vectors derived from ColexNet+ and matrix factorization to initialize subword embeddings, achieving superior performance with lower computational and environmental costs than existing pretraining methods. The latter introduces language and script embeddings during pretraining, fostering more language-neutral contextualized representations and boosting crosslingual transfer. The final series of studies provides viable methods for breaking script barriers in PLMs in order to improve crosslingual transfer across languages written in different scripts. We introduce transliteration-based learning objectives: transliteration contrastive modeling (TCM) and transliteration language modeling (TLM). These approaches enhance crosslingual alignment, especially for languages with underrepresented scripts, without requiring parallel data. Our methods significantly improve crosslingual transfer performance, revealing the potential of transliteration for improving alignment. In addition, we systematically investigate how and why transliteration-based approaches achieve better crosslingual alignment. This thesis contributes to multilingual NLP by expanding the coverage of languages supported by NLP models, improving the efficiency and effectiveness of multilingual pretraining, enhancing the crosslingual alignment of existing models, and advancing transfer performance across languages and scripts.