MorphIRe · Morphologically-informed representations for natural language processing
Horizon 2020 — Marie Skłodowska-Curie Actions
- Duration
- 2019-04-01 → 2021-03-31
- EU contribution
- €207,312
- Participants
- 1
- Scheme
- MSCA-IF-EF-ST
Lines connect the coordinator with its partners.
Results in brief
Morphologically-informed representations for natural language processing
Natural language processing (NLP) is a technology that we encounter frequently in the digital world: for example, it is involved when we use an automatic translation service, or when typing a question into a search engine and getting back an answer extracted from the web. While this often works remarkably well for languages like English, the performance of such systems is significantly worse when it involves less-researched languages like Basque, Finnish, or Polish. This is an important societal issue, as it contributes to a "digital language divide" where speakers who are not proficient in English are put at a disadvantage. The MorphIRe project worked on closing this gap by developing techniques that perform better on a broader range of the world's languages. Today's NLP models mostly work with artificial intelligence and machine learning: techniques that require large amounts of training data---e.g. sets of questions with their correct answers---which are then fed into an algorithm that "learns" to perform the task. Importantly, the techniques that are widely used today are indifferent to which language is being used---whether the task is performed on English or on Basque, the algorithms work exactly the same. In particular, they do not take into account the word-internal (i.e., "morphological") structure of these languages: whereas English tends to use separate words to express different grammatical and semantic concepts, morphologically richer languages like Basque can express these concepts within a single word form (compare English "because of the rain" with Basque "euriagatik"). The MorphIRe project provides direct evidence that we shouldn't ignore the morphological structure of languages when building NLP models, as it contributes to errors that current state-of-the-art NLP models make. It also proposes a new algorithm for word segmentation that better corresponds to morphological structure. By highlighting these problems in today's NLP models and working towards concrete solutions that can be integrated into these models, the MorphIRe project makes an important contribution towards improving NLP technology for a wider range of languages.
Data: CORDIS, © European Union
Project objective
The morphological structure of a word plays an important role in determining its function and meaning, yet it is often disregarded by current machine learning models aimed at natural language processing (NLP). State-of-the-art NLP models typically rely on word-level or character-level representations. This arguably works well for English, the dominant language in NLP research, since it is morphologically simple, but poses a challenge for morphologically-rich languages like Basque, Estonian, or Kurdish. As a consequence, the current state of the art is biased against these languages, preventing us from building better NLP technology for them.The MorphIRe project aims to learn morphologically-informed representations for NLP. It proposes to explore the fine-grained morphological analysis of word forms in order to learn representations that are grounded in morphemes, the smallest grammatical unit of language. Using these representations as input to NLP models is expected to improve their performance particularly for morphologically-rich languages. To this end, MorphIRe will make use of deep learning with neural network architectures both to learn the representations and to apply them to state-of-the-art models for a variety of NLP tasks, such as language modelling and dependency parsing.The impact of MorphIRe is twofold: 1) Learning input representations that can be used in a variety of models encourages reusability of the results and promises that improvements will carry over to future NLP research. 2) Through improving the state of the art on morphologically-rich languages, speakers of these languages will ultimately benefit from better NLP technology. This way, MorphIRe has the potential for making both a scientific and a societal impact.
Original text from CORDIS.
Participants
- KOBENHAVNS UNIVERSITET · KOBENHAVNCoordinatorDenmark
Links
Data: CORDIS, © European Union
