H2020Individual fellowship2020–2022

SEBAMAT · Semantics-Based Machine Translation

Horizon 2020 — Marie Skłodowska-Curie Actions

Duration
2020-04-01 → 2022-03-31
EU contribution
€165,085
Participants
1
Scheme
MSCA-IF

Lines connect the coordinator with its partners.

Results in brief

Semantics-Based Machine Translation

Most current machine translation systems are corpus-based. They typically take the semantics of a text only in so far into account as they are implicit in the underlying text corpora. This is also true for the recent neural machine translation systems which, in comparison to standard phrase-based systems, tend to have the focus even more on fluency rather than adequacy. However, the question is whether it is possible to improve the use of semantic knowledge. For example, it has been suggested that future machine translation systems should use information of the type "who is doing what to whom, when and why", which may require the identification of the semantic roles of the items occurring in a sentence. To move forward in the direction of semantics-based machine translation, we propose to implement and evaluate three different approaches: The first approach is based on state of the art machine translation but considers word senses rather than words. That is, a word sense disambiguation system is used to determine the word senses in large parallel text corpora. Then a neural machine translation system is trained on the word-sense-disambiguated rather than the original parallel corpora. Our second approach uses role labeling for identifying the semantic roles of the words in a sentence. In this case the neural machine translation system is trained on a corpus which was annotated with semantic roles. With both word sense disambiguation and semantic role labeling it is hoped that the respective annotation software does a better job than what neural machine translation is doing implicitly, and that this may improve translation quality. In contrast, the third approach tries to reduce the data acquisition bottleneck as encountered in the case of low-resource languages. It uses multilingual neural machine translation systems to translate between language pairs where no parallel data is available.

Data: CORDIS, © European Union

Project objective

Most current machine translation systems are either rule-based or corpus-based. They typically take the semantics of a text only in so far into account as they are implicit in the underlying text corpora or dictionaries. This is also true for the recent neural machine translation systems, which - in comparison to standard phrase-based systems, tend to have the focus even more on fluency rather than adequacy. However, it has been pointed out that it is unlikely to be able to bring machine translation quality to the next level as long as the systems do not make better use of semantic knowledge. For example, according to Kevin Knight future machine translation systems should use information of the type ""who is doing what to whom and when"", i.e. involving the identification of the semantic roles of the items occurring in a sentence. To move forward in this direction, we propose to implement and evaluate three different approaches: The first approach is based on state of the art machine translation but considers word senses rather than words. That is, a word sense disambiguation system is used to identify the word senses in large parallel text corpora. Then, in analogy to standard word alignment, the word senses are aligned across languages, and the resulting multilingual sense dictionaries are used in conjunction with the word sense disambiguation systems for translating new texts. Our second approach uses role labeling for identifying the semantic roles of the words in a sentence. The roles are aligned across languages, and this information is then used to improve the translation process. The third approach is based on an algorithm which computes the semantic similarity between phrases. It considers the translation task as finding semantically similar phrases across languages.""

Original text from CORDIS.

Participants

  • ATHINA-EREVNITIKO KENTRO KAINOTOMIAS STIS TECHNOLOGIES TIS PLIROFORIAS, TON EPIKOINONION KAI TIS GNOSIS · MAROUSSICoordinatorGreece

Links

Data: CORDIS, © European Union