FP7Индивидуална стипендия2014–2016

LATEST · Advanced LAnguage TEchnology Platform for TranSlaTors (LATEST)

7РП — „Хора“ (Действия „Мария Кюри“)

Период
2014-02-01 → 2016-01-31
Финансиране от ЕС
223 002 €
Участници
1
Схема
MC-IEF

Линиите свързват координатора с партньорите.

Накратко на български

Технологичната платформа LATEST автоматично разпознава устойчиви словосъчетания (например глагол и съществително) и предлага преводите им. Това помага на преводачите и синхронните преводачи да разбират и превеждат по-точно тези специфични езикови структури.

Този кратък обзор е генериран от изкуствен интелект

Кратко обяснение, генерирано от езиков модел по текста на CORDIS. Оригиналът е по-долу.

Резултати накратко

Advanced LAnguage TEchnology Platform for TranSlaTors (LATEST)

The objective of LATEST which drew on the input from and collaboration with translators, was the development of original methodology for, and the implementation of, a novel LAnguage TEchnology platform for tranSlaTors (LATEST) to automatically identify multiword expressions (collocations) and provide their translations, thus assisting translators and interpreters to understand and translate them. LATEST regards the task of translating multiword expressions as a two-stage process. The first stage is the extraction of multiword expressions (MWEs) in each of the languages; the second stage is a matching procedure for the extracted MWEs in each language which proposes the translation equivalents. The methodology which works for any pair of languages, is based on a knowledge-poor approach which does not depend on translation resources such as dictionaries, translation memories or parallel corpora which can be time consuming to develop or difficult to acquire, being expensive or proprietary. The only information comes from comparable corpora, inexpensively compiled. The implementation covers English and Spanish and focuses on a particular subclass of multiword expressions (MWEs) verb-noun expressions (collocations). In the MWE extraction phase, statistical association measures were employed to quantify the strength of the relationship between two words and to propose that a combination of a verb and a noun above a specific threshold would be a (candidate for) multiword expression. In the MWE translation phase, distributional similarity methods were applied the premise being that MWEs have the same or very similar contexts as their translation equivalents. The comparable corpora compiled for this project covered the newswire genre which was determined by the fact that newswire is a widespread genre and available in different languages. As the approach was developed as language independent, the intention was to have it tested for different languages after the lifetime of the project. Extensive evaluation experiments were conducted both for the performance in the extraction and translation phase, with detailed comparison of the various measures provided and the interannotator agreement of all annotators computed. The effect of the quality of the comparable corpora and as well as their size was investigated as well. The evaluation results point to a very interesting finding and sheds light for the first time on the following. It is the quality of the comparable corpora that is more important than the size of the data for the performance of automatic translation of MWEs. LATEST achieved its specific research objectives and research training objectives. The completion of the research training objectives and the acquisition of new skills, were instrumental in the fellow successfully refocusing his research into a new area, that of (computational) phraseology where he achieved significant novel results. This is evidenced by the publications related to the project and invitations to give talks at international conferences. Wide range of dissemination and outreach activities contributed to publicising the project and making the outputs easy to understand.

Текст от CORDIS, на английски · Данни: CORDIS, © Европейски съюз

Цел на проекта

In today's multicultural world of multilingual communication and exponential growth of information, specialised texts and terminology, the need for correct understanding and reliable translation is greater than ever. Most Natural Language Processing (NLP) translation tools have been developed by NLP researchers and the practical needs of translators have been virtually ignored. Translators (and interpreters) face the everyday challenge of translating emerging terms which do not appear in dictionaries, and the fact that a very high proportion of these new terms are multiword units, which need to be recognised and translated as such, can only complicate the situation. Unfortunately, translator-orientated practical tools to assist in such scenarios are a scarce commodity, if one at all. This project, which will build on the input from and collaboration with translators, aims to fill in this gap. Its objective is the development of original methodology for, and the implementation of, a novel LAnguage TEchnology platform for tranSlaTors (LATEST). LATEST will go beyond words in that it will automatically identify multiword terms (collocations), and assist translators and interpreters to understand and translate them. It will explain the meaning in specific context by simply clicking on words or expressions. It will also provide the best translation in context by selecting the most appropriate translation from a set of possible translations in the target language.The innovative methodology underlying the functionalities of the tools to be developed will be based on advanced NLP approaches. A valuable by-product of the project will be the development of new resources (comparable corpora) as well as the development of a tool for automatic compilation of such corpora, which are essential for the operation of LATEST and for the work (translation, documentary search) of translators and interpreters in general.

Оригинален текст от CORDIS (на английски).

Участници

Връзки

Данни: CORDIS, © Европейски съюз