H2020Индивидуална стипендия2017–2020

ML-TEXTSUM · Multi-language text summarization

„Хоризонт 2020“ — Действия „Мария Склодовска-Кюри“

Период
2017-09-01 → 2020-08-31
Финансиране от ЕС
265 840 €
Участници
2
Схема
MSCA-IF

Линиите свързват координатора с партньорите.

Накратко на български

Системата за многоезично резюмиране на текстове създава кратки копия на документи на същия или различен език. Това помага на хората да се справят с огромното количество информация и да достъпват важни данни независимо от езика, на който са написани.

Този кратък обзор е генериран от изкуствен интелект

Кратко обяснение, генерирано от езиков модел по текста на CORDIS. Оригиналът е по-долу.

Резултати накратко

Multi-language text summarization

In our daily life, we are submerged by huge amounts of text, coming from different sources such as emails, news, reports, analyses, and so on. The availability of unprecedented volumes of data represents both a challenge and an opportunity. On one hand, information overload can have severe consequences. For example, scientists might miss a relevant reference to their work due to the increased peace of publication, losing months (if not years) of work; intelligence systems might miss a security thread buried down vast amounts of data, etc. On the other hand, there is widespread agreement that the effective harnessing of text and data mining techniques is important to the performance of advanced economies, such as those of the European Union (EU). With respect to the development of text and data mining techniques, the EU faces an added challenge due to its rich cultural heritage. Multilingualism is a core value of the European Union, as integral to Europe as the freedom of movement, the freedom of residence and the freedom of expression. Hence, it is only admissible to tackle text understanding challenges from a multi-language perspective, in order to ensure that knowledge is distributed independently of spoken language. The objective of this project is to develop a system for efficient and accurate multi-lingual text summarization. That is, given as input a text document, the system will output a summary of the document in the same or in a different language. The availability of such system shall allow citizens, regardless of their language, to better handle the information overload and to gain access to critically distilled information (e.g., what is a certain newspaper’s opinion on the same topic this year? Are male/female athletes portrayed differently by the media?). Conclusion ---------- This project was terminated after 13 months. In this time significant progress has been made in the foundations and its computational aspects. This resulted in several publications in top venues that I describe below. Compared to the plan outlined in the proposal, the work accomplished in this time is more theoretical in nature. We proposed novel optimization techniques that pave the ground to develop more refined modeling techniques for the text summarization problem, although due to its early termination, this perspective remains to be fully exploited.

Текст от CORDIS, на английски · Данни: CORDIS, © Европейски съюз

Цел на проекта

In our daily life, we are submerged by huge amounts of text, coming from different sources such as emails, news, reports, and so on. The availability of unprecedented volumes of data represents both a challenge and an opportunity. On one hand, it can lead to information overload, a phenomenon that limits one’s capacity to understand an issue and act in the presence of too much information. On the other hand, the effective harnessing of this information has undeniable economical potential. Furthermore, In the European context, special needs to be put to multilingualism to guarantee global access to high quality information.The objective of this application is to develop ML-TEXTSUM, a system for efficient and accurate multi-lingual text summarization. That is, given as input a text document, the system will output a summary of the document in the same or in a different language. Building on recent breakouts in machine learning and natural language processing, I propose a novel architecture for ML-TEXTSUM that will be able to produce high quality summaries while at same time remain modular enough so that new languages can be added with minimal effort. The availability of such system shall allow citizens, regardless of their language, to better handle the information overload and to gain access to critically distilled information (e.g., what is a certain newspaper’s opinion on the same topic this year? Are male/female athletes portrayed differently by the media?). The project is characterized by the interplay of multiple disciplines: the proposed architecture requires to master a combination of natural language processing and machine learning techniques. At the same time, the formidable scale of this system will require the development of novel distributed optimization methods. This interplay will be achieved thanks to my past and future collaborations, my solid background in optimization and machine learning, as well as through the acquisition of new ad-hoc skills.

Оригинален текст от CORDIS (на английски).

Участници

Връзки

Данни: CORDIS, © Европейски съюз