ARTIFACTO · Analyzing and Recognizing Time, Factuality, and Opinion in Text
7РП — „Хора“ (Действия „Мария Кюри“)
- Период
- 2009-07-01 → 2013-06-30
- Финансиране от ЕС
- 100 000 €
- Участници
- 1
- Схема
- MC-IRG
Линиите свързват координатора с партньорите.
Накратко на български
Технологиите за обработка на езика анализират времето, фактичността и мненията в текстове, например дали едно събитие е сигурно или вероятно. Това помага за пренасянето на софтуерни инструменти от английски към испански и каталански език.
Кратко обяснение, генерирано от езиков модел по текста на CORDIS. Оригиналът е по-долу.
Резултати накратко
Analyzing and Recognizing Time, Factuality, and Opinion in Text
ARTiFactO rationale. The overall goal of this research is two-fold: on the one hand, exploring methodologies on porting language processing technology from one language to another, while on the other, contributing specific technological infrastructure to the languages targeted here: Catalan (CA) and Spanish (SP). The project focuses on technology for the processing of three semantic domains: time, event factuality, and opinion information in natural language text. The choice of these three areas is motivated, firstly, by methodological reasons: conceptual systems such as time, factuality, and opinion are expressed by means of well delimited fragments of the general grammar and lexicon of any given language. At the same time, the properties and basic relations among elements in each of these systems (e.g., temporal relations of ordering: before, after, simultaneous; factuality degrees: possible, probable, certain; perspective attitudes: be in favor, against, etc.) are shared across languages. These factors make these systems perfect test-beds for exploring the cross-linguistic porting of Natural Language Processing (NLP) technology. Secondly, the choice is also motivated given that most of the research and technology developed on these areas had been carried out mainly for English. There is therefore a need for technological resources of this type in other languages as well. Project objectives. The objectives in the present project are established along two different dimensions: Methodological. Furthering the knowledge on methods and techniques for cross-lingual porting of NLP technology. Contemplating the following tasks: Task 1.a Exploring techniques for porting NLP technology considering language pairs of different degrees of similarity (EN – SP, SP – CA) Task 1.b Exploring alternatives to standard NLP components in order to cover technological gaps in less-resourced languages (e.g., dependency parsing). Practical. Building the necessary infrastructure for recognizing and analyzing time, factuality, and opinion information in Catalan and Spanish text, which will be incorporated as an active component in the information extraction system developed for these two languages at the hosting institution. It involves the following tasks: Task 2.a Setting the description model for each system (in CA & SP). Task 2.b Creating corpora annotated with time, factuality, and opinion information (CA& SP). Task 2.c Building event recognizers (for CA & SP). Task 2.d Building analyzers of temporal information: time expressions extractor/normalize and temporal relations analyzers (for CA & SP). Task 2.e Building factuality profilers (for CA & SP). Task 2.f Building opinion analyzers (for CA & SP). These methodological and practical dimensions expand into the following objectives: Objective 1. Defining description models for the systems of time, factuality and opinion as expressed in Catalan and Spanish, the 2 languages targeted by the project (Task 2.a) Objective 2. Corpus building (Task 2.b) Objective 3. Creating an event recognizer for Catalan and Spanish (Task 2.c) Objective 4. Creating time analyzers for Catalan and Spanish (Task 2.d) Objective 5. Creating factuality and opinion analyzers for Catalan and Spanish (Tasks 2.e-f) Objective 6. Final wrap-up of the system (Tasks 2.c-f) The results of the current research were expected to include: • The ARTiFactO system. A program that recognizes events, sorts them along the temporal axis, and identifies the factuality degrees and opinions assigned to them by relevant sources. • A set of description models (specification scheme and annotation guidelines) for the addressed systems on Catalan and Spanish, to contribute towards future cross-lingual standards. • A set of corpora for the systems of time, factuality, and opinion in Catalan and Spanish. • Analysis of techniques for cross-lingual transport of lexicons and grammars, taking into account language pairs with different degrees of similarity (e.g., CA–SP, and SP–EN). • Analysis of NLP techniques and tools to be used as alternatives in the case of less-resourced languages (for example, the case of dependency parsing). These achievements were expected to benefit: • The host institution, by enhancing its information extraction (IE) systems. • Research on semantic IE, regarding topics such as: designing adequate cross-lingual descriptive models for NLP applications, linking equivalent corpus resources, expanding NLP technology to other languages. • The EU as a multilingual community. Investigations on techniques for building NLP resources through cross-lingual porting help obviate linguistic borders, enhance communication among different linguistic communities, and ultimately contribute towards granting information access to all citizens.
Текст от CORDIS, на английски · Данни: CORDIS, © Европейски съюз
Цел на проекта
This project aims at exploring methodologies on porting natural language processing (NLP) technology from one language to another (Objective 1), while contributing specific technological infrastructure to the target language (Objective 2). Research on cross-linguistic porting of technology is increasingly popular due to the advantages of reusing NLP tools of costly production. Moreover, it is a key issue in the EU, given its multilinguality and the challenges this poses regarding resource sharing across linguistic borders.The project focuses on technology for processing time, factuality, and opinion expressed in text. Time information concerns “what happened when” and involves ordering events along the temporal axis, factuality has to do with the degree of certainty of events, and opinion deals with attitudes and their sources. These are basic levels for NLP tasks concerning text understanding, but started to receive some attention only recently. They are expressed through well-delimited grammar and lexicon fragments and share conceptual structures across languages, which makes them ideal test-beds for exploring the cross-language porting of NLP tools.The project will result in a suite of tools for analyzing the 3 levels of information in Catalan and Spanish, to be integrated as part of the information extraction system in the host. Objective 1 will be achieved through experiments exploring the feasibility of different procedures for porting lexical and parsing resources. Objective2 involves research at theoretical (setting description models for the phenomena) and applied levels (developing specific resources, to be done based on the experiments for Objective1).Most previous work on the area has been carried out in the USA -in the group of the researcher and other collaborating centers. By reintegrating her, the EU strengthens its scientific role in the area, guarantees future international collaborations, and allows her to consolidate her scientific profile.
Оригинален текст от CORDIS (на английски).
Участници
- FUNDACIO BARCELONA MEDIA · BARCELONAКоординаторИспания
Връзки
Данни: CORDIS, © Европейски съюз
