H2020Individual fellowship2015–2017

WFL · Morphology beyond inflection. Building a wordformation based dictionary for Latin

Horizon 2020 — Marie Skłodowska-Curie Actions

Duration
2015-11-01 → 2017-10-31
EU contribution
€180,277
Participants
1
Scheme
MSCA-IF-EF-ST

Lines connect the coordinator with its partners.

Results in brief

Morphology beyond inflection. Building a wordformation based dictionary for Latin

In the past two decades there has been a considerable increase in the creation of computational linguistic resources for the investigation of classical languages, which have updated the state of the art almost to the same level as that of the resources currently available for modern languages. However, among the existing linguistic resources, we currently lack, for Latin (and Ancient Greek - and indeed even for the majority of modern languages), a morphological derivational dictionary that connects lexical elements on the basis of Word Formation Rules (WFRs). In linguistics, there are two kinds of morphological rules: 1. inflectional, which relate to different forms of the same lexeme (i.e. singular vs. plural, present vs. past tense); 2. word formation rules, which relate to different lexemes, e.g. love vs. lover. There are two types of word formation processes in Latin: derivation and compounding. Derivation can be further split into: 1. Affixation, where one or more morphemes, called affixes, can be attached to the base of a word. Affixation can be of two types, and can involve (or not) a change of part of speech: a. Prefixation: where the affix is attached before the base. b. Suffixation: where the affix is attached after the base. 2. Conversion, where the derived word incurs only in a change of part of speech without the addition of any affix. Compounding is the formation of a new lexeme from two or more lexemes. The WFL project has consisted in the compilation of a derivational morphological lexicon of the Latin language, Word Formation Latin (WFL), which connects lexical elements on the basis of word formation rules (WFRs), through the use of computational linguistic methods. The resulting lexicon has been integrated into the most recent version of the morphological analyser and lemmatiser for Latin Lemlat (www.lemlat3.eu), and can be browsed in its own dedicated website at http://wfl.marginalia.it. Enriching textual data with derivational morphology tagging promises to provide strong outcomes. It can organise the lexicon at higher level than words, by building word formation based sets of lexemes sharing a common ancestor. Moreover, information on word formation can act as a bridge between morphology and semantics, since core semantic properties are shared at different levels by words built by a common word formation process. The scope of WFL is to assign a WFR to each morphologically-complex lexeme (i.e. one word morphologically derived from another word) and to link each complex lexeme to its ancestor. All those lexemes that share a common (not derived) ancestor belong tothe same “word formation family”. For instance, the noun bellatrix ‘she who wages war’, the verb rebello ‘to revolt, rebel’, and the adjective bellicosus ‘fond of war’ all belong to the word formation family whose ancestor is noun bellum ‘war’. The semi-automatic insertion of lemmas into the WFL database establishes input-output relations for a set of lemmas matching the features that characterise each WFR.

Data: CORDIS, © European Union

Project objective

The proposed project aims at the compilation of a derivational morphological dictionary of the Latin language, which connects lexical elements on the basis of wordformation rules, through the use of computational linguistic methods. The final resource will be both a standalone tool accessible through its own website, and interconnected with the Index Thomisticus Treebank (IT-TB), a syntactically annotated corpus of texts of Thomas Aquinas, currently based at the Centro Interdisciplinare di Ricerche per la Computerizzazione dei Segni dell’Espressione (CIRCSE), at the Universita’ Cattolica del Sacro Cuore in Milan. The integration with the IT-TB will be operated through the embedding of the dictionary data within the morphological layer of annotation of the treebank, using TEI (Text Encoding Initiative) P5 conformant XML encoding which will ensure easy sharing and linking of the data across a variety of potentially related projects, and will make sure, through the use of an internationally recognised standard for the encoding of textual data, such data will be re-usable and expandable in the future

Original text from CORDIS.

Participants

  • UNIVERSITA CATTOLICA DEL SACRO CUORE · MILANOCoordinatorItaly

Links

Data: CORDIS, © European Union