HEИндивидуална стипендия2024–2026

PhLex · Phylogenetic Lexification: Patterns of meaning through time

„Хоризонт Европа“ — Действия „Мария Склодовска-Кюри“

Период
2024-06-01 → 2026-05-31
Финансиране от ЕС
195 915 €
Участници
2
Схема
HORIZON-TMA-MSCA-PF-EF

Линиите свързват координатора с партньорите.

Накратко на български

Смисловите модели в езиците се изследват, за да се види дали промените в значенията на думите следват предвидими пътища във времето. Това помага за анализа на застрашени езици, за които съществуват само кратки списъци с думи.

Този кратък обзор е генериран от изкуствен интелект

Кратко обяснение, генерирано от езиков модел по текста на CORDIS. Оригиналът е по-долу.

Резултати накратко

Phylogenetic Lexification: Patterns of meaning through time

Languages encode knowledge, culture, and identity, yet more than half of the world’s languages are currently endangered, and many are represented only through short, fragmentary wordlists. These sparse records restrict our ability to understand how languages evolve, how meanings change, and how linguistic diversity develops over time. Traditional historical linguistic research relies heavily on large, well-documented datasets and on phonological correspondences, meaning that the world’s most under-documented and often Indigenous languages are systematically excluded from global comparative research. This creates an uneven evidence base and limits the development of methods capable of supporting communities working to document or revitalise their languages. The Phylogenetic Lexification (PhLex) project was established to address these challenges by exploring whether patterns of meaning (rather than only word forms) can provide reliable signals of linguistic history. While meaning has long been considered too variable or unpredictable to be used systematically, new research suggests that semantic structures may persist across many generations of speakers, even when the words themselves change. If semantic patterns can be shown to carry historical information, they offer a promising and inclusive pathway for analysing languages for which extensive documentation is not available. PhLex advances this idea by developing two complementary data resources: EvoLex, a comparative lexical dataset that collates cognate sets across a chosen family of related languages, and EvoSem, a semantic tool that models patterns of meaning alignment, colexification, and semantic divergence. Together, these resources enable large-scale investigation of how meanings shift and how such changes relate to known historical relationships among languages. Although the proposal originally focused on Oceanic languages, the project pivoted to Australian Indigenous languages when it became clear that these offered a feasible and ethically appropriate dataset and stronger opportunities for collaboration. This change did not alter the scientific objectives, but it significantly increased the societal relevance of the work by contributing to a region where documentation gaps are particularly acute. The overarching objective of PhLex is to determine whether meaning-based patterns, identified through lexification and re-lexification, retain a detectable phylogenetic signal. To achieve this, the project pursued three linked goals: 1. Build a comprehensive database of meaning correspondences across a set of related languages (EvoLex/EvoSem). 2. Use phylogenetic techniques to measure the strength of the historical signal in these semantic patterns. 3. Compare a meaning-based model with a traditional cognate-based approach, and assess whether a combined, hybrid approach can enhance historical inference. These objectives respond to a broader strategic need within European and international research: developing computational and data-driven methods capable of handling incomplete or heterogeneous datasets. Such methods support Open Science goals by increasing the reusability of linguistic data and by extending analytical tools to cases where conventional methods cannot be applied. They also align with the EU’s wider societal priorities concerning cultural heritage, Indigenous knowledge, and equitable access to scientific infrastructure. By focusing on a historically under-studied language family, PhLex tackles both methodological and societal challenges. Scientifically, it tests the evolution of meaning through computational historical linguistics. Societally, it develops resources that may support Indigenous language centres and researchers by improving the accessibility and comparability of lexical information, always following ethical and community-informed protocols for handling sensitive materials. While the project is not expected to produce direct economic outcomes, its cultural and research impacts are potentially significant: improving the representation of under-documented languages in global comparative work and enabling new forms of historical and typological research. The expected impact of PhLex is therefore twofold. First, it contributes directly to the scientific understanding of how languages diversify by introducing semantic data into phylogenetic modelling. Second, it lays the foundations for future comparative work on languages that lack extensive documentation, thereby expanding the inclusiveness and reach of linguistic science. As analyses continue beyond the funded period, the methods and datasets established through the project will support further publications, ongoing collaborations, and long-term contributions to the study of language change worldwide.

Текст от CORDIS, на английски · Данни: CORDIS, © Европейски съюз

Цел на проекта

In historical linguistics, scholars often grapple with limited data. Nearly half of the world's ~7,000 languages are endangered. Many of these languages are under-documented, often reduced to mere short wordlists. This is a pressing challenge, as the scarcity hampers our ability to understand linguistic evolution and diversity. Traditional methods often rely on the form of linguistic units, such as phonemes or morphemes. The PhLex project challenges this by emphasising that semantics, while less explored, is not intrinsically unyielding to analysis. PhLex aims to innovate by incorporating semantic analysis into traditional methods, enhancing both our grasp of linguistic evolution, and thus better utilise under-documented language data. To validate this methodology, the PhLex project will focus on the Oceanic language family, spoken across Polynesia, Melanesia, and Micronesia. The first phase involves creating a comprehensive database that explores how these languages articulate concepts related to the semantic domain of perception, like seeing, hearing, and feeling—an area with notable cross-linguistic variation. Subsequent analysis using advanced statistical models traces the historical evolution of the semantic structures, thus applying ‘phylogenetics’ to ‘lexification’, allowing for a comparative assessment against traditional cognate-only methods. The PhLex project’s ultimate objective is to enhance historical linguistic methods by testing a hybrid model that integrates semantic and cognate-based approaches. This model aims to optimise the use of limited linguistic data to provide a nuanced understanding of language origins and evolution. The project thereby opens a new avenue for determining genetic relatedness between languages, an essential factor in understanding linguistic diversification and aiding language preservation efforts in the face of under-documentation.

Оригинален текст от CORDIS (на английски).

Участници

  • ECOLE NORMALE SUPERIEURE · ParisКоординаторФранция
  • MAX-PLANCK-GESELLSCHAFT ZUR FORDERUNG DER WISSENSCHAFTEN EV · MUNCHENГермания

Връзки

Данни: CORDIS, © Европейски съюз