H2020Individual fellowship2017–2019

ThReDS · A Theory of Reference for Distributional Semantics

Horizon 2020 — Marie Skłodowska-Curie Actions

Duration
2017-07-01 → 2019-06-30
EU contribution
€158,122
Participants
1
Scheme
MSCA-IF

Lines connect the coordinator with its partners.

Results in brief

A Theory of Reference for Distributional Semantics

The overarching goal of ThReDS was to build a computer system that can 'refer', i.e. generate descriptions of concepts or entities in a way that uniquely identifies them for a human (e.g. Harry Potter = "the wizard with the round glasses", Hermione = "the best student at Hogwarts", etc). Having such a system would let us understand better how humans communicate with each other, and help us build artificial agents that can converse with us. Before an artificial agent can talk, it needs to learn about the world, just as a child would do. In linguistic and computational terms, this means acquiring *representations* of the things the agent is exposed to. In the field of Distributional Semantics, such computational representations have traditionally been built from raw text data (sometimes enriched with visual information) and take the form of a 'vector', that is, a mathematical model of the way a particular word is used by human beings, as experienced by the agent. Such vectors can be found in many everyday applications like search engines, recommendation systems and conversational agents. So far, however, they have only been constructed for *concepts* (e.g. 'student', 'owl', 'broom') rather than individual entities ('Harry Potter', 'Hedwig', 'Harry's Nimbus 2000'). This is because current algorithms need considerable amounts of data to learn properly, and references to individual entities are much less frequent in raw text than generic occurrences of words. Further, those raw vector representations are not suitable to refer from, because they do not explicitly encapsulate the properties of the concept or individual that a human would use to identify them (e.g. 'wearing glasses' for Harry Potter). In order to make vectors compatible with so-called 'Referring Expression Generation' systems, that is, algorithms that can produce successful references to things in the world, a translation must be found to a more formal and structured representation of meaning, which in theoretical linguistics has its incarnation in 'Model-theoretic Semantics'. ThReDS tackled two challenges: a) the computational extraction of representations of entities from raw text, concentrating on the small data issue; b) the theoretical account of how raw exposure to linguistic data (distributional semantics information) can shape the agent's representation of the world (their model-theoretic semantics).

Data: CORDIS, © European Union

Project objective

One of the most fundamental human faculties is 'reference': the capacity to 'talk about things'. This extraordinary ability is at the core of many forms of human exchange, from asking for the salt at the dinner table to collaboratively building a solar probe. It involves using linguistic signs to identify things in the world and bring them to the mind of another. Reference is poorly understood: in particular, we do not know how humans build a shared linguistic representation of their environment, which they use to link words and world. My goal is to build a computational model of the way people acquire world knowledge from language and translate knowledge back into language. My overall framework includes three steps: 1) creating a representation of the way people 'talk about things', using distributional semantics (DS: a computational approach to modelling word usage); 2) automatically mapping the distributional model onto a partial set-theoretic model (a formal knowledge representation expressing shared beliefs about the world); 3) using the set-theoretic model to generate unobserved linguistic expressions which refer. The pipeline will be evaluated via an online game where a computer has to produce references to well-known concepts and individuals for a human tester. This work will significantly advance the state-of-the-art in linguistics: while DS has enjoyed considerable success in modelling lexical phenomena, it is currently showing its limits in explaining referential aspects of meaning. Conversely, referential semantics is still far from fully explaining the cognitive aspects of concept acquisition and reuse. The proposed investigation requires a very novel integration of computational semantics (my area of expertise) and formal linguistics (in which my host is an internationally recognised expert). The collaboration will give us the chance to lead a burgeoning area of research aiming at integrating reference into DS.

Original text from CORDIS.

Participants

  • UNIVERSIDAD POMPEU FABRA · BarcelonaCoordinatorSpain

Links

Data: CORDIS, © European Union