H2020Individual fellowship2019–2022

IMAGINE · Informing Multi-modal lAnguage Generation wIth world kNowledgE

Horizon 2020 — Marie Skłodowska-Curie Actions

Duration
2019-06-14 → 2022-04-23
EU contribution
€232,394
Participants
2
Scheme
MSCA-IF

Lines connect the coordinator with its partners.

Results in brief

IMAGINE – Informing Multi-modal lAnguage Generation wIth world kNowledgE

Recently we have witnessed unprecedented improvements in the quality of computational models of human language, including models that use both images and text. Such models can, for instance, answer different questions about the content of images, i.e., Visual Question Answering (VQA). Despite apparently being able to solve complicated tasks like VQA or visual commonsense reasoning (VCR), we do not know the extent of the capabilities of these vision and language (V&L) models. Moreover, V&L models have recently been found to suffer from drawbacks such as generalising poorly on unseen or rare cases. For example, VQA models often learn to answer questions starting with "How many ..." with the number "2", since in its training data this is the most frequent answer to this type of question. When used in practice, these models' poor generalisation capabilities can also lead to different types of biased predictions or to the unfair representation of minority groups, for instance. An important reason these models do not generalise well is the fact that the data they are trained on is biased, and they cannot efficiently "understand" and utilise human-curated knowledge present in structured knowledge graphs. In the IMAGINE project my main goal is to incorporate world knowledge to better learn state-of-the-art V&L models so that they better generalise to unseen or rare cases, and also so that we can mitigate issues related to bias and unfairness. I investigate ways to make V&L models transparently connect to knowledge graphs, and whether that leads to less bias and better generalisation. I also currently devise better datasets to train vision & language models, and better benchmarks to evaluate these models, as orthogonal strategies to gauge their capabilities.

Data: CORDIS, © European Union

Project objective

Deep neural networks have caused lasting change in the fields of natural language processing and computer vision. More recently, much effort has been directed towards devising machine learning models that bridge the gap between vision and language (V&L). In IMAGINE, I propose to lead this even further and to integrate world knowledge into natural language generation models of V&L. Such knowledge is easily taken for granted and is necessary to perform even simple human-like reasoning tasks. For example, in order to properly answer the question “What are the children doing?” about an image which shows parents with children playing in a park, a model should be able to (a) tell children from parents (e.g. children are considerably shorter), and infer that (b) because they are in a park, laughing, and with other children, they are very likely playing.Much of this knowledge is presently available in large-scale machine-friendly multi-modal knowledge bases (KBs) and I will leverage these to improve multiple natural language generation (NLG) tasks that require human-like reasoning abilities. I will investigate (i) methods to learn representations for KBs that incorporate text and images, as well as (ii) methods to incorporate these KB representations to improve multiple NLG tasks that reason upon V&L. In (i) I will research how to train a model that learns KB representations (e.g. learning that children are young adults and likely do not work) jointly with the component that understands the image content (e.g. identifies people, animals, objects and events in an image). In (ii) I will investigate how to jointly train NLG models for multiple tasks together with the KB entity linking, so that these models benefit from one another by sharing parameters (e.g. a model that answers questions about an image benefits from the training data of a model that describes the contents of an image), and also benefit from the world knowledge representations in the KB.

Original text from CORDIS.

Participants

Links

Data: CORDIS, © European Union