HEIndividual fellowship2022–2024

FairER · Fairness in Language Models: Equally right for the right reasons

Horizon Europe — Marie Skłodowska-Curie Actions

Duration
2022-09-01 → 2024-08-31
EU contribution
€214,934
Participants
1
Scheme
HORIZON-TMA-MSCA-PF-EF

Lines connect the coordinator with its partners.

Results in brief

Fairness in Language Models: Equally right for the right reasons

Large-scale pre-trained language models have revolutionized the field of natural language processing (NLP). They carry great potential and promise to solve many real-life applications such as translation, search, question answering and many more. Recently, a new generation of language models, so called Large Language Models (LLMs) have been developed and released in applications like ChatGPT which reached 100 million active users within 2 months, in comparison Google Translated reached that threshold after 78 months. It is therefore important to get a deeper understanding of how those models work and in particular in what scenarios they might not work as expected. Machine learning-based algorithms, such as LLMs are calibrated and heavily rely on predefined training data to solve particular tasks. In order to generalize well, i.e., to work on a variety if not all possible unseen datasets, a large amount of training data is needed. One of the problems that arises with this training procedure is that the datasets are too big to be curated and no one knows the datasets in detail. Furthermore, decisions made by those algorithms are often dicult to trace back so they mainly function as black boxes and it is not trivial to explain their decisions. Biases such as stereotypes about certain demographics that appear in training datasets will then be forwarded to the models and influence their decisions. It is therefore of critical importance to thoroughly understand those models in order to prevent them from harming certain demographics, often those who already suer from implicit biases in society. Language models not only need to be correct, they need to be “right for the right reasons”. As those models are meant to interact with humans and base their decisions on a human-like reasoning, I argue that in order to understand those models in-depth, it is important to investigate those models further with respect to how well they align with human behaviour regarding dierent demographics. Therefore I want to investigate a) the reasoning behind a decision, i.e., align attention and gradient-based importance attributes by models with human fixation patterns with the help of eye-tracking and b) whether the performance in a task aligns dierently between models and humans for dierent demographics and languages c) explain a model’s decisions with respect to gradient-based explainability methods to further open the blackbox and make models more transparent. This way, we can also trace back a model’s decision to the input which explains what part of the data is responsible for a certain outcome. Finally, this research needs d) to be carried out in a multilingual setting and extended to languages other than English.

Data: CORDIS, © European Union

Project objective

Most of us use technology related to natural language processing (NLP) such as Google Search or virtual assistants in phones and other devices on a daily basis. Large-scale pre-trained language models hereby play a crucial role as they often form the basis of those technologies. Those models are trained on a large amount of training data (e.g. the entire English Wikipedia and the Brown corpus) which makes it impossible to curate the training corpus and potential stereotypes and biases will be implemented into the model, often without researchers noticing. This can lead to problematic and unfair behaviour towards certain demographics, often those who already suffer from implicit biases in society.With FairER, I aim to get a deeper understanding of the inner workings of these language models. In particular, I want to investigate how well their solution strategies align with those of humans and whether this depends on certain demographic attributes such as gender, race, age but also reading abilities and level of education. I will also probe those language models for fairness and inclusiveness, i.e., find out whether the performance of an NLP application depends on demographic attributes of the user. Furthermore, I will conduct this project in a multilingual setting and apply interpretability methods to better understand the rationale behind a models decision.The main impact of FairER will be a better understanding of how language models treat different demographics. These insights will help to improve the fairness and inclusiveness of NLP applications. Furthermore, the datasets I will record and publish along with the code will encourage other researchers to replicate my findings and continue this line of research. Ultimately, this project will have both a scientific and societal impact on the NLP community and users of NLP applications.

Original text from CORDIS.

Participants

  • KOBENHAVNS UNIVERSITET · KOBENHAVNCoordinatorDenmark

Links

Data: CORDIS, © European Union