H2020Individual fellowship2020–2022

LanguageOfDNA · Deciphering the Language of DNA to Identify Regulatory Elements and Classify Transcripts Into Functional Classes

Horizon 2020 — Marie Skłodowska-Curie Actions

Duration
2020-06-01 → 2022-09-28
EU contribution
€156,981
Participants
1
Scheme
MSCA-IF

Lines connect the coordinator with its partners.

Results in brief

Deciphering the Language of DNA to Identify Regulatory Elements and Classify Transcripts Into Functional Classes

What is the problem/issue being addressed? We have developed new deep learning language models for genomic sequences based on transformer architecture. These models can be used to learn the structure of DNA, identify the functional elements and predict their function. Moreover, we have collected and curated benchmark datasets (publicly available in genomic_benchmarks GitHub repository) that can be used to judge any new methods developed in the future. We uploaded our neural network models into the public repository Hugging Face Hub. Finally, during our outreach activities, we have taught newcomers to the machine learning field in general and deep learning in particular to finetune pre-trained neural networks to their downstream tasks. Why is it important for society? Understanding DNA functional elements and their function contribute to understanding the molecular mechanisms and can provide insight into the molecular mechanisms underlying disease. More generally, transformer embeddings provide a way to understand DNA on a deeper level. They map categorical variables (sequence of A, C, G, and T) into a continuous space. This can be useful for a variety of tasks, including prediction, clustering, and similarity measurement. Objectives The goal is to provide easy to use neural network models that even people that are not machine learning experts can use to train for a specific task, e.g. classification of genomic sequences or identification of elements in the genome. This project enables me to return from the industry to the research as planned. I also partially covered the cost of the computation (neural network training). Thanks to the collaboration with my supervisor, I was given a chance to develop my soft skills and reach maturity as an independent researcher. As a proof of that I have received my first grant as a principal investigator (PI) that will start in January 2023 and acquired two PhD students (starting in Fall 2022). The fellowship opened the door to CEITEC for me in particular and Brno’s Biology/Genetics community in general (I had never had an academic job in his city before).

Data: CORDIS, © European Union

Project objective

The genomics era dawned about two decades ago with the completion of a multi-billion project sequencing the complete human genome. Today a similar task is within reach of any modestly equipped lab, due to the advances in sequencing techniques. Thousands of new species are now having their genome sequenced per year. A volume of produced genomic data challenges the interpretation capacity of classical statistical methods, opening the doors for novel machine learning approaches.A genomic sequence can be conceptually seen as a close parallel to a human language. Both utilize information (nucleotides/codons and phonemes/syllables) to encode and transmit a signal that can be faithfully decoded, with attention to error minimization, at the receiving end. Genomic messages are a product of multiple and often contradictory evolutionary pressures and are aimed to be decoded at the same time by many different actors in variable ways. For example, a genomic sequence could encode for a protein product, thus displaying a three-nucleotide / codon-based language model. However, it has also subtexts of the regulation (a codon sequence can include motifs aimed at RNA binding proteins), structural information (functional RNA folding patterns pressuring sequences to a specific direction) and so on.The main challenge of applying machine learning models to the identification of genomic function is to find creative ways to untangle these multiple layers of subtexts and focus on each type of message separately. We will adapt algorithms recently developed for the processing of human languages and use them for the classification of RNA transcripts into functional classes and the classification of untranslated functional genomic regions (enhancers, transcription factor binding sites). We will create ready-to-use datasets to benchmark existing and future methods in this field and make all DNA/RNA language models publicly available.

Original text from CORDIS.

Participants

  • Masarykova univerzita · BrnoCoordinatorCzechia

Links

Data: CORDIS, © European Union