H2020Individual fellowship2018–2020

DeepRNA · Discovering functional protein-RNA interactions through data integration and machine learning.

Horizon 2020 — Marie Skłodowska-Curie Actions

Duration
2018-03-01 → 2020-02-29
EU contribution
€158,122
Participants
1
Scheme
MSCA-IF-EF-ST

Lines connect the coordinator with its partners.

Results in brief

Discovering functional protein-RNA interactions through data integration and machine learning.

RNA-binding proteins are implicated across a wide spectrum of human genetic disorders, with molecular mechanisms ranging from aggregation of proteins and RNAs to defects in splicing, localisation and translation. Examples include heterogeneous and life-threatening genetic disorders such as Diamond-Blackfan anaemia, retinitis pigmentosa, spinocerebellar ataxia, and amyotrophic lateral sclerosis (ALS) among others. The DeepRNA project targeted genetic disease predisposition via disease-associated variants in the human transcriptome and was enabled by recent data on expression quantitative trait loci (eQTLs) and experimentally determined RNA-protein and RNA-RNA interactions. These data were complemented with high-quality protein-RNA interaction predictions carried out in the host group, which had a strong track record in computing and validating ribonucleoprotein associations. The project has expanded the human protein–RNA interactome in a genome-wide manner beyond experimental data, which is available for only 352 of the 1,542 recently described RNA-binding proteins. The information in this extended interactome should contribute to progress towards precision medicine. The deliverables of DeepRNA were designed to be of direct use in the clinical assessment of potentially pathogenic genomic variants. It is my hope that this will directly help to improve the diagnostics performance and value of clinical analysis products, as well as delivering wider inspiration for artificial intelligence applications in biology and RNA-protein interaction network research, increasing the competitiveness of European research and innovation in these fields. A short international secondment at a leading and highly innovative genomic machine learning group, the Kundaje lab at Stanford University (USA), also gave me an opportunity to disseminate the results of my project and to acquire detailed knowledge on advanced machine learning methods applicable to genomic data. The key deliverable of the project, RNAct, a functionally annotated comprehensive reconstruction of the human RNA-protein interaction network rooted in the authoritative Ensembl and UniProt resources, is intended to be of long-term usefulness to researchers across sectors and disciplines and to complement the already excellent profile of European public resources in genomics. It has been published in Nucleic Acids Research and is accessible at https://rnact.crg.eu. It is now fully integrated as an external database in UniProt, the authoritative protein information database.

Data: CORDIS, © European Union

Project objective

RNA-binding proteins are implicated across a wide spectrum of human genetic disorders, with molecular mechanisms ranging from aggregation of proteins and RNAs to defects in splicing and translation. Examples include heterogeneous and life-threatening genetic disorders such as Diamond-Blackfan anaemia, spinocerebellar ataxia and amyotrophic lateral sclerosis (ALS) among others.The DeepRNA project targets genetic diseases via disease-associated variants in the human transcriptome and is enabled by recent data on expression quantitative trait loci (eQTLs) and experimentally determined RNA-protein and RNA-RNA interactions. The data will be complemented with high-quality RNA-protein interaction predictions carried out in the host group that has a strong track record in computing and validating RNA-protein associations.To my knowledge the use of eQTL variants to study RNA-protein interactions is a novel approach and is useful for developing new tools for personalised medicine. My approach will expand the human interactome in a genome-wide manner beyond experimental data, which is currently available for only 352 of the 1,542 recently described RNA-binding proteins. Complementing experimentally determined interactions with predictions will allow me to expand my analyses of the human interactome to the genomic scale while maintaining accuracy, by using the experimental dataset as a gold standard. I will employ methods such as graph partitioning and graph neural network encoding to rationalise the effects of disease-associated variants on the human interaction network, and make quantitative predictions of polymorphisms associated with genetic diseases, thereby aiding personalised medicine.I am confident that this fellowship will equip me with the domain knowledge, independence and transferrable skills to confidently and creatively build an interdisciplinary, globally recognised research team within Europe that will focus on medically relevant human signalling systems.

Original text from CORDIS.

Participants

  • FUNDACIO CENTRE DE REGULACIO GENOMICA · BarcelonaCoordinatorSpain

Links

Data: CORDIS, © European Union