DeepTextNet · Deep learning-based text mining for interpretation of omics data
Horizon 2020 — Marie Skłodowska-Curie Actions
- Duration
- 2021-11-01 → 2023-10-31
- EU contribution
- €207,312
- Participants
- 1
- Scheme
- MSCA-IF
Lines connect the coordinator with its partners.
Results in brief
Deep learning-based text mining for interpretation of omics data
In the rapidly evolving field of biomedical research, understanding the intricate network of molecular interactions is crucial. These interactions are the foundation of biological processes and diseases. However, deciphering this complex web has been a significant challenge due to the sheer volume and complexity of available data. This project aimed to tackle this challenge by harnessing the power of state-of-the-art text mining and network analysis techniques. The insights gained from understanding molecular interactions have far-reaching implications for society, particularly in healthcare. By unraveling the details of these interactions, we can better comprehend disease mechanisms, discover potential drug targets, and develop more effective treatments. This project's outcomes are not just a leap forward for scientific understanding but also a step towards improving health outcomes and the quality of life. The project had two primary goals. The first was to develop a next-generation text-mining technology using deep neural networks to extract molecular interactions from biomedical literature. This technology aims to transform how we gather and interpret complex biological data, making it more efficient and comprehensive. The second goal was to create advanced molecular networks and develop a novel methodology for integrated network analysis on large-scale omics data. This approach was intended to provide a deeper understanding of the molecular interactions and their implications in biological systems, thereby enhancing our ability to analyze and interpret vast amounts of omics data. Throughout the project, I focused on achieving these objectives while overcoming various challenges and adapting to unforeseen circumstances. The results have not only advanced the state-of-the-art in bioinformatics but also set the stage for future innovations in understanding and utilizing molecular interaction networks.
Data: CORDIS, © European Union
Project objective
The academic community and the pharmaceutical industry use omics technologies to produce big data at an incredibly increasing rate but are faced with major challenges when it comes to their interpretation. Key for this interpretation is the association between individual entities, which in a biological context means creating molecular networks. These associations cannot be derived from the omics data alone, but rely heavily on pre-generated networks created by text mining of millions of scientific articles. One of the most popular sources of such networks is the STRING database, which currently serves ~100,000 users monthly.Many of these users work with omics data and a major obstacle, which limits potential benefits for them, is that literature-derived networks are made up of ""functional associations"", stating only that two molecules do something together, but neither the interaction type nor the direction. Hence, our hypothesis is that state-of-the-art computational approaches will be able to exploit new possibilities in network biology that emerge from big data. The key objective of DeepTextNet is to extract novel information from the biomedical literature on the type and direction of gene/protein associations. Specifically, a new paradigm will be realized by building a next generation text mining technology for relation extraction of molecular interactions that explicitly utilizes deep learning and, in contrast to current methodology, makes use of big data for training as opposed to small manually curated datasets. This new strategy for obtaining comprehensive molecular networks with both type and direction for the interactions is precisely what is currently missing for the interpretation of omics data. We expect the impact to be high and wide, as on top of applying this strategy on omics datasets as part of the project, the new technology will feed directly into STRING, which is used globally and integrated into workflows in both academia and industry.""
Original text from CORDIS.
Participants
- KOBENHAVNS UNIVERSITET · KOBENHAVNCoordinatorDenmark
Links
Data: CORDIS, © European Union
