H2020Doctoral network2021–2024

ALPACA · ALgorithms for PAngenome Computational Analysis

Horizon 2020 — Marie Skłodowska-Curie Actions

Duration
2021-01-01 → 2024-12-31
EU contribution
€3,725,035
Participants
13
Scheme
MSCA-ITN

Lines connect the coordinator with its partners.

Results in brief

ALgorithms for PAngenome Computational Analysis

Graph-based, instead of sequence-based data structures have decisive benefits with respect to storage, primary analysis, comparison and knowledge extraction when dealing with large, biologically coherent collections of genomes ("pan-genomes"). As a few prominent examples, consider the systematic exploration of the genetic foundations of microbial resistance, the identification of rare diseases, or the complexity of cancer, both on the individual level and on the level of cancer types and subtypes. With genome data rapidly amassing, the urgent need for a shift in computational paradigms, from ordinary sequence-based to graph-based representations of genome collections is no longer deniable: beyond the general acknowledgment of the movement, from which "computational pan-genomics", as a computer science centric area of genomics research emerged, high-impact journals are publishing special genome graph collections thereby recognizing the importance of computer science. The main objective of the project is leading the paradigm shift from sequence- to graph-based representations of genomes. We will provide new graph-based representations of evolutionarily related collections of genomes, together with the computational operations that implement their practical benefits, which is instrumental for leveraging the potential of the big genome data (and preventing serious congestion of resources). We will obtain decisive improvements in terms of 1) Redundancy reduction and data compression, 2) The convenient highlighting of commonalities and differences, 3) Visualization, and 4) Comprehensive annotation. To amplify the benefits of those improvements we will provide software implementations of quality competitive with sequence-based software packages in terms of computational complexity. The schematic Figure 1 highlights a large set of operations for which efficient and reliable computational frameworks and algorithms are necessary. In summary, the ITN will raise a new class of researchers who master the complexity of the era of computational pan-genomics, and thus required to bring along an innovative, unique set of skills: being both highly interdisciplinary and multi-specialized, able to address problems ranging from fundamental algorithms and data structures to software development, big data management and analysis, statistics and machine learning, all closely entangled with genomics and genetics, and bioinformatics in general, and able to bring together the academic and industrial sectors engaged in related business. This explains why interdisciplinary and intersectoral training of ESRs is a compelling necessity for the future development of computational pan-genomics.

Data: CORDIS, © European Union

Project objective

Genomes are strings over the letters A,C,G,T, which represent nucleotides, the building blocks of DNA. In view of ultra-large amounts of genome sequence data emerging from ever more and technologically rapidly advancing genome sequencing devices—in the meantime, amounts of sequencing data accrued are reaching into the exabyte scale—the driving, urgent question is: how can we arrange and analyze these data masses in a formally rigorous, computationally efficient and biomedically rewarding manner?Graph based data structures have been pointed out to have disruptive benefits over traditional sequence based structures when representing pan-genomes, sufficiently large, evolutionarily coherent collections of genomes. This idea has its immediate justification in the laws of genetics: evolutionarily closely related genomes vary only in relatively little amounts of letters, while sharing the majority of their sequence content. Graph-based pan-genome representations that allow to remove redundancies without having to discard individual differences, make utmost sense. In this project, we will put this shift of paradigms—from sequence to graph based representations of genomes—into full effect. As a result, we can expect a wealth of practically relevant advantages, among which arrangement, analysis, compression, integration and exploitation of genome data are the most fundamental points. In addition, we will also open up a significant source of inspiration for computer science itself.For realizing our goals, our network will (i) decisively strengthen and form new ties in the emerging community of computational pan-genomics, (ii) perform research on all relevant frontiers, aiming at significant computational advances at the level of important breakthroughs, and (iii) boost relevant knowledge exchange between academia and industry. Last but not least, in doing so, we will train a new, “paradigm-shift-aware” generation of computational genomics researchers.

Original text from CORDIS.

Participants

  • UNIVERSITAET BIELEFELD · BielefeldCoordinatorGermany
  • CENTRE NATIONAL DE LA RECHERCHE SCIENTIFIQUE CNRS · ParisFrance
  • EUROPEAN MOLECULAR BIOLOGY LABORATORY · HeidelbergGermany
  • GENETON S.R.O. · BratislavaSlovakia
  • HEINRICH-HEINE-UNIVERSITAET DUESSELDORF · DusseldorfGermany
  • HELSINGIN YLIOPISTO · HelsinkiFinland
  • INSTITUT NATIONAL DE RECHERCHE EN INFORMATIQUE ET AUTOMATIQUE · Le Chesnay CedexFrance
  • INSTITUT PASTEUR · ParisFrance
  • STICHTING NEDERLANDSE WETENSCHAPPELIJK ONDERZOEK INSTITUTEN · UtrechtNetherlands
  • THE CHANCELLOR MASTERS AND SCHOLARS OF THE UNIVERSITY OF CAMBRIDGE · CAMBRIDGEUnited Kingdom
  • UNIVERSITA DI PISA · PisaItaly
  • UNIVERSITA' DEGLI STUDI DI MILANO-BICOCCA · MilanoItaly
  • UNIVERZITA KOMENSKEHO V BRATISLAVE · BRATISLAVA 1Slovakia

Links

Data: CORDIS, © European Union