PANGAIA · Pan-genome Graph Algorithms and Data Integration
Horizon 2020 — Marie Skłodowska-Curie Actions
- Duration
- 2020-01-01 → 2025-10-31
- EU contribution
- €1,140,800
- Participants
- 15
- Scheme
- MSCA-RISE
Lines connect the coordinator with its partners.
Results in brief
Pan-genome Graph Algorithms and Data Integration
PANGAIA is a network of researchers who are working on laying down the algorithmic foundations of graph pangenomics, a new computational field that aims to renew genome informatics by replacing the notion of linear reference genome with a much richer graph representation. In fact, the traditional view of a human genome as a linear sequence fails to capture human diversity and misses the structure of a population genome, where common and rare genomic variations, such as polymorphisms, duplications, insertions, and deletions are manifest. The notion of graph pangenomes has been introduced to overcome this limitation, effectively facilitating the comparison of thousands (or even millions) of genomes. The main goal of PANGAIA is to leverage the notion of graph pangenomes to address numerous questions arising from the unprecedented availability of sequencing data. This data, for the first time in human history, has the potential to serve as a valuable source of information for advancing human health. In pursuit of this objective, PANGAIA addresses questions such as: How can we produce a better representation of viral genomes? How can we support the investigation of genetic causes of diseases? How can we contribute to understanding human genomic diversity? How can we support research on antibiotic resistance? PANGAIA focuses on three scientific topics: 1) developing methods for constructing pangenomes from vertebrate, viral, and bacterial genomes. 2) Studying measures of similarity and dissimilarities to compare pangenome graphs, including developing new algorithms to compare pangenomes to detect significant genomic variants. 3) Translating the results on the previous topics into actual human health advances.
Data: CORDIS, © European Union
Project objective
Genomes are strings over the letters A,C,G,T, which represent nucleotides, the building blocks of DNA. In view of ultra-large amounts of genome sequence data emerging from ever more and technologically rapidly advancing genome sequencing devices—in the meantime, amounts of sequencing data accrued are reaching into the exabyte scale—the driving, urgent question is: how can we arrange and analyze these data masses in a formally rigorous, computationally efficient and biomedically rewarding manner?Graph based data structures have been pointed out to have disruptive benefits over traditional sequence based structures when representing pan-genomes, sufficiently large, evolutionarily coherent collections of genomes. This idea has its immediate justification in the laws of genetics: evolutionarily closely related genomes vary only in relatively little amounts of letters, while sharing the majority of their sequence content. Graph-based pan-genome representations that allow to remove redundancies without having to discard individual differences, make utmost sense. In this project, we will put this shift of paradigms—from sequence to graph based representations of genomes—into full effect. As a result, we can expect a wealth of practically relevant advantages, among which arrangement, analysis, compression, integration and exploitation of genome data are the most fundamental points. In addition, we will also open up a significant source of inspiration for computer science itself.
Original text from CORDIS.
Participants
- UNIVERSITA' DEGLI STUDI DI MILANO-BICOCCA · MilanoCoordinatorItaly
- CORNELL UNIVERSITY · IthacaUnited States
- GENETON S.R.O. · BratislavaSlovakia
- HUNAN UNIVERSITY · CHANGSHAChina
- ILLUMINA CAMBRIDGE LIMITED · CambridgeUnited Kingdom
- INSTITUT PASTEUR · ParisFrance
- Masarykova univerzita · BrnoCzechia
- NATIONAL UNIVERSITY CORPORATION THE UNIVERSITY OF TOKYO · TOKYOJapan
- STICHTING NEDERLANDSE WETENSCHAPPELIJK ONDERZOEK INSTITUTEN · UtrechtNetherlands
- Simon Fraser University · BurnabyCanada
- THE PENNSYLVANIA STATE UNIVERSITY · University ParkUnited States
- THE REGENTS OF THE UNIVERSITY OF CALIFORNIA · OaklandUnited States
- UNIVERSITA DI PISA · PisaItaly
- UNIVERSITAET BIELEFELD · BielefeldGermany
- UNIVERZITA KOMENSKEHO V BRATISLAVE · BRATISLAVA 1Slovakia
Links
- View on CORDIS
- DOI: 10.3030/872539
- https://ec.europa.eu/research/participants/documents/downloadPublic?documentIds=080166e50ec1b94b&appId=PPGMS
- https://ec.europa.eu/research/participants/documents/downloadPublic?documentIds=080166e511385ecb&appId=PPGMS
- https://ec.europa.eu/research/participants/documents/downloadPublic?documentIds=080166e51138639e&appId=PPGMS
- https://ec.europa.eu/research/participants/documents/downloadPublic?documentIds=080166e51138639f&appId=PPGMS
- https://ec.europa.eu/research/participants/documents/downloadPublic?documentIds=080166e5245286d0&appId=PPGMS
- https://ec.europa.eu/research/participants/documents/downloadPublic?documentIds=080166e5245288fa&appId=PPGMS
- https://ec.europa.eu/research/participants/documents/downloadPublic?documentIds=080166e5cc9a464c&appId=PPGMS
- https://ec.europa.eu/research/participants/documents/downloadPublic?documentIds=080166e5ee35fe7e&appId=PPGMS
- https://ec.europa.eu/research/participants/documents/downloadPublic?documentIds=080166e5ff710736&appId=PPGMS
- https://www.pangenome.eu
Data: CORDIS, © European Union
