EvoDCode · Deep representational learning of the evolutionary DNA code in the vertebrate pallium
Horizon Europe — Marie Skłodowska-Curie Actions
- Duration
- 2025-04-01 → 2027-03-31
- EU contribution
- €200,400
- Participants
- 1
- Scheme
- HORIZON-TMA-MSCA-PF-EF
Lines connect the coordinator with its partners.
Results in brief
Deep representational learning of the evolutionary DNA code in the vertebrate pallium
Vertebrate genomes generally consist of billions of nucleotides and encode tens of thousands of genes. Yet, while these genes could theoretically be expressed in almost any combination, evolution has resulted in a large, but finite number of stable configurations with distinct downstream functions. These configurations form the basis of cell types, groups of cells that share a core identity. However, our understanding of how an organisms’ cell types are encoded in the genome is still lacking. While genes themselves can be homologous between species to a certain extent, the regulatory DNA is much less conserved. The expression of genes is regulated by enhancer elements that bind transcription factors, which can be located far away from the target gene. To investigate the regulatory logic that encodes cell types in the genome, a systematic approach must be taken towards sequence phylogeny, where we identify which sequences are active in each of the different types of cells across a wide variety of species, before asking what the defining features of these sequences are. The rapidly advancing field of artificial intelligence holds great promise for comparative genomics and Convolutional Neural Networks (CNNs) and DNA language models have already been successfully used to model gene regulatory logic in an interpretable way by revealing the transcription factor binding sites within cis-regulatory elements and their co-regulatory relationships. These models require large amounts of training data, and the introduction of the single-cell Assay for Transposase Accessible Chromatin sequencing (scATAC-seq) has enabled us to collect the large amounts of data stratified by cell type, which are required to train these artificial intelligence models. The overarching goal of this proposal is to better understand how the genome sequence underlies cell identity in the pallium across vertebrate species. I hypothesize that gene regulatory logic can be directly learned from the genomic sequence and used to model and predict cell types. I specifically focus on the pallium, a part of the brain that is strongly tied to species-specific behaviour and underwent strong divergent evolution. In humans the dorsal pallium is expanded into the cerebral cortex, providing most of our expanded cortical abilities, while the avian dorsal pallium consists of only a single cortical layer (Wulst). Our understanding of how these large differences came to be is still limited. Using modern single-cell epigenomic methods we can study how evolutionary changes impact gene regulation by sampling across a wide set of vertebrate species and using this data to model cell type evolution.
Data: CORDIS, © European Union
Project objective
Vertebrate genomes consist of tens of thousands of genes and yet they produce only a limited number of stable cell types. Our understanding of how cell types are encoded in the genome is still lacking. The expression of genes is controlled by gene regulatory elements, which are evolutionarily much less conserved than genes are and can be located far away in the genome. My main goal is to better understand how the genomic sequence underlies cell identity in the pallium across vertebrate species and I hypothesize that gene regulatory logic can be directly learned from the genomic sequence and used to predict cell types. Recent advances in the field of single cell sequencing have made the generation of large epigenomic datasets possible, while novel machine learning models like DNA language models are providing unprecedented insights into the regulatory logic of the genome. In addition, for many non-model species high quality reference genomes are becoming available as there is an increased awareness for the need to preserve biodiversity. I will conduct single cell multiome sequencing, supplemented with low-cost single-cell ATAC-seq using the HyDrop platform - developed in the host-lab - to profile regulatory elements across cell types in the pallium from multiple species, including mammals, birds, lizards and fish. My experience in generating and annotating single cell data from the brain will aid in the analysis and alignment of the generated datasets. Next, I will use this data to study the relationships between the genome of a species and the identified cell types by training species-aware DNA language models and using these models to identify species-specific and conserved regulatory logic. These models will not only be interesting from an evolutionary perspective, but also aid in ongoing efforts to develop synthetic enhancers used to target highly specific cell types, which can be applied in many therapeutic applications.
Original text from CORDIS.
Participants
- VIB VZW · ZWIJNAARDE - GENTCoordinatorBelgium
Links
- View on CORDIS
- DOI: 10.3030/101202429
- https://ec.europa.eu/research/participants/documents/downloadPublic?documentIds=080166e5225c9c2c&appId=PPGMS
Data: CORDIS, © European Union
