Metagenome binning · Accurate reconstruction of microbial genomes from the environment
„Хоризонт Европа“ — Действия „Мария Склодовска-Кюри“
- Период
- 2023-08-01 → 2025-07-31
- Финансиране от ЕС
- 189 687 €
- Участници
- 1
- Схема
- HORIZON-TMA-MSCA-PF-EF
Линиите свързват координатора с партньорите.
Накратко на български
Метагеномното групиране изследва начини за точно възстановяване на микробни геноми от природни проби, например чрез разграничаване на сходни видове бактерии. По-точните алгоритми помагат за по-доброто идентифициране на редки микроорганизми и правилното разпределяне на генетичните им фрагменти.
Кратко обяснение, генерирано от езиков модел по текста на CORDIS. Оригиналът е по-долу.
Резултати накратко
Accurate reconstruction of microbial genomes from the environment
The primary goal of the proposed project is to improve metagenome binning. It is a computational process in which metagenomic contigs are grouped together based on their presumed genomic origin. State-of-the-art tools perform binning in two stages: (i) computing the distance or similarities in the abundance and k-mer frequency profiles of contigs, and (ii) clustering contigs based on these similarities. Binning tools differ in their similarity measures and clustering algorithms. The current challenges in metagenome binning are i) incorrect binning of conserved regions due to cross mapping of reads to conserved regions, ii) poor recovery of low abundance species and iii) SCMGs-based assessment of intermediate bins which results in too optimistic measures of bin quality. We proposed a new binning algorithm designed to address these issues in three key improvements. First, a linear mixture model is applied to account for cross-mapping of reads. Second, Poisson statistics is applied to effectively process low read counts. Third, the refinement process during clustering using analyses on read counts and k-mer frequencies without applying SCMGs. The algorithm uses a Bayesian theory to derive novel distance measures to identify contigs belonging to the same genome and probabilistic assignment of contigs to genomic bins to improve clustering accuracy and completeness. Overall, the algorithm has several important advantages over the existing methods: i) it promises to be more accurate in binning conserved regions due to the mixture modeling that solves the problem of cross-mappability of reads, ii) it models the read and k-mer count distributions with the appropriate (Poisson) statistics to effectively recover genomes in low abundance and iii) it does not have to rely on single-copy marker genes, permitting an unbiased quality assessment.
Текст от CORDIS, на английски · Данни: CORDIS, © Европейски съюз
Цел на проекта
Metagenome-assembled genomes (MAGs) obtained from metagenomics are of fundamental value to understanding diverse ecological niches of microbes such as the human gut, with applications in medicine, biotechnology, and climate science. However, the quality of MAGs constructed with state-of-the-art tools is often unsatisfactory and worse than the self-reported quality. The main source of error is binning, a computational step that groups sequences assembled from short sequencing reads (contigs) into species-wise bins. The two chief challenges are accurately binning (1) genomes with low abundance and (2) highly conserved regions. Due to cross-mapping of reads, the contigs from conserved regions appear to have abundances equal to the sum of the abundances of the related species or strains. As conventional binning tools all rely on clustering contigs according to their abundances across samples, conserved regions end up forming separate bins. Besides, most existing methods optimise quality measures (purity and completeness based on conserved marker genes) and assess the final quality on these very measures, leading to highly optimistic results. I aim to solve these problems by developing a binning algorithm that applies i) linear mixture models using non-negative matrix factorization to account for cross-mapping,ii) Poisson statistics to accurately model low abundance, and iii) Bayesian statistics-based multinomial clustering to calculate bin numbers. Importantly, it does not require marker gene-based quality measures for binning.By improving the binning of low-abundance and highly conserved contigs, this approach should yield more high-quality MAGs, thereby enhancing a multitude of downstream metagenomic analyses for all areas of microbiome research.
Оригинален текст от CORDIS (на английски).
Участници
- MAX-PLANCK-GESELLSCHAFT ZUR FORDERUNG DER WISSENSCHAFTEN EV · MUNCHENКоординаторГермания
Връзки
- Виж в CORDIS
- DOI: 10.3030/101111457
- https://ec.europa.eu/research/participants/documents/downloadPublic?documentIds=080166e507c5bd42&appId=PPGMS
- https://ec.europa.eu/research/participants/documents/downloadPublic?documentIds=080166e51e66b241&appId=PPGMS
Данни: CORDIS, © Европейски съюз
