FP6Реинтеграция2005–2006

MOLECULARCLOCK TOOLS · Development of tools for the analysis of molecular clock and their application for a case of Hepatitis C Virus epidemics with known infection history

6РП — Действия „Мария Кюри“

Период
2005-06-01 → 2006-05-31
Финансиране от ЕС
40 000 €
Участници
1
Схема
ERG

Линиите свързват координатора с партньорите. За проекти отпреди 2014 г. CORDIS не винаги дава точни координати. Тези точки са на ниво град или държава.

Накратко на български

Статистически методи за анализ на молекулярния часов помагат за определяне на датата на заразяване с хепатит С. Тези инструменти подобряват точността при проследяване на епидемии и помагат за решаване на епидемиологични и криминалистични въпроси.

Този кратък обзор е генериран от изкуствен интелект

Кратко обяснение, генерирано от езиков модел по текста на CORDIS. Оригиналът е по-долу.

Резултати накратко

Final Activity Report Summary - MOLECULARCLOCK TOOLS (Development of tools for the analysis of molecular clock and their application for a case of Hepatitis C Virus epidemics with known ... history)

The research concentrated in two areas, namely the analysis of molecular evolution of Hepatits C virus (HCV) and the development of statistical methods to test phylogenetic trees. We used sequences of the HCV E1E2 region to explore methods to estimate the infection events in HCV epidemics. This research was continued beyond the project completion. The method was applied for a very large sequences' data set from almost 300 patients. Our analysis showed that the problems of heterogeneity of evolutionary rates among sites and the presence of selection could be resolved to allow for the use of the molecular clock at a short timescale, provided that the number of calibration sites was sufficient, especially when using the relaxed clock model, which allowed for different mutation rates for different branches in the tree. Using cross-validation, we further showed that the Bayesian method provided highly accurate infection date estimation. This method also offered the opportunity to use molecular clock methods to solve epidemiological and forensic questions. The results demonstrated that methods that avoided the issue of possible positive selection, such as site stripping, did not provide an advantage. However, we also proved that the sites under selection did not carry temporal information. In a parallel analysis we compared the amino acid composition between the different time point samples of one patient. Positions for which this composition was significantly different between the two time points of a single patient were identified. 23 patients were analysed in total. The sequences under analysis also corresponded to the E1E2 region, including the hypervariable region 1 (HVR1) and the HVR2. Interestingly, we identified a third region which presented similar features to both HVR1 and HVR2, previously described in E2 protein, even though the variability degree was slightly lower. This could be explained by the reduced exposition characterising the antigenic site that was included in this new region, namely mAb7/16b, according to a structural model proposed for E2 protein. The new region was termed HVR4. Furthermore, several statistical procedures were proposed to test trees and construct confidence sets of topologies. Unfortunately, in some situations these tests gave contradictory results. In addition, some tests had computational problems. For example, the expected likelihood weights test and Swofford-Olsen-Wadden-Hillis were very intensive computationally, while the generalised least-squares (LS) test required the calculation of the covariance matrix and its inverse might be impossible for data sets that contained many closely related sequences. The weighted LS method differed from the generalised LS in that it treated the distances between taxa as if they were independent to allow for a more efficient calculation of the test statistic. Moreover, we compared the performance of the weighted LS method with the generalised LS and other methods of topology testing that depended on the character data using biological sequences. The first data set we considered was that of mammalian mitochondrial protein sequences. We then proceeded to investigate the size of confidence sets of trees using the eight-taxon data sets constructed from the data sets deposited in the European Molecular Biology Laboratory (EMBL) Align database. We also analysed a data set of a large number of short viral sequences in which testing of alternative phylogenies was vital in including or excluding patients from a hepatitis C outbreak. We showed that the weighted LS method could provide a computationally efficient approximation to the generalised LS statistic, particularly useful in the exploratory analysis of the size of the confidence sets of trees when assessing the phylogenetic signal in the data, and in case other methods were not available. We provided an example of such a data set, i.e. on deoxyribonucleic acid DNA-DNA hybridisation data obtained from four species of sand dollars with sea biscuit as an outgroup. The weighted LS method that was developed as part of the project was applied in several analyses of sequences' evolution from a wide spectrum of organisms, ranging from viruses to fish, and thus far resulted in four manuscripts. One was accepted by the time of the project completion, two were submitted and another two were in preparation. In addition, two papers HCV and two papers on evolution were prepared.

Текст от CORDIS, на английски · Данни: CORDIS, © Европейски съюз

Цел на проекта

The strategic directions of multidisciplinary research at the Institute of Oceanology of the Polish Academy of Sciences include genetics and molecular biology of marine organisms, fields of research where progress increasingly depends on expertise in computational biology. This project will allow making the first step in introducing this much-needed expertise in the Institute and indeed in the region, which we hope will benefit several projects at the Institute (molecular systematics of marine invertebrate s, population genetics of fish and molluscs, marine microbiology). However, considering the short timeline, the utility of tools developed during the project will first be tested on a dataset of Hepatitis C Virus sequences obtained from an epidemic with kn own infection history. These data allow constructing a true phylogenetic tree with extensive temporal information, as the times for both the external nodes and crucial internal nodes are known. HCV displays one the highest rates of evolution detected in an y life form (reaching 10-2 substitutions/site/year. Moreover, extreme heterogeneity of substitution rate is observed in some of the regions of HCV genome. The project will concentrate on resolving the following issues: (1) the level of heterogeneity of the rate of substitution among sites, (2) the use of complex models (for example, models which allow multiple substitution matrices) for possibly better description of the data, (3) the effects of the presence of positive selection on the molecular clock, (4) the nature of amino acid substitutions that occur, especially in the context of the quasispecies hypothesis and the possibility of multiple infections. By analysing sequences coming from a major human pathogen, we hope to provide insight into the molecular biology of HCV and other RNA viruses, relevant not only for molecular evolution, but also for molecular medicine, molecular epidemiology and forensics.

Оригинален текст от CORDIS (на английски).

Участници

  • INSTITUTE OF OCEANOLOGY, POLISH ACADEMY OF SCIENCES · SOPOTКоординаторНиво градПолша

Връзки

Данни: CORDIS, © Европейски съюз