Promoter predictions · Bioinformatic analysis of transcription regulation: a modeling approach
FP7 — People (Marie Curie Actions)
- Duration
- 2011-03-01 → 2015-02-28
- EU contribution
- €100,000
- Participants
- 1
- Scheme
- MC-IRG
Lines connect the coordinator with its partners.
Results in brief
Bioinformatic analysis of transcription regulation: a modeling approach
The main goal of the project was to develop novel bioinformatic methods and generate fundamental understanding necessary for analyzing transcription regulation. Specifically, the project improved methods for TSS predictions, and gained quantitative understanding of the mechanisms of transcription initiation and promoter specificity. The project also modeled dynamics of gene expression regulation, in particular those involved in defense against bacterial viruses and antibiotic (microcine) production. The main results are briefly summarized below: We first systematically determined specificities of the promoter elements outside of the canonical -35 and -10 box – in particular the so called -15 element, which was previously not included in the transcription start site searches. We established that strengths of sigma 70 promoter elements complement each other, so as to achieve a sufficient level of transcription activity; we related this result with a recently proposed `mix-and-match' model of promoter recognition. We used the new alignment to improve the information-theory method for promoter recognition, which significantly (~50%) reduces the number of false positives. We furthermore systematically investigated the importance of different sequence elements in determining the promoter specificity. Furthermore, we investigated kinetic properties of E. coli genomic segments. While we find that RNA polymerase DNA-binding domains are designed to reduce the number of poised promoters, their number is still significant in E. coli intergenic regions, which significantly contributes to false positives in promoter searches. We also found examples of a significant underrepresentation of the regulatory elements in genomic sequences. Based on the notion of the significant binding score deviations from the random ensemble, we developed a new method for detecting direct targets (target genes) of a given regulator, which is based on the Kolmogorov-Smirnov procedure, leading to a significantly higher prediction accuracy. We also developed a new procedure for detecting promoters transcribed by bacteriophage encoded sigma factors (including those from ECF sigma subfamily). Contrary to the current paradigm, by analyzing bacteriophage and canonical (SigmaE and SigmaW) ECF sigmas, we find quantitative and qualitative evidence of strong mix-and-matching in this subfamily. As an example of the dynamics of gene expression, we modeled CRISPR transcript processing, and showed that this system functions as a strong linear amplifier. Based on this analysis, we proposed a synthetic gene circuit [4] that can produce a large amount of product, from small amounts of potentially toxic substrate. As another example of the dynamics of gene regulation, we modeled transcription regulation of mccA and mccB [7], whose products have a crucial role in microcine C (McC) synthesis. Overall, the project resulted in 9 papers that are published in the leading international journals (the average impact factor of 4.1), four papers submitted for publication, and one manuscript in preparation, where the fellow is the first or the senior author on 11 of these papers. In addition, the fellow was recently promoted to an Associate Professor at the Faculty of Biology, University of Belgrade, where he also initiated his group, and is currently mentoring three PhD students (for more details see www.bio.bg.ac.rs/Marko_Djordjevic_web_site/).
Data: CORDIS, © European Union
Project objective
Transcription is both the first step and a major regulatory checkpoint in gene expression. Transcription start sites are locations in genome where RNA polymerase initiates transcription, while transcription binding sites are locations where transcription factors bind to regulate transcription. Knowledge of both transcription start sites, and transcription factor binding sites, is crucial for understanding transcription. However, methods for bioinformatic detection of these sites, which are mainly based on information theory, are typically characterized by low accuracy. Major underlying problems are: i) transcription initiation is a complex process that is characterized by both binding of RNA polymerase and opening of two DNA strands, ii) transcription factor binding sites usually have to be discovered/aligned within longer DNA fragments, which is often technically demanding and unreliable, iii) discovery of direct target genes of a transcription factor is complicated by random occurrence of binding sites that have high binding energy, but are not functional in regulating transcription.The main goal of our proposal is to develop bioinformatic methods for accurate detection of transcription signals. To address the above problems, we will use biophysical modeling to i) Develop a novel method for transcription start site detection in bacteria, which is based on explicit calculation of transcription initiation rates and takes into account both RNA polymerase binding and opening of two DNA strands, ii) Develop a method for inferring transcription factor-DNA interaction parameters directly from DNA fragments selected through high-throughput in-vitro selection experiments, iii) Develop a method for detection of target genes of a transcription factor, which detects an overrepresentation of binding energy distribution upstream of genes. We expect that these methods will significantly improve accuracy of analyzing transcription regulation.
Original text from CORDIS.
Participants
- FACULTY OF BIOLOGY OF THE UNIVERSITY OF BELGRADE · BeogradCoordinatorSerbia
Links
Data: CORDIS, © European Union
