NovoFold · De novo protein discovery as a tool for understanding the folding conundrum
Horizon 2020 — Marie Skłodowska-Curie Actions
- Duration
- 2019-01-03 → 2021-06-05
- EU contribution
- €183,455
- Participants
- 1
- Scheme
- MSCA-IF-EF-ST
Lines connect the coordinator with its partners.
Results in brief
De novo protein discovery as a tool for understanding the folding conundrum
The protein folding problem describes the question of how an amino acid sequence of a protein relates to its 3D structure. This conundrum is among the great challenges in biochemistry and has drawn the attention of scientists for decades. Solving the folding problem would not only revolutionize structure prediction but also protein design, which has vast scientific, technological, and medical implications. Most recently, AI and deep learning have revolutionised the prediction of protein structures from amino acid sequences, which has been pioneered by Alpha Fold from Google. The success of this strategy is at least in part due to the availability of large amounts of structural data and sequence data, which is essential for deep learning approaches. This highlights the increasing importance of data driven approaches in biology and the availability of large high quality data sets. The NovoFold project aims at exploring protein sequence space beyond the natural proteome. The natural proteome is a result of natural evolution and therefore limited in diversity. As thus it is not comprehensive of all possible folded amino acid sequences and protein folds. In fact, protein designers have successfully created new folds that are unprecedented in nature. However, designers and their developed algorithms are biased towards recreating variations of existing structures. In contrast, the NovoFold technology enables broad random searches in sequence space, potentially discovering new folds that lie beyond designers’ imagination. Rather than searching for particular bioactivities (e.g. binding, catalysis etc.) the NovoFold assay will experimentally identify sequences, whose primary property is folding. With a throughput of billions to trillions of sequences, the NovoFold platform has the potential to provide further data sets for data driven approaches to solve the protein folding problem. The NovoFold technology interfaces mRNA display with a protein folding sensor, based on the ribosome in conjunction with the arrest peptide SecM. The SecM arrest peptide sequence has been previously used to study protein folding of various proteins, a field pioneered by von Heijne and coworkers. To date, using arrest peptides it has been possible to identify folding intermediates and further analyse co-translational protein folding. Hereby, the ribosome arrests at the last position of the GIRAGP arrest sequence motif, through perturbation of the peptidyl transfer centre. However, if a force is applied to the nascent peptide chain, the ribosome can resume protein synthesis. This force can be applied mechanically or through a protein sequence upstream of the arrest motif which folds within the exit tunnel of the ribosome and exerts a pulling force. Typically, such arrest peptide based folding assays are performed one at a time in in vitro translation extracts. Combining the arrest peptide technology with mRNA display and sequencing would allow to investigate many proteins (and their mutants) in parallel. In mRNA display, a RNA template, which is 3’ covalently linked to puromycin is ribosomally translated in vitro. Once a stop codon close to the end of the template is reached, puromycin is inserted, leading to a covalent linkage between peptide and RNA. The NovoFold technology takes advantage of this, by only linking folded proteins to their cognate mRNA. This is simply achieved by placing a stop codon further downstream of the arrest peptide. Expressing an N-terminal affinity tag, proteins can be panned, leading only to the recovery of cDNA of folded proteins. Using mRNA display it is possible to investigate up to a trillion different sequences. This assay, is very direct, obliterating the use of proteases, which typically have a biased substrate scope and have been previously used in similar assay formats. Apart from identifying de novo proteins and providing large data sets for studying protein folding, this assay could also be applied to the exploration of proteomes beyond the 20 canonical amino acids, towards engineering of xenobiological systems. The objectives of the MSCA action were to implement the technology and show its applicability to random/naïve protein libraries.
Data: CORDIS, © European Union
Project objective
Order is a prerequisite for activity, which is the essence of the sequence-structure-function paradigm of bio-macromolecules. Proteins are a case in point, displaying intricate architectures and impressive functions. Understanding how the primary sequence of these complex macromolecules relates to their 3D structure and vice versa is of fundamental interest and the basis for broad technological exploitation of proteins. The proposed research aims to investigate the folding space of proteins, gaining new insights that will help in solving the folding conundrum. Specifically, the frequency and sequence patterns of ordered polypeptides and their physical and structural properties will be experimentally analysed. Based on these data sets, exploratory data analysis can be performed, providing unprecedented insights into protein folding. Generation of quantifiable high quality data is thus of utmost importance for this project. One of the main objectives of this proposal is therefore the development of an ultrahigh-throughput experimental platform for discovery of de novo proteins with unparalleled precision, using state-of-the-art molecular biology methods. Based on cutting-edge mRNA display technology, I will isolate folded polypeptides from naïve libraries with up to 1E13 members in the size range of 50-100 amino acids. Modern bioinformatic analysis, modelling and biophysical characterisation will be performed to analyse the folding space of proteins and derive new empirical folding rules. Given the increasing importance of protein engineering in fundamental research and industry, the results of this study will be of interest to a wide range of scientists. The fellowship will provide training in state-of-the-art techniques such as statistical data analysis and structural biology. The results of this research project will form the basis for future efforts to exploit the biotechnological potential of newly discovered proteins.
Original text from CORDIS.
Participants
- UNIVERSITY OF BRISTOL · BRISTOLCoordinatorUnited Kingdom
Links
Data: CORDIS, © European Union
