MOTIF VECTORS · Hierarchical Motif Vectors for Protein Alignment and Functional Classification
FP7 — People (Marie Curie Actions)
- Duration
- 2009-11-11 → 2012-11-10
- EU contribution
- €75,000
- Participants
- 1
- Scheme
- MC-IRG
Lines connect the coordinator with its partners.
Results in brief
Hierarchical Motif Vectors for Protein Alignment and Functional Classification
The two principal tasks of computational analysis of proteins based on their amino acid sequences are the determination of proteins related to one another either by function or by evolutionary development, and the identification of specific amino acid combinations, or motifs, that determine the protein function. Both of these two tasks have largely been addressed by matching amino acid sequences across different proteins. Sequence alignment algorithms compute similarity scores between the amino acid sequences of different proteins that are then used to identify protein subgroups with high within-group similarity. Amino acid combinations observed consistently among proteins of a specific functional subgroup constitute the sequence motifs of the related subgroup, the presence of which indicates membership of the protein in that functional group. The primary challenge in the computational analysis of amino acid sequences is the combinatorial complexity inherent in representing amino acid sequences as words composed of letters from an alphabet of twenty, with each letter corresponding to a different amino acid. Faced with the daunting prospect of evaluating potentially millions of possible amino acid combinations for functional specificity, we introduce a numerical alternative that characterizes protein structure from amino acid sequences via numerical means using techniques from multi-scale signal decomposition and statistical learning. The proposed framework is based on a notion of hierarchical motif vectors that capture the numerical variation of the local physico-chemical composition along a protein’s amino acid sequence. This allows using an extensive library of vector space data processing methods for rigorously computing the similarity of the corresponding amino acid sequence motifs, both in the alignment of amino acid sequences as well as the identification of motifs specific to functional protein groups. This project starts with developing global and local alignment methods for sequences of motif vectors to establish correspondence between the underlying amino acid sequences. Next, it identifies hierarchical motif vectors that possess functional or structural specificity in select protein groups via quasi-supervised statistical learning. Finally, it formulates a protein function recogniton strategy based on group-specific hierarchical motif vectors. The experimental results on local as well as global motif vector alignment indicate that the motif vectors adequately characterize the physico-chemical composition along amino acid sequences and allow associating segments sharing similar amino acid configurations at short, mid and long range neighborhoods along their respective sequences. This allows establishing associations between amino acid sequence segments that share similar functions due to amino acid configurations that share similar their physico-chemical properties. Results on the prediction of N-glycosylation at consensus sequence sites also confirm that the hierarchical motif vectors accurately characterize the physico-chemical configurations at and around amino acid sites for functional significance. Furthermore, the quasi-supervised learning strategy can sort through the prospective sites of activity and identify the ones with real functional potential based on their respective motif vectors. The quasi-supervised learning strategy is especially fitting to biomedical information processing tasks where a relatively small collection of instances are available with experimentally verified attributes against the backdrop of a very large number of unknown prospects. The quasi-supervised learning algorithm successfully separates the probable prospects from the unlikely ones automatically with no user intervention.
Data: CORDIS, © European Union
Project objective
This proposal introduces hierarchical motif vectors for numerical analysis of sequence motifs, and develops a novel framework for alignment and functional classification of proteins. Hierarchical motif vectors will be computed using multi-scale decompositions of property sequences obtained by converting amino acid sequences into numeric sequences of various amino acid properties. These hierarchical motif vectors will capture the variations of amino acid properties in the vicinity of each amino acid in the sequence of a given protein. We will develop alignment algorithms for amino acid sequences that match their hierarchical motif vectors. We will also use unsupervised statistical learning algorithms to identify hierarchical motif vectors specific to functional protein groups, notably the antigen binding proteins, transcription factors, growth factors, and glycosylation proteins. We will then apply these methods to protein classification, using the overlap scores from the hierarchical motif vector-based sequence alignment as well as the presence and extent of hierarchical motif vectors specific to the protein group in consideration. We will validate all methods developed in this project against existing sequence alignment, motif detection, and protein classification algorithms in the literature. Among the innovations of the project is the use of hierarchical motif vectors for characterization of local physico-chemical variations along an amino acid sequence. This allows analyzing sequence motifs by general machine learning methods via the embedded vector space arrangement. Next, sequence alignment can be tuned to different amino acid properties at various scales, improving the potential for sequence alignment-based protein similarity in functional classification. Furthermore, group-specific hierarchical motif vectors will be identified as those that occur exclusively among the members of a protein group, increasing their likelihood of bearing functional specificity.
Original text from CORDIS.
Participants
- IZMIR INSTITUTE OF TECHNOLOGY · İzmirCoordinatorTürkiye
Links
Data: CORDIS, © European Union
