FP7Individual fellowship2008–2010

METAGENOGRIDS · Algorithmics for Metagenomics on Grids

FP7 — People (Marie Curie Actions)

Duration
2008-07-17 → 2010-07-16
EU contribution
€118,572
Participants
1
Scheme
MC-IOF

Lines connect the coordinator with its partners.

Results in brief

Algorithmics for Metagenomics on Grids

The initial objective of this project was to address the problems of metagenomics assembly and metagenomics annotation on grid computing platforms. A preliminary study showed that algorithmic solutions to metagenomics annotation were already known. This result was obtained by changing the point of view on the problem and by showing its similarity with a problem of steady-state optimization of request processing. The three partners involved were not able to initiate a successful collaboration on metagenomics assembly. It was thus decided to refocus the project and to propose algorithmic tools generic for a large range of bioinformatics applications. Virtualization enables to hide most architectural and system characteristics from users. It is therefore appealing to users of distributed computing that are not computer scientists. This is especially true of bioinformaticians who have huge computational needs. We therefore addressed the problem of the efficient use of virtual machines at large scale. More specifically, we propose an alternate solution to batch scheduling for solving the problem of resource allocation management on computational clusters, by taking advantage of the capabilities offered by virtual machines. The aim of this new task was to enhance platform utilization (and thus to more efficiently use these platforms) while ensuring some Quality of Service to users (which is not done in today's systems). After an initial study of the complexity of this problem, we first addressed the problems of resource allocation for applications having constant resource consumption rates and infinite execution times. This unrealistic scenario served as a theoretical framework on which to acquire knowledge and intuitions. We then moved to the study of applications having constant resource consumption rates, but which are submitted over time (online scenario) and whose execution times are unknown (non-clairvoyant scenario). In this context, we designed several job scheduling algorithms. We presented results obtained in simulations for synthetic and real-world High Performance Computing (HPC) workloads, in which we compared our proposed algorithms with standard batch scheduling algorithms. We found that our approach provides drastic performance improvements over batch scheduling in terms of max-stretch minimization and significant improvement for platform utilization. In particular, we identified a few promising algorithms that perform well across most experimental scenarios. Our results demonstrate that virtualization technology coupled with lightweight scheduling strategies affords dramatic improvements in Quality of Service for HPC workloads while using less resources. The obtained results bear great promises to one day be able to design resource management systems delivering significantly better platform utilization than current systems, while ensuring some Quality of Service. To reach such a goal, the previous study should be extended to applications whose resource consumption rates change during the applications' executions. However ambitious this goal is, the foundations that were laid in this study appear to be strong enough to enable to fulfill it. As a side project, we considered bag-of-tasks applications which are especially important in bioinformatics as they encompass applications such as parameter sweep applications. We studied a problem of steady-state optimization of bag-of-tasks application in a probabilistic context (in a non-clairvoyant setting, tasks computation and communication times are defined by distributions). To the best of our knowledge, this work is the first work on steady-state scheduling in a probabilistic context. This work shows that, in such a context, static approaches can still be designed, and still deliver better performance than system-like approaches, even in a non-clairvoyant setting.

Data: CORDIS, © European Union

Project objective

Numerous critical problems in bioinformatics and computational biology are computationally intensive and could benefit from being solved on a computational Grid, that is a network of geographically distributed computers. Most works that combine bioinformatics and Grids focus on biological problems that can be trivially adapted to Grids, staying clear of more complex problems. In this fellowship we propose to address, in an interdisciplinary research, the resolution on Grids of two complex and demanding problems: metagenomics assembly and metagenomics annotation. Metagenomics is the study of the genomes of an entire community of microbes; metagenomics assembly is the problem of assembling the genome of each of the microbial species in the studied community. The applicant, a Grid computing expert, proposes to work with experts in bioinformatics and experts in Grid computing to design new algorithmic solutions to solve these two metagenomics problems. In this proposal, the applicant would simultaneously visit for a year, and work with, Prof. Guylaine Poisson's Bioinformatics Lab (BiL) and Prof. Henri Casanova's Concurrent Research Group (CoRG), both located at the University of Hawai`i at Manoa, USA.

Original text from CORDIS.

Participants

  • INSTITUT NATIONAL DE RECHERCHE EN INFORMATIQUE ET AUTOMATIQUE · Le Chesnay CedexCoordinatorFrance

Links

Data: CORDIS, © European Union