DAMA · Extreme-Scale Data Management
Horizon 2020 — Marie Skłodowska-Curie Actions
- Duration
- 2018-11-01 → 2020-10-31
- EU contribution
- €185,076
- Participants
- 1
- Scheme
- MSCA-IF
Lines connect the coordinator with its partners.
Results in brief
Extreme-Scale Data Management
Resources from large-scale high-performance computing (HPC) platforms or large data centers are shared between concurrent applications. Users submit jobs to a batch scheduler, and this resource manager assigns computing nodes to applications according to their requested processing power. In these architectures, the access to persistent data happens through a shared I/O infrastructure including a parallel file system (PFS) deployment over a set of dedicated servers. Differently from processing power, and despite being shared, this data access resource is NOT arbitrated. Hence each application, together with used I/O libraries, will work to achieve its own peak I/O performance, without considering the interference on other concurrent applications. That will often result on contention in the access to the shared infrastructure, which decreases the global I/O performance, makes the applications execution times longer, and therefore wastes expensive computing resources. While contention is an important issue to global I/O performance, the individual performance obtained by applications depend strongly on their access pattern and on the optimization techniques they use. An example is the use of collective operations from the MPI-IO library, which improve performance in many cases, but decrease performance in others. Moreover, its success requires the correct tuning of parameters such as the buffer size and the number of aggregators. Despite some heuristics implemented into the library, most of the responsibility of using collective operations or not (and of choosing the parameters) still belongs to the developers and users. That is a problem because parallel I/O performance depends on a large number of variables and explaining it is not usually a trivial task. We argue it is not reasonable to ask from users and developers to be proficient on parallel I/O in order to achieve high performance. One of the main arguments for this is the fact I/O access patterns known to have poor performance - such as generating small sparse requests - are still frequently observed in large production machines. Hence the main objective of this project is to provide a data management layer that works in the context of the whole machine. A middleware - the data manager - will be responsible for all data accesses from the applications, and will work towards two goals: - improving global metrics of performance by avoiding contention; - improving individual application performance by adapting the used libraries, parameters, and optimization techniques. To achieve that, the data manager requires information about applications, which are not typically available in the stateless HPC I/O stack.
Data: CORDIS, © European Union
Project objective
This project is concerned with the I/O challenges that arise from the convergence between high performance computing (HPC) and big data, two very different paradigms. This convergence is an important topic for the scientific community today, and extreme-scale machines are expected to observe a heterogeneous workload composed of traditional scientific applications and data analytics tasks. The goal of this action is to provide data management for extreme-scale computing environments for the convergence scenario, to benefit both types of workload. The methodology will be an experimental one, and the instrument will be the development of an I/O middleware, the data manager. The data manager will combine storage capacity available in the supercomputer, including NVRAM devices, transparently. Its activities will be optimized by minimizing data movement and applying coordination to avoid performance interference due to concurrency. The most important characteristic of this project among the state-of-the-art is the intelligence to learn and predict applications needs, so storage capacity and data can be available at a close location before the user needs them.The action will benefit from the researcher's experience on parallel I/O for HPC, allied to the host laboratory expertise in in-situ processing, big data, and machine learning. Through this two-year fellowship, the researcher will have the opportunity to expand her knowledge while conducting highly innovative research, what will improve her perspectives for future employment.
Original text from CORDIS.
Participants
- INSTITUT NATIONAL DE RECHERCHE EN INFORMATIQUE ET AUTOMATIQUE · Le Chesnay CedexCoordinatorFrance
Links
Data: CORDIS, © European Union
