H2020Individual fellowship2018–2020

LHCBIGDATA · Exploiting big data and machine learning techniques for LHC experiments

Horizon 2020 — Marie Skłodowska-Curie Actions

Duration
2018-07-02 → 2020-07-01
EU contribution
€180,277
Participants
1
Scheme
MSCA-IF-EF-ST

Lines connect the coordinator with its partners.

Results in brief

Exploiting big data and machine learning techniques for LHC experiments

The aim of the LHCBIGDATA project is to provide the Large Hadron Collider (LHC) community with the necessary tools to deploy Machine Learning (ML) solutions. The tools under development are experiment-independent to promote the exchange of common solutions among the various LHC communities. The benefits of such an approach are being applied to a real world use case, the optimization of the computing operations for the CMS experiment. To set the scale, a typical LHC experiment manages a computing infrastructure of more than 100000 cores spread over more than one hundred computing centres around the world. The two biggest experiments (ATLAS and CMS) have collected and produced around 1 EB of data since LHC start. Operating such an infrastructure still has a very large human cost (of the order of 50-100 FTEs per year per experiment). The following objectives have been identified for this project: • Development of a scalable ML framework and integration of several architectures (CPU, GPUs, FPGAs); • Promotion of ML techniques within the LHC community and beyond, by creating links with the local scientific community, and within international collaborations in the field of High Energy Physics (HEP); • Application to the use case of the optimisation of CMS computing operations.

Data: CORDIS, © European Union

Project objective

Large international scientific collaborations will face in the near future unprecedented computing and data challenges. The analysis of multi-PetaByte datasets at CMS, ATLAS, LHCb and Alice, the four experiments at the Large Hadron Collider (LHC), requires a global federated infrastructure of distributed computing resources. The HL-LHC, the High Luminosity upgrade of the LHC, is expected to deliver 100 times more data than the LHC, with corresponding increase of event sizes, volumes and complexity. Modern techniques for big data analytics and machine learning (ML) are needed to cope with such unprecedented data stream. Critical areas that will strongly benefit from ML are data analysis, detector operation including calibration and monitoring, and computing operations. Aim of this project is to provide the LHC community with the necessary tools to deploy ML solutions through the use of open cloud technologies such as the INDIGO-DataCloud services. Heterogeneous technologies (systems based on multi-cores, GPUs, ...) and opportunistic resources will be integrated. The developed tools will be experiment-independent to promote the exchange of common solutions among the various LHC experiments. The benefits of such an approach will be demonstrated in a real world use case, the optimization of the computing operations for the CMS experiment. In addition, once available, the tools to deploy ML as a service can be easily transferred to other scientific domains that have the need to treat large data streams.

Original text from CORDIS.

Participants

  • ISTITUTO NAZIONALE DI FISICA NUCLEARE · FrascatiCoordinatorItaly

Links

Data: CORDIS, © European Union