PMOHR · Probabilistic modelling of electronic health records
„Хоризонт 2020“ — Действия „Мария Склодовска-Кюри“
- Период
- 2016-10-01 → 2019-09-30
- Финансиране от ЕС
- 269 858 €
- Участници
- 2
- Схема
- MSCA-IF-GF
Линиите свързват координатора с партньорите.
Накратко на български
Вероятностното машинно обучение се използва за създаване на разбираеми модели от електронни здравни досиета. Това помага за откриване на скрити закономерности в медицинските данни и създаване на системи за клинична подкрепа.
Кратко обяснение, генерирано от езиков модел по текста на CORDIS. Оригиналът е по-долу.
Резултати накратко
Probabilistic modelling of electronic health records
Machine learning is a field of computer science with the aim of discovering statistical patterns in large datasets and making predictions using these patterns. While there is a growing interest on methods that can provide accurate predictions on unseen data, comparatively less effort is invested on designing interpretable machine learning models. Interpretable models have the benefit to provide human-understandable patterns that can be interpreted by domain experts. This enables knowledge discovery in the scientific disciplines, and it also allows reasoning about causality and making counterfactual predictions. In probabilistic machine learning, we first encode our assumptions about the data structure in the form of a model that has latent variables, which represent the hidden patterns. We then learn the latent variables using an inference algorithm. The results of the inference procedure can be used to make predictions and explore the data collection. Crucially, we can impose an interpretable structure in the model design phase. Finally, in probabilistic machine learning we can also apply model testing approaches to verify the properties of the posited model and whether if fails capture certain properties of the data. PMOHR is an interdisciplinary project focused on the design of interpretable models through probabilistic machine learning, with the ultimate goal of modelling Electronic Health Records (EHRs). Applying machine learning tools to EHR data can help not only design clinical support systems, but it may also lead to uncover unknown patterns from the data and even form causal theories. However, medical data, and EHRs in particular, present several challenges that prevent us from applying probabilistic modelling tools, because the datasets are large and heterogeneous. The objectives of PMOHR are to develop both probabilistic models and inference algorithms that are suitable for modelling EHR data. This new set of tools can then be applied to make predictions and to analyse medical datasets. In particular, PMOHR uses both publicly available EHR data and EHR data from the New York Presbyterian Hospital.
Текст от CORDIS, на английски · Данни: CORDIS, © Европейски съюз
Цел на проекта
The growing worldwide adoption of Electronic Health Records (EHR) enables new research opportunities to analyse massive amounts of medical information, motivated by the promise of improving health systems while providing significant budget savings. Biomedical research increasingly uses machine learning methods as a data-driven approach to learn complex comorbidity patterns of diseases, study drug interactions, and form predictions. The analysis of EHRs may not only lead to knowledge discovery, but it also facilitates personalised medical treatment and early diagnosis of the diseases through the design of clinical support systems.However, current approaches for the analysis of EHRs are still in their early stages. The two main technical challenges that need to be addressed are integration of heterogeneous data and scalability to massive datasets. Most of the existing methods are tailored to homogeneous data and, therefore, to a single source of information, and hence they cannot handle EHR datasets. Scalability also represents a difficulty for most of the current machine learning techniques, which are limited to the analysis to moderate-sized datasets.In this project, we will develop novel tools for the analysis of heterogeneous EHR data. Our approach will be based on probabilistic modelling techniques, since they are an effective approach for understanding real-world data in many areas of science. We will make use of Bayesian nonparametric modelling techniques, coupled with stochastic variational inference to allow for scalable inference. Probabilistic models, including BNPs, are amenable to both descriptive and predictive analysis at the same time. We will collaborate with the Department of Biomedical Informatics, who will provide their knowledge about the problem, allowing for good model formulations and results analysis.
Оригинален текст от CORDIS (на английски).
Участници
- THE CHANCELLOR MASTERS AND SCHOLARS OF THE UNIVERSITY OF CAMBRIDGE · CAMBRIDGEКоординаторОбединеното кралство
- TRUSTEES OF COLUMBIA UNIVERSITY IN THE CITY OF NEW YORK · New YorkСъединени щати
Връзки
Данни: CORDIS, © Европейски съюз
