H2020Индивидуална стипендия2021–2023

Stein-ML · Stein’s method and functional inequalities in machine learning

„Хоризонт 2020“ — Действия „Мария Склодовска-Кюри“

Период
2021-09-01 → 2023-11-30
Финансиране от ЕС
186 451 €
Участници
2
Схема
MSCA-IF

Линиите свързват координатора с партньорите.

Накратко на български

Математическите методи за оценка на грешките при приблизителните изчисления в машинното обучение се изследват чрез примери от медицината и астрофизиката. Това помага за по-добра проверка на надеждността и точността на алгоритмите, които анализират сложни данни.

Този кратък обзор е генериран от изкуствен интелект

Кратко обяснение, генерирано от езиков модел по текста на CORDIS. Оригиналът е по-долу.

Резултати накратко

Stein’s method and functional inequalities in machine learning

Mathematical statistics is a ubiquitous tool in modern data analysis. However, in typical applications, theoretically supported statistical frameworks cannot be used directly - for instance because they lack closed-form solutions or are too computationally expensive. Because of this, approximate inference techniques have been developed in the recent years as a way to speed up the learning process of algorithms based on statistics. However, the quality of the associated approximations is still not well-understood. Researchers and practitioners therefore look for tools to measure both the efficiency of the associated inference and sampling methods and the size of the error they generate when using approximations. Indeed, such tools are still not widely available in the literature for many of the commonly used models, therefore leaving researchers and practitioners unable to assess whether their methods are simultaneously fast and robust enough for their purposes. At the same time, due to the broad availability of probabilistic programming languages, approximate inference has been applied to a wide variety of problems, related to medical imaging, detection of gravitational waves or, recently, modelling of infectious diseases. Underestimation of uncertainty or inaccuracy of point estimates in those applications could undermine the reliability of the associated research or the quality of measures introduced as a result of it. The aim of this research project was to advance the development of quality measures for approximations in machine learning and statistics, using the rich theoretical machinery of mathematical analysis and probability. The project has been concluded with five papers, each targeting a specific approximation. Specifically, together with collaborators, we have achieved the following goals. We have constructed fully computable quality guarantees for the Laplace approximation of a Bayesian posterior, with respect to a variety of useful divergences. We have also constructed a new functional-data goodness-of-fit test for Gaussian Process targets and measures absolutely continuous with respect to Gaussians. Moreover, we have provided a novel targeted accuracy diagnostic for distributional approximations. Furthermore, in another paper, we have proved a functional version of the celebrated de Jong Theorem describing the asymptotic behaviour of U-statistics. Finally, we have proved new Berry-Esseen bounds for vector-valued statistics of binomial processes, with respect to the convex distance.

Текст от CORDIS, на английски · Данни: CORDIS, © Европейски съюз

Цел на проекта

The project aims to develop quality measures for approximations in machine learning and statistics, using tools of probability and functional analysis, such as Stein's method and functional inequalities. Approximate inference techniques have been used in the recent years as a way to speed up the learning process, which is particularly important in the era of big data. It is, however, necessary for researchers to be able to measure the error of the associated approximations. Indeed, wrong variance or mean estimates in applications related, for instance, to modelling infectious diseases, may have highly negative outcomes. In this project, I will concentrate on three specific aspects of this problem. I will firstly propose tools for measuring the quality of posterior approximations in Gaussian Process inference. In order to do this, I will use the theory of Stein discrepancies which has already been successfully applied, in the context of Bayesian inference, to finite-dimensional distributions. I will combine it with the recent developments in probability theory related to Stein's method for infinite-dimensional measures. Secondly, I will construct a tool for a simultaneous study of the rate of convergence and the output quality of MCMC schemes based on discretising diffusion processes. Both those objects may be analysed using the infinitesimal generator of the underlying diffusion. Indeed, for the former we may apply the associated log-Sobolev or Poincare inequalities and, for the latter, utilise the associated Stein operator. The resulting tool will help users choose (or construct) an algorithm which is simultaneously fast and robust. Thirdly, I will construct a Gaussian-Process goodness-of-fit test, allowing users to test whether the given data come from a marginal of a particular GP. In order to do this, I will use infinite-dimensional Stein’s method together with techniques used recently to construct kernel goodness-of-fit tests based on Stein discrepancies.

Оригинален текст от CORDIS (на английски).

Участници

Връзки

Данни: CORDIS, © Европейски съюз