DESlRE · Data-Efficient Scalable Reinforcement Learning for Practical Robotic Environments
„Хоризонт 2020“ — Действия „Мария Склодовска-Кюри“
- Период
- 2018-04-01 → 2020-03-31
- Финансиране от ЕС
- 159 461 €
- Участници
- 1
- Схема
- MSCA-IF-EF-ST
Линиите свързват координатора с партньорите.
Накратко на български
Алгоритмите за обучение с подкрепление се изследват, за да работят реалните роботи също добре, както в компютърните симулации. Това помага за по-доброто разбиране на управлението в непредвидима среда и вземането на социални решения.
Кратко обяснение, генерирано от езиков модел по текста на CORDIS. Оригиналът е по-долу.
Резултати накратко
Data-Efficient Scalable Reinforcement Learning for Practical Robotic Environments
The initial aim of the project was to develop algorithms suitable for challenging control tasks. Current algorithms that perform well in simulation typically transfers suboptimally to real test or robots. For example, one of the hot topics in research is termed- sim-to-real transfer, which aims to transfer amazing feats that algorithms can achieve in computer simulations to test time performance. If we can understand how to perform control in a highly stochastic environment, many problems in social decision-making can be solved. The overall objectives are to advance our understanding of the difficulty of such an application of algorithms and developing new ones.
Текст от CORDIS, на английски · Данни: CORDIS, © Европейски съюз
Цел на проекта
The robotics industry is in the process of greater adoption of machine learning. Recent reinforcement learning (RL) and AI breakthroughs, such as AlphaGo, rely on collecting large amounts of data. Such methods are unsuitable for real robots which often can only afford a few trials. Moreover, some states are unsafe to explore, e.g. running over a cliff. Conversely, works such as PILCO combine Bayesian models with model-based RL to improve data efficiency. Those frameworks typically thrive in small data regimes. The goal of this project is to develop RL algorithms that scale to high dimensions while learning with less data. The main pillars of our methodology are RL, recurrent networks, Bayesian methods, embodied exploration, and optimization. To tackle the data efficiency, we adopt model-based RL approaches. We plan to combine representation learning and dynamics in a single model, leading to high predictive power and low-dimensional internal state spaces. Notably, we use methods that can learn disentangled representations, e.g. infoGAN. In practical robots, effective exploration is a real problem in current approaches. We want to leverage recent works in embodied exploration by the host group which allows various real-world robots to explore their capabilities in minutes of interaction. I received my Ph.D. for work in optimization with Dr. William Hager. I also conducted postdoctoral research in machine learning. The Autonomous Learning group is led by Dr. Georg Martius, who has previously studied artificial intrinsic motivation, the self-organized exploration of sensorimotor coordination via information theory, and internal model learning. He also developed the robotics environment LPZRobots. I will gain extensive experience in practical robotics, embodied exploration, and information theory through the collaboration and mature as an advanced AI researcher. Both Dr. Martius and I have a track record of publishing code online. We will continue this effort.
Оригинален текст от CORDIS (на английски).
Участници
- MAX-PLANCK-GESELLSCHAFT ZUR FORDERUNG DER WISSENSCHAFTEN EV · MUNCHENКоординаторГермания
Връзки
Данни: CORDIS, © Европейски съюз
