FP7Реинтеграция2009–2013

PLURELEARN · Plural Reinforcement Learning

7РП — „Хора“ (Действия „Мария Кюри“)

Период
2009-11-01 → 2013-10-31
Финансиране от ЕС
100 000 €
Участници
1
Схема
MC-IRG

Линиите свързват координатора с партньорите.

Накратко на български

Алгоритмите за обучение на машини се развиват, за да комбинират инструкции от експерти с опит чрез проба и грешка. Това помага за по-доброто управление на сложни системи, при които има голяма неопределеност и много променливи данни.

Този кратък обзор е генериран от изкуствен интелект

Кратко обяснение, генерирано от езиков модел по текста на CORDIS. Оригиналът е по-долу.

Резултати накратко

Plural Reinforcement Learning

The objective of the proposed research is to establish a new paradigm for learning in large-scale complex dynamic systems under uncertainty. Our goal is to develop algorithms, theory, and applications that use plurality of learning approaches and models in a synergetic way. In order to do that we defined the following specific objectives: 1. Develop a learning approach that combines learning from a teacher and learning by trial and error. 2. Devise a structure discovery methodology for reasoning about uncertainty in high dimensional Markov processes. 3. Come up with approaches for algorithm selection and mini-strategies. We have made good progress on all three specific objectives which we detail below. All our research was within the Markov decision processes (MDP) formulation focusing on the reinforcement learning paradigm. Regarding objective 1: we showed in a couple of papers how to use a tutor or expert advice in RL algorithms. Specifically, we considered problems where linear constraints can be added to the learning algorithm representing additional knowledge. We also argued that this knowledge may have to be relaxed, and outlined ways to do it. Overall, we developed new algorithms for the problem of learning from a plurality of sources and showed that it works in medium scale applications. Regarding objective 2: we showed that the problem of structure discovery is much more complex that thought. In fact, we showed that classical approaches for determining which of several models fits the data best when the data has dependencies (as in MDPs) are bound to fail. We then moved on to developing new approaches that are consistent. Overall, we developed the theoretical and applied aspects of model selection and structure discovery and showed it is much harder to detect dynamic structure than expected. As a remedy to model uncertainty we proposed to use robustness to uncertainty and developed two approaches for mitigating risk. The first is based on policy gradients and is geared towards problem where a simulator is available. The second is based on a robust optimization approach where the focus is on couple uncertainties between states. Both approaches tackle different aspects of the optimization problem and facilitate reasoning about uncertainty. Regarding objective 3: We considered various approaches to select mini-strategies (also known as “options” or macro-actions). These are strategies the constitute fragments of the overall policy and whose combination may lead to improved performance. In one line of work, we showed how to modify the option as it runs and generate new and improved options. We can view this process as a “model iteration” since the option model is continuously modified leading. In another line of work we considered the problem of option generation. We showed that it is possible to use “randomly generated” options to expedite both planning and learning. The options model improves with time and (random) options are selected and de-selected continuously. We showed how this new model leads to improved performance in both theory and practice. To summarize, the funded research resulted in a new framework for planning and learning in data-driven stochastic environments. This approach allows combing plurality of information sources (data, simulator, expert and a tutor) for learning and planning faster and more accurately. The research opens up opportunities for large-scale optimization of dynamic systems and may have a significant impact on the scale of problems that we can solve.

Текст от CORDIS, на английски · Данни: CORDIS, © Европейски съюз

Цел на проекта

We propose a new paradigm for learning in complex high-dimensions dynamic environments. Our goal is to develop algorithms, theory, and applications that use plurality of learning approaches and models in a synergetic way. Our paradigm considers the task of learning a control policy by combining trial and error in the style of reinforcement learning with learning from a competent teacher whose interaction with the environment can be observed. Instead of using the teacher for imitation, our paradigm is focused on learning good representations of the world-model. We consider four specific issues in the new paradigm: (i) The usage of iteration and reiteration between learning from a teacher and reinforcement learning. (ii) Learning representation and structure from the teacher. (iii) Optimizing policies based on learned representations and reasoning about model uncertainty. (iv) Learning sub-strategies from a teacher and when and how to use them. We will develop algorithms and theory pertaining to the new paradigm and will apply it in two challenging domains: a fighter jet simulator and a network operating center simulator.

Оригинален текст от CORDIS (на английски).

Участници

  • TECHNION - ISRAEL INSTITUTE OF TECHNOLOGY · HaifaКоординаторИзраел

Връзки

Данни: CORDIS, © Европейски съюз