2018

Model-Based Value Estimation for Efficient Model-Free Reinforcement Learning

Feinberg, Vladimir, Wan, Alvin, Stoica, Ion et al.

Understand

Recent model-free reinforcement learning algorithms have proposed incorporating learned dynamics models as a source of additional data with the intention of reducing sample complexity.

  • Such methods hold the promise of incorporating imagined data coupled with a notion of model uncertainty to accelerate the learning of continuous control tasks.
  • Unfortunately, they rely on heuristics that limit usage of the dynamics model.
  • We present model-based value expansion, which controls for uncertainty in the model by only allowing imagination to fixed depth.

Built on

  • Integrated architectures for learning, planning, and reacting based on approximating dynamic programming

    Sutton, Richard S · 1990

    Earlier work this paper cites.

  • Incremental multi-step Q-learning

    Peng, Jing and Williams, Ronald J · 1994

    Earlier work this paper cites.

  • Off-policy actor-critic

    Original

    Degris, Thomas, White, Martha, and Sutton, Richard S · 2012

    Earlier work this paper cites.

  • Mujoco: A physics engine for model-based control

    Todorov, Emanuel, Erez, Tom, and Tassa, Yuval · 2012

    Earlier work this paper cites.

  • Deterministic policy gradient algorithms

    Silver, David, Lever, Guy, Heess, Nicolas, Degris, Thomas, Wierstra, Daan, and Riedmiller, Martin · 2014

    Earlier work this paper cites.

  • Learning continuous control policies by stochastic value gradients

    Heess, Nicolas, Wayne, Gregory, Silver, David, Lillicrap, Tim, Erez, Tom, and Tassa, Yuval · 2015

    Earlier work this paper cites.

Similar

  • OpenAI gym

    Original

    Brockman, Greg, Cheung, Vicki, Pettersson, Ludwig, Schneider, Jonas, Schulman, John, Tang, Jie, and Zaremba, Wojciech · 2016

    Cited alongside, same era.

  • Continuous deep q-learning with model-based acceleration

    Gu, Shixiang, Lillicrap, Timothy, Sutskever, Ilya, and Levine, Sergey · 2016

    Cited alongside, same era.

  • Continuous control with deep reinforcement learning

    Lillicrap, Timothy P, Hunt, Jonathan J, Pritzel, Alexander, Heess, Nicolas, Erez, Tom, Tassa, Yuval, Silver, David, and Wierstra, Daan · 2016

    Cited alongside, same era.

  • On the sample complexity of the linear quadratic regulator

    Original

    Dean, Sarah, Mania, Horia, Matni, Nikolai, Recht, Benjamin, and Tu, Stephen · 2017

    Cited alongside, same era.

  • Uncertainty-driven imagination for continuous deep reinforcement learning

    Kalweit, Gabriel and Boedecker, Joschka · 2017

    Cited alongside, same era.

Then

  • Bridging the gap between value and policy based reinforcement learning

    Nachum, Ofir, Norouzi, Mohammad, Xu, Kelvin, and Schuurmans, Dale · 2017

    Later among the works it cites.

  • Value prediction network

    Oh, Junhyuk, Singh, Satinder, and Lee, Honglak · 2017

    Later among the works it cites.

  • Parameter space noise for exploration

    Plappert, Matthias, Houthooft, Rein, Dhariwal, Prafulla, Sidor, Szymon, Chen, Richard Y, Chen, Xi, Asfour, Tamim, Abbeel, Pieter, and Andrychowicz, Marcin · 2017

    Later among the works it cites.

  • Imagination-augmented agents for deep reinforcement learning

    Racanière, Sébastien, Weber, Théophane, Reichert, David, Buesing, Lars, Guez, Arthur, Rezende, Danilo Jimenez, Badia, Adrià Puigdomènech, Vinyals, Oriol, Heess, Nicolas, Li, Yujia, et al · 2017

    Later among the works it cites.

  • Model-ensemble trust-region policy optimization

    Kurutach, Thanard, Clavera, Ignasi, Duan, Yan, Tamar, Aviv, and Abbeel, Pieter · 2018

    Closest in time.

Beyond the bibliography

alphaXiv searches the wider corpus for related work and actual follow-ups.

Open on alphaXiv

alphaXiv is searching for related work…