Understand
Recent model-free reinforcement learning algorithms have proposed incorporating learned dynamics models as a source of additional data with the intention of reducing sample complexity.
- Such methods hold the promise of incorporating imagined data coupled with a notion of model uncertainty to accelerate the learning of continuous control tasks.
- Unfortunately, they rely on heuristics that limit usage of the dynamics model.
- We present model-based value expansion, which controls for uncertainty in the model by only allowing imagination to fixed depth.
Built on
Integrated architectures for learning, planning, and reacting based on approximating dynamic programming
Sutton, Richard S · 1990
Earlier work this paper cites.
Incremental multi-step Q-learning
Peng, Jing and Williams, Ronald J · 1994
Earlier work this paper cites.
Degris, Thomas, White, Martha, and Sutton, Richard S · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Todorov, Emanuel, Erez, Tom, and Tassa, Yuval · 2012
Earlier work this paper cites.
Deterministic policy gradient algorithms
Silver, David, Lever, Guy, Heess, Nicolas, Degris, Thomas, Wierstra, Daan, and Riedmiller, Martin · 2014
Earlier work this paper cites.
Learning continuous control policies by stochastic value gradients
Heess, Nicolas, Wayne, Gregory, Silver, David, Lillicrap, Tim, Erez, Tom, and Tassa, Yuval · 2015
Earlier work this paper cites.
Similar
Brockman, Greg, Cheung, Vicki, Pettersson, Ludwig, Schneider, Jonas, Schulman, John, Tang, Jie, and Zaremba, Wojciech · 2016
Cited alongside, same era.
Continuous deep q-learning with model-based acceleration
Gu, Shixiang, Lillicrap, Timothy, Sutskever, Ilya, and Levine, Sergey · 2016
Cited alongside, same era.
Continuous control with deep reinforcement learning
Lillicrap, Timothy P, Hunt, Jonathan J, Pritzel, Alexander, Heess, Nicolas, Erez, Tom, Tassa, Yuval, Silver, David, and Wierstra, Daan · 2016
Cited alongside, same era.
On the sample complexity of the linear quadratic regulator
Dean, Sarah, Mania, Horia, Matni, Nikolai, Recht, Benjamin, and Tu, Stephen · 2017
Cited alongside, same era.
Uncertainty-driven imagination for continuous deep reinforcement learning
Kalweit, Gabriel and Boedecker, Joschka · 2017
Cited alongside, same era.
Then
Bridging the gap between value and policy based reinforcement learning
Nachum, Ofir, Norouzi, Mohammad, Xu, Kelvin, and Schuurmans, Dale · 2017
Later among the works it cites.
Value prediction network
Oh, Junhyuk, Singh, Satinder, and Lee, Honglak · 2017
Later among the works it cites.
Parameter space noise for exploration
Plappert, Matthias, Houthooft, Rein, Dhariwal, Prafulla, Sidor, Szymon, Chen, Richard Y, Chen, Xi, Asfour, Tamim, Abbeel, Pieter, and Andrychowicz, Marcin · 2017
Later among the works it cites.
Imagination-augmented agents for deep reinforcement learning
Racanière, Sébastien, Weber, Théophane, Reichert, David, Buesing, Lars, Guez, Arthur, Rezende, Danilo Jimenez, Badia, Adrià Puigdomènech, Vinyals, Oriol, Heess, Nicolas, Li, Yujia, et al · 2017
Later among the works it cites.
Model-ensemble trust-region policy optimization
Kurutach, Thanard, Clavera, Ignasi, Duan, Yan, Tamar, Aviv, and Abbeel, Pieter · 2018
Closest in time.
Beyond the bibliography
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…