Fetching the paper…
Reading the bibliography…
For over a decade, model-based reinforcement learning has been seen as a way to leverage control-based domain knowledge to improve the sample-efficiency of reinforcement learning agents.
Integrated architectures for learning, planning, and reacting based on approximating dynamic programming
Richard S Sutton · 1990
Earlier work this paper cites.
Python reference manual
Guido Van Rossum and Fred L Drake Jr · 1995
Earlier work this paper cites.
A guide to NumPy , volume 1
Travis E Oliphant · 2006
Earlier work this paper cites.
Matplotlib: A 2d graphics environment
John D Hunter · 2007
Earlier work this paper cites.
Python for scientific computing
Travis E Oliphant · 2007
Earlier work this paper cites.
Fitted q-iteration in continuous action-space mdps
András Antos, Csaba Szepesvári, and Rémi Munos · 2008
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
Brian D Ziebart, Andrew L Maas, J Andrew Bagnell, and Anind K Dey · 2008
Earlier work this paper cites.
Double q-learning
Hado V Hasselt · 2010
Earlier work this paper cites.
Algorithms for reinforcement learning
Csaba Szepesvári · 2010
Earlier work this paper cites.
Modeling purposeful adaptive behavior with the principle of maximum causal entropy
Brian D Ziebart · 2010
Earlier work this paper cites.
Pilco: A model-based and data-efficient approach to policy search
Marc Deisenroth and Carl E Rasmussen · 2011
Earlier work this paper cites.
The numpy array: a structure for efficient numerical computation
Stefan Van Der Walt, S Chris Colbert, and Gael Varoquaux · 2011
Earlier work this paper cites.
Python for data analysis: Data wrangling with Pandas, NumPy, and IPython
Wes McKinney · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Earlier work this paper cites.
Reinforcement learning in robotics: A survey
Jens Kober, J Andrew Bagnell, and Jan Peters · 2013
Earlier work this paper cites.
Guided policy search
Sergey Levine and Vladlen Koltun · 2013
Earlier work this paper cites.
Learning phrase representations using rnn encoder-decoder for statistical machine translation
Kyunghyun Cho, Bart Van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio · 2014
Earlier work this paper cites.
SciPy: Open source scientific tools for Python
Eric Jones, Travis Oliphant, and Pearu Peterson · 2014
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
Martin L Puterman · 2014
Earlier work this paper cites.
Bias in natural actor-critic algorithms
Philip Thomas · 2014
Earlier work this paper cites.
Taming the noise in reinforcement learning via soft updates
Roy Fox, Ari Pakman, and Naftali Tishby · 2015
Earlier work this paper cites.
Learning continuous control policies by stochastic value gradients
Nicolas Heess, Gregory Wayne, David Silver, Timothy Lillicrap, Tom Erez, and Yuval Tassa · 2015
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Earlier work this paper cites.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Earlier work this paper cites.
Continuous deep q-learning with model-based acceleration
Shixiang Gu, Timothy Lillicrap, Ilya Sutskever, and Sergey Levine · 2016
Earlier work this paper cites.
Jupyter notebooks-a publishing format for reproducible computational workflows
Thomas Kluyver, Benjamin Ragan-Kelley, Fernando Pérez, Brian E Granger, Matthias Bussonnier, Jonathan Frederic, Kyle Kelley, Jessica B Hamrick, Jason Grout, Sylvain Corlay, et al · 2016
Earlier work this paper cites.
Reproducibility of benchmarked deep reinforcement learning tasks for continuous control
Riashat Islam, Peter Henderson, Maziar Gomrokchi, and Doina Precup · 2017
Earlier work this paper cites.
Optimal and autonomous control using reinforcement learning: A survey
Bahare Kiumarsi, Kyriakos G Vamvoudakis, Hamidreza Modares, and Frank L Lewis · 2017
Cited alongside, same era.
Path integral networks: End-to-end differentiable optimal control
Masashi Okada, Luca Rigazio, and Takenobu Aoshima · 2017
Cited alongside, same era.
Survey of model-based reinforcement learning: Applications on robotics
Athanasios S Polydoros and Lazaros Nalpantidis · 2017
Cited alongside, same era.
Maximum a posteriori policy optimisation
Abbas Abdolmaleki, Jost Tobias Springenberg, Yuval Tassa, Remi Munos, Nicolas Heess, and Martin Riedmiller · 2018
Cited alongside, same era.
Differentiable mpc for end-to-end planning and control
Brandon Amos, Ivan Jimenez, Jacob Sacks, Byron Boots, and J Zico Kolter · 2018
Cited alongside, same era.
Dream to control: Learning behaviors by latent imagination
Danijar Hafner, Timothy Lillicrap, Jimmy Ba, and Mohammad Norouzi · 2019
Later among the works it cites.
When to trust your model: Model-based policy optimization
Michael Janner, Justin Fu, Marvin Zhang, and Sergey Levine · 2019
Later among the works it cites.
Deep learning for video game playing
Niels Justesen, Philip Bontrager, Julian Togelius, and Sebastian Risi · 2019
Later among the works it cites.
Model-based reinforcement learning for atari
Lukasz Kaiser, Mohammad Babaeizadeh, Piotr Milos, Blazej Osinski, Roy H Campbell, Konrad Czechowski, Dumitru Erhan, Chelsea Finn, Piotr Kozakowski, Sergey Levine, et al · 2019
Later among the works it cites.
Stabilizing off-policy q-learning via bootstrapping error reduction
Aviral Kumar, Justin Fu, Matthew Soh, George Tucker, and Sergey Levine · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sample-efficient reinforcement learning with stochastic ensemble value expansion
Jacob Buckman, Danijar Hafner, George Tucker, Eugene Brevdo, and Honglak Lee · 2018
Cited alongside, same era.
A lyapunov-based approach to safe reinforcement learning
Yinlam Chow, Ofir Nachum, Edgar Duenez-Guzman, and Mohammad Ghavamzadeh · 2018
Cited alongside, same era.
Deep reinforcement learning in a handful of trials using probabilistic dynamics models
Kurtland Chua, Roberto Calandra, Rowan McAllister, and Sergey Levine · 2018
Cited alongside, same era.
Safe exploration in continuous action spaces
Gal Dalal, Krishnamurthy Dvijotham, Matej Vecerik, Todd Hester, Cosmin Paduraru, and Yuval Tassa · 2018
Cited alongside, same era.
Model-based value expansion for efficient model-free reinforcement learning
V Feinberg, A Wan, I Stoica, MI Jordan, JE Gonzalez, and S Levine · 2018
Cited alongside, same era.
Reinforcement learning in financial markets-a survey
Thomas G Fischer · 2018
Cited alongside, same era.
Addressing function approximation error in actor-critic methods
Scott Fujimoto, Herke van Hoof, and David Meger · 2018
Cited alongside, same era.
Later among the works it cites.
Stochastic latent actor-critic: Deep reinforcement learning with a latent variable model
A. X. Lee, A. Nagabandi, P. Abbeel, and S. Levine · 2019
Later among the works it cites.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al · 2019
Later among the works it cites.
Self-supervised exploration via disagreement
Deepak Pathak, Dhiraj Gandhi, and Abhinav Gupta · 2019
Later among the works it cites.
When to use parametric models in reinforcement learning?
Hado P van Hasselt, Matteo Hessel, and John Aslanides · 2019
Later among the works it cites.
Grandmaster level in starcraft ii using multi-agent reinforcement learning
Oriol Vinyals, Igor Babuschkin, Wojciech M Czarnecki, Michaël Mathieu, Andrew Dudzik, Junyoung Chung, David H Choi, Richard Powell, Timo Ewalds, Petko Georgiev, et al · 2019
Later among the works it cites.
Benchmarking model-based reinforcement learning
T. Wang, X. Bao, I. Clavera, J. Hoang, Y. Wen, E. Langlois, S. Zhang, G. Zhang, P. Abbeel, and J. Ba · 2019
Later among the works it cites.
Exploring model-based planning with policy networks
Tingwu Wang and Jimmy Ba · 2019
Later among the works it cites.
Behavior regularized offline reinforcement learning
Yifan Wu, George Tucker, and Ofir Nachum · 2019
Later among the works it cites.
Hydra - a framework for elegantly configuring complex applications
Omry Yadan · 2019
Later among the works it cites.
Improving sample efficiency in model-free reinforcement learning from images
Denis Yarats, Amy Zhang, Ilya Kostrikov, Brandon Amos, Joelle Pineau, and Rob Fergus · 2019
Later among the works it cites.
Model-augmented actor-critic: Backpropagating through paths
Ignasi Clavera, Violet Fu, and Pieter Abbeel · 2020
Closest in time.
Infinite-horizon differentiable model predictive control
Sebastian East, Marco Gallieri, Jonathan Masci, Jan Koutník, and Mark Cannon · 2020
Closest in time.
Implementation matters in deep policy gradients: A case study on ppo and trpo
Logan Engstrom, Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras, Firdaus Janoos, Larry Rudolph, and Aleksander Madry · 2020
Closest in time.
On the role of planning in model-based deep reinforcement learning
Jessica B Hamrick, Abram L Friesen, Feryal Behbahani, Arthur Guez, Fabio Viola, Sims Witherspoon, Thomas Anthony, Lars Buesing, Petar Veličković, and Théophane Weber · 2020
Closest in time.
Sunrise: A simple unified framework for ensemble learning in deep reinforcement learning
Kimin Lee, Michael Laskin, Aravind Srinivas, and Pieter Abbeel · 2020
Closest in time.
Iterative amortized policy optimization
Joseph Marino, Alexandre Piché, Alessandro Davide Ialongo, and Yisong Yue · 2020
Closest in time.
Mastering atari, go, chess and shogi by planning with a learned model
Julian Schrittwieser, Ioannis Antonoglou, Thomas Hubert, Karen Simonyan, Laurent Sifre, Simon Schmitt, Arthur Guez, Edward Lockhart, Demis Hassabis, Thore Graepel, et al · 2020
Closest in time.
Planning to explore via self-supervised world models
Ramanan Sekar, Oleh Rybkin, Kostas Daniilidis, Pieter Abbeel, Danijar Hafner, and Deepak Pathak · 2020
Closest in time.
Dynamics-aware unsupervised skill discovery
Archit Sharma, Shane Gu, Sergey Levine, Vikash Kumar, and Karol Hausman · 2020
Closest in time.
D2rl: Deep dense architectures in reinforcement learning
Samarth Sinha, Homanga Bharadhwaj, Aravind Srinivas, and Animesh Garg · 2020
Closest in time.
Local search for policy iteration in continuous control
Jost Tobias Springenberg, Nicolas Heess, Daniel Mankowitz, Josh Merel, Arunkumar Byravan, Abbas Abdolmaleki, Jackie Kay, Jonas Degrave, Julian Schrittwieser, Yuval Tassa, et al · 2020
Closest in time.
Soft actor-critic (sac) implementation in pytorch
Denis Yarats and Ilya Kostrikov · 2020
Closest in time.