Fetching the paper…
Reading the bibliography…
This paper considers the problem of learning a model in model-based reinforcement learning (MBRL).
Benchmarking model-based reinforcement learning
Tingwu Wang, Xuchan Bao, Ignasi Clavera, Jerrick Hoang, Yeming Wen, Eric Langlois, Shunshi Zhang, Guodong Zhang, Pieter Abbeel, and Jimmy Ba · 1907
Earlier work this paper cites.
Neural policy gradient methods: Global optimality and rates of convergence
Lingxiao Wang, Qi Cai, Zhuoran Yang, and Zhaoran Wang · 1909
Earlier work this paper cites.
Integrated architectures for learning, planning, and reacting based on approximating dynamic programming
Richard S. Sutton · 1990
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J. Williams · 1992
Earlier work this paper cites.
Efficient learning and planning within the Dyna framework
Jing Peng and Ronald J. Williams · 1993
Earlier work this paper cites.
Stable function approximation in dynamic programming
Geoffrey Gordon · 1995
Earlier work this paper cites.
Integral probability metrics and their generating classes of functions
Alfred Müller · 1997
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S. Sutton, David McAllester, Satinder Singh, and Yishay Mansour · 2000
Earlier work this paper cites.
Infinite-horizon policy-gradient estimation
Jonathan Baxter and Peter L. Bartlett · 2001
Earlier work this paper cites.
A natural policy gradient
Sham Kakade · 2001
Earlier work this paper cites.
On actor-critic algorithms
Vijay R. Konda and John N. Tsitsiklis · 2001
Earlier work this paper cites.
Simulation-based optimization of Markov reward processes
Peter Marbach and John N. Tsitsiklis · 2001
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Sham Kakade and John Langford · 2002
Earlier work this paper cites.
Reinforcement learning for humanoid robotics
Jan Peters, Vijayakumar Sethu, and Stefan Schaal · 2003
Earlier work this paper cites.
Policy search by dynamic programming
J. Andrew Bagnell, Sham Kakade, Andrew Y. Ng, and Jeff Schneider · 2004
Earlier work this paper cites.
Interpolation-based Q-learning
Csaba Szepesvári and William D. Smart · 2004
Earlier work this paper cites.
A basic formula for online policy gradient algorithms
Xi-Ren Cao · 2005
Earlier work this paper cites.
Tree-based batch mode reinforcement learning
Damien Ernst, Pierre Geurts, and Louis Wehenkel · 2005
Earlier work this paper cites.
Bayesian policy gradient algorithms
Mohammad Ghavamzadeh and Yaakov Engel · 2007
Earlier work this paper cites.
Learning near-optimal policies with Bellman-residual minimization based fitted policy iteration and a single sample path
András Antos, Csaba Szepesvári, and Rémi Munos · 2008
Earlier work this paper cites.
Finite-time bounds for fitted value iteration
Rémi Munos and Csaba Szepesvári · 2008
Earlier work this paper cites.
Natural actor-critic
Jan Peters and Stefan Schaal · 2008
Earlier work this paper cites.
Dyna-style planning with linear function approximation and prioritized sweeping
Richard S. Sutton, Csaba Szepesvári, Alborz Geramifard, and Michael Bowling · 2008
Cited alongside, same era.
Natural actor–critic algorithms
Shalabh Bhatnagar, Richard S. Sutton, Mohammad Ghavamzadeh, and Mark Lee · 2009
Cited alongside, same era.
Regularized fitted Q-iteration for planning in continuous-space Markovian Decision Problems
Amir-massoud Farahmand, Mohammad Ghavamzadeh, Csaba Szepesvári, and Shie Mannor · 2009
Cited alongside, same era.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio · 2010
Cited alongside, same era.
Analysis of a classification-based policy iteration algorithm
Alessandro Lazaric, Mohammad Ghavamzadeh, and Rémi Munos · 2010
Cited alongside, same era.
Algorithms for Reinforcement Learning
Csaba Szepesvári · 2010
Value-aware loss function for model-based reinforcement learning
Amir-massoud Farahmand, André M.S. Barreto, and Daniel N. Nikovski · 2017
Later among the works it cites.
Value prediction network
Junhyuk Oh, Satinder Singh, and Honglak Lee · 2017
Later among the works it cites.
Asymptotic bias of stochastic gradient search
Vladislav B Tadić, Arnaud Doucet, et al · 2017
Later among the works it cites.
Self-correcting models for model-based reinforcement learning
Erik Talvitie · 2017
Later among the works it cites.
Boosted fitted Q-iteration
Samuele Tosatto, Matteo Pirotta, Carlo D’Eramo, and Marcello Restelli · 2017
Later among the works it cites.
Equivalence between wasserstein and value-aware model-based reinforcement learning
Kavosh Asadi, Evan Cater, Dipendra Misra, and Michael L. Littman · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Approximate policy iteration: A survey and some new methods
Dimitri P. Bertsekas · 2011
Cited alongside, same era.
Value pursuit iteration
Amir-massoud Farahmand and Doina Precup · 2012
Cited alongside, same era.
Finite-sample analysis of least-squares policy iteration
Alessandro Lazaric, Mohammad Ghavamzadeh, and Rémi Munos · 2012
Cited alongside, same era.
Approximate modified policy iteration
Bruno Scherrer, Mohammad Ghavamzadeh, Victor Gabillon, and Matthieu Geist · 2012
Cited alongside, same era.
A survey on policy search for robotics
Marc Peter Deisenroth, Gerhard Neumann, and Jan Peters · 2013
Cited alongside, same era.
Reinforcement learning with misspecified model classes
Joshua Joseph, Alborz Geramifard, John W. Roberts, Jonathan P. How, and Nicholas Roy · 2013
Cited alongside, same era.
Iterative value-aware model learning
Amir-massoud Farahmand · 2018
Later among the works it cites.
TreeQN and ATreeC: Differentiable tree planning for deep reinforcement learning
Gregory Farquhar, Tim Rocktaeschel, Maximilian Igl, and Shimon Whiteson · 2018
Later among the works it cites.
Addressing function approximation error in actor-critic methods
Scott Fujimoto, Herke Van Hoof, and David Meger · 2018
Later among the works it cites.
Recurrent world models facilitate policy evolution
David Ha and Jürgen Schmidhuber · 2018
Later among the works it cites.
Optimality and approximation with policy gradient methods in Markov Decision Processes
Alekh Agarwal, Sham M. Kakade, Jason D. Lee, and Gaurav Mahajan · 2019
Later among the works it cites.
Global optimality guarantees for policy gradient methods
Jalaj Bhandari and Daniel Russo · 2019
Later among the works it cites.
Information-theoretic considerations in batch reinforcement learning
Jinglin Chen and Nan Jiang · 2019
Later among the works it cites.
Neural proximal/trust region policy optimization attains globally optimal policy
Boyi Liu, Qi Cai, Zhuoran Yang, and Zhaoran Wang · 2019
Later among the works it cites.
Algorithmic framework for model-based deep reinforcement learning with theoretical guarantees
Yuping Luo, Huazhe Xu, Yuanzhi Li, Yuandong Tian, Trevor Darrell, and Tengyu Ma · 2019
Later among the works it cites.
Monte carlo gradient estimation in machine learning
Shakir Mohamed, Mihaela Rosca, Michael Figurnov, and Andriy Mnih · 2019
Later among the works it cites.
Mastering atari, go, chess and shogi by planning with a learned model
Julian Schrittwieser, Ioannis Antonoglou, Thomas Hubert, Karen Simonyan, Laurent Sifre, Simon Schmitt, Arthur Guez, Edward Lockhart, Demis Hassabis, Thore Graepel, et al · 2019
Later among the works it cites.
Model-based reinforcement learning with value-targeted regression
Alex Ayoub, Zeyu Jia, Csaba Szepesvari, Mengdi Wang, and Lin F Yang · 2020
Closest in time.
Gradient-aware model-based policy search
Pierluca D’Oro, Alberto Maria Metelli, Andrea Tirinzoni, Matteo Papini, and Marcello Restelli · 2020
Closest in time.
Adaptive trust region policy optimization: Global convergence and faster rates for regularized MDPs
Lior Shani, Yonathan Efroni, and Shie Mannor · 2020
Closest in time.
Improving sample complexity bounds for actor-critic algorithms
Tengyu Xu, Zhe Wang, and Yingbin Liang · 2020
Closest in time.