Fetching the paper…
Reading the bibliography…
The shortcomings of maximum likelihood estimation in the context of model-based reinforcement learning have been highlighted by an increasing number of papers.
A note on certainty equivalence in dynamic planning
H. Theil · 1957
Earlier work this paper cites.
Contraction mappings in the theory underlying dynamic programming
Eric V. Denardo · 1967
Earlier work this paper cites.
Discrete-time markovian decision processes with an unknown parameter-average return criterion
Masami Kurano · 1972
Earlier work this paper cites.
Estimation and control in markov chains
P. Mandl · 1974
Earlier work this paper cites.
Estimation et controle des chaines de markov sur des espaces arbitraires
J. P. Georgin · 1978
Earlier work this paper cites.
Adaptive control of markov chains, i: Finite parameter set
V. Borkar and P. Varaiya · 1979
Earlier work this paper cites.
Optimal harvesting with imprecise parameter estimates
D. Ludwig and C.J. Walters · 1982
Earlier work this paper cites.
Neuronlike adaptive elements that can solve difficult learning control problems
Andrew G Barto, Richard S Sutton, and Charles W Anderson · 1983
Earlier work this paper cites.
Adaptive control of discounted markov decision chains
O. Hernández-Lerma and S. I. Marcus · 1985
Earlier work this paper cites.
Estimation and control in discounted stochastic dynamic programming
Schäl Manfred · 1987
Earlier work this paper cites.
Maximum likelihood estimation of discrete control processes
John Rust · 1988
Earlier work this paper cites.
Learning control of finite markov chains with an explicit trade-off between estimation and control
M. Sato, K. Abe, and H. Takeda · 1988
Earlier work this paper cites.
Model error concepts in control design
RE Skelton · 1989
Earlier work this paper cites.
Learning sequential decision rules using simulation models and competition
John J. Grefenstette, Connie Loggia Ramsey, and Alan C. Schultz · 1990
Earlier work this paper cites.
Dyna, an integrated architecture for learning, planning, and reacting
Richard S Sutton · 1991
Earlier work this paper cites.
Perturbation and stability theory for markov control problems
M. Abbad and J.A. Filar · 1992
Earlier work this paper cites.
Self-improving reactive agents based on reinforcement learning, planning and teaching
Long-Ji Lin · 1992
Earlier work this paper cites.
Q-learning
Christopher JCH Watkins and Peter Dayan · 1992
Earlier work this paper cites.
Reverse accumulation and attractive fixed points
Bruce Christianson · 1994
Earlier work this paper cites.
Linear-quadratic control: an introduction
Peter Dorato, Vito Cerone, and Chaouki Abdallah · 1994
Earlier work this paper cites.
Python tutorial , volume 620
Guido Van Rossum and Fred L Drake Jr · 1995
Earlier work this paper cites.
Optimization of computer simulation models with rare events
Reuven Y Rubinstein · 1997
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David McAllester, Satinder Singh, and Yishay Mansour · 1999
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Michael Kearns and Satinder Singh · 2002
Earlier work this paper cites.
Convex optimization
Stephen Boyd, Stephen P Boyd, and Lieven Vandenberghe · 2004
Earlier work this paper cites.
Robust dynamic programming
Garud N. Iyengar · 2005
Earlier work this paper cites.
Robust control of markov decision processes with uncertain transition matrices
Arnab Nilim and Laurent El Ghaoui · 2005
Earlier work this paper cites.
Towards a unified theory of state abstraction for mdps
Lihong Li, Thomas J Walsh, and Michael L Littman · 2006
Earlier work this paper cites.
A guide to NumPy , volume 1
Travis E Oliphant · 2006
Cited alongside, same era.
Matplotlib: A 2d graphics environment
John D Hunter · 2007
Cited alongside, same era.
Python for scientific computing
Travis E Oliphant · 2007
Cited alongside, same era.
Evaluating Derivatives
Andreas Griewank and Andrea Walther · 2008
Cited alongside, same era.
Maximum entropy inverse reinforcement learning
Brian D Ziebart, Andrew L Maas, J Andrew Bagnell, and Anind K Dey · 2008
Cited alongside, same era.
Double q-learning
Hado Hasselt · 2010
Cited alongside, same era.
Rectified linear units improve restricted boltzmann machines
Deep reinforcement learning in a handful of trials using probabilistic dynamics models
Kurtland Chua, Roberto Calandra, Rowan McAllister, and Sergey Levine · 2018
Later among the works it cites.
Addressing function approximation error in actor-critic methods
Scott Fujimoto, Herke Hoof, and David Meger · 2018
Later among the works it cites.
Lecture notes on statistical reinforcement learning
Nan Jiang · 2018
Later among the works it cites.
Reinforcement learning and control as probabilistic inference: Tutorial and review
Sergey Levine · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Later among the works it cites.
A lagrangian method for inverse problems in reinforcement learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Vinod Nair and Geoffrey E Hinton · 2010
Cited alongside, same era.
Reward design via online gradient ascent
Jonathan Sorg, Richard L Lewis, and Satinder Singh · 2010
Cited alongside, same era.
Distributionally robust markov decision processes
Huan Xu and Shie Mannor · 2010
Cited alongside, same era.
Closing the learning-planning loop with predictive state representations
Byron Boots, Sajid M. Siddiqi, and Geoffrey J. Gordon · 2011
Cited alongside, same era.
The numpy array: a structure for efficient numerical computation
Stefan Van Der Walt, S Chris Colbert, and Gael Varoquaux · 2011
Cited alongside, same era.
The implicit function theorem: history, theory, and applications
Steven G Krantz and Harold R Parks · 2012
Cited alongside, same era.
Pierre-Luc Bacon, Florian Schäfer, Clement Gehring, Animashree Anandkumar, and Emma Brunskill · 2019
Later among the works it cites.
Deep equilibrium models
Shaojie Bai, J Zico Kolter, and Vladlen Koltun · 2019
Later among the works it cites.
The value function polytope in reinforcement learning
Robert Dadashi, Adrien Ali Taiga, Nicolas Le Roux, Dale Schuurmans, and Marc G Bellemare · 2019
Later among the works it cites.
FAX: differentiating fixed point problems in JAX, 2019
Clement Gehring, Pierre-Luc Bacon, and Florian Schaefer · 2019
Later among the works it cites.
Learning latent dynamics for planning from pixels
Danijar Hafner, Timothy Lillicrap, Ian Fischer, Ruben Villegas, David Ha, Honglak Lee, and James Davidson · 2019
Later among the works it cites.
When to trust your model: Model-based policy optimization
Michael Janner, Justin Fu, Marvin Zhang, and Sergey Levine · 2019
Later among the works it cites.
Model-based reinforcement learning for atari
Lukasz Kaiser, Mohammad Babaeizadeh, Piotr Milos, Blazej Osinski, Roy H Campbell, Konrad Czechowski, Dumitru Erhan, Chelsea Finn, Piotr Kozakowski, Sergey Levine, et al · 2019
Later among the works it cites.
Meta-learning with implicit gradients
Aravind Rajeswaran, Chelsea Finn, Sham M Kakade, and Sergey Levine · 2019
Later among the works it cites.
Mastering atari, go, chess and shogi by planning with a learned model
Julian Schrittwieser, Ioannis Antonoglou, Thomas Hubert, Karen Simonyan, Laurent Sifre, Simon Schmitt, Arthur Guez, Edward Lockhart, Demis Hassabis, Thore Graepel, et al · 2019
Later among the works it cites.
Policy-aware model learning for policy gradient methods
Romina Abachi, Mohammad Ghavamzadeh, and Amir-massoud Farahmand · 2020
Later among the works it cites.
The differentiable cross-entropy method
Brandon Amos and Denis Yarats · 2020
Later among the works it cites.
Model-based reinforcement learning with value-targeted regression
Alex Ayoub, Zeyu Jia, Csaba Szepesvari, Mengdi Wang, and Lin F Yang · 2020
Later among the works it cites.
The DeepMind JAX Ecosystem, 2020
Igor Babuschkin, Kate Baumli, Alison Bell, Surya Bhupatiraju, Jake Bruce, Peter Buchlovsky, David Budden, Trevor Cai, Aidan Clark, Ivo Danihelka, Claudio Fantacci, Jonathan Godwin, Chris Jones, Tom Hennigan, Matteo Hessel, Steven Kapturowski, Thomas Keck, Iurii Kemaev, Michael King, Lena Martens, Vladimir Mikulik, Tamara Norman, John Quan, George Papamakarios, Roman Ring, Francisco Ruiz, Alvaro Sanchez, Rosalia Schneider, Eren Sezener, Stephen Spencer, Srivatsan Srinivasan, Wojciech Stokowiec, and Fabio Viola · 2020
Later among the works it cites.
Gradient-aware model-based policy search
Pierluca D’Oro, Alberto Maria Metelli, Andrea Tirinzoni, Matteo Papini, and Marcello Restelli · 2020
Later among the works it cites.
The value equivalence principle for model-based reinforcement learning
Christopher Grimm, André Barreto, Satinder Singh, and David Silver · 2020
Later among the works it cites.
Mastering atari with discrete world models
Danijar Hafner, Timothy Lillicrap, Mohammad Norouzi, and Jimmy Ba · 2020
Later among the works it cites.
Flax: A neural network library and ecosystem for JAX, 2020
Jonathan Heek, Anselm Levskaya, Avital Oliver, Marvin Ritter, Bertrand Rondepierre, Andreas Steiner, and Marc van Zee · 2020
Later among the works it cites.
Objective mismatch in model-based reinforcement learning
Nathan Lambert, Brandon Amos, Omry Yadan, and Roberto Calandra · 2020
Later among the works it cites.
Optimizing millions of hyperparameters by implicit differentiation
Jonathan Lorraine, Paul Vicol, and David Duvenaud · 2020
Later among the works it cites.
A game theoretic framework for model based reinforcement learning
Aravind Rajeswaran, Igor Mordatch, and Vikash Kumar · 2020
Later among the works it cites.
Mopo: Model-based offline policy optimization
Tianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon, James Zou, Sergey Levine, Chelsea Finn, and Tengyu Ma · 2020
Later among the works it cites.
Mt-opt: Continuous multi-task robotic reinforcement learning at scale
Dmitry Kalashnikov, Jacob Varley, Yevgen Chebotar, Benjamin Swanson, Rico Jonschkowski, Chelsea Finn, Sergey Levine, and Karol Hausman · 2021
Closest in time.