Fetching the paper…
Reading the bibliography…
The effectiveness of model-based versus model-free methods is a long-standing question in reinforcement learning (RL).
On the Statistical Treatment of Linear Stochastic Difference Equations
Henry B. Mann and Abraham Wald · 1943
Earlier work this paper cites.
The Stabilizing Solution of the Discrete Algebraic Riccati Equation
B. Molinari · 1975
Earlier work this paper cites.
The expectation of products of quadratic forms in normal variables: the practice
Jan R. Magnus · 1979
Earlier work this paper cites.
Mixing properties of ARMA processes
Abdelkader Mokkadem · 1988
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J. Williams · 1992
Earlier work this paper cites.
Probability and Measure
Patrick Billingsley · 1995
Earlier work this paper cites.
Linear Least-Squares Algorithms for Temporal Difference Learning
Steven J. Bradtke and Andrew G. Barto · 1996
Earlier work this paper cites.
PAC Adaptive Control of Linear Systems
Claude-Nicolas Fiechter · 1997
Earlier work this paper cites.
Metric Entropy of the Grassmann Manifold
Alain Pajor · 1998
Earlier work this paper cites.
Least-Squares Temporal Difference Learning
Justin Boyan · 1999
Earlier work this paper cites.
Stochastic Approximation and Recursive Algorithms and Applications
Harold Kushner and George Yin · 2003
Earlier work this paper cites.
Least-Squares Policy Iteration
Michail G. Lagoudakis and Ronald Parr · 2003
Earlier work this paper cites.
On the Markov chain central limit theorem
Galin L. Jones · 2004
Earlier work this paper cites.
On the Kronecker Product
Kathrin Schäcke · 2004
Earlier work this paper cites.
The Schur Complement and its Applications , volume 4 of Numerical Methods and Algorithms
Fuzhen Zhang · 2005
Earlier work this paper cites.
PAC Model-Free Reinforcement Learning
Alexander L. Strehl, Lihong Li, Eric Wiewiora, John Langford, and Michael L. Littman · 2006
Earlier work this paper cites.
Reinforcement learning of motor skills with policy gradients
Jan Peters and Stefan Schaal · 2008
Cited alongside, same era.
Reinforcement Learning in Finite MDPs: PAC Analysis
Alexander L. Strehl, Lihong Li, and Michael L. Littman · 2009
Cited alongside, same era.
Near-optimal Regret Bounds for Reinforcement Learning
Thomas Jaksch, Ronald Ortner, and Peter Auer · 2010
Cited alongside, same era.
Regret Bounds for the Adaptive Control of Linear Quadratic Systems
Yasin Abbasi-Yadkori and Csaba Szepesvári · 2011
Cited alongside, same era.
Efficient Reinforcement Learning for High Dimensional Linear Quadratic Systems
Morteza Ibrahimi, Adel Javanmard, and Benjamin Van Roy · 2012
Cited alongside, same era.
Query Complexity of Derivative-Free Optimization
Kevin G. Jamieson, Robert D. Nowak, and Benjamin Recht · 2012
Cited alongside, same era.
Evolution Strategies as a Scalable Alternative to Reinforcement Learning
Tim Salimans, Jonathan Ho, Xi Chen, Szymon Sidor, and Ilya Sutskever · 2017
Later among the works it cites.
Asymptotic and finite-sample properties of estimators based on stochastic gradients
Panos Toulis and Edoardo M. Airoldi · 2017
Later among the works it cites.
Model-Free Linear Quadratic Control via Reduction to Expert Prediction
Yasin Abbasi-Yadkori, Nevena Lazić, and Csaba Szepesvári · 2018
Closest in time.
Improved Regret Bounds for Thompson Sampling in Linear Quadratic Control Problems
Marc Abeille and Alessandro Lazaric · 2018
Closest in time.
Model-Based Reinforcement Learning via Meta-Policy Optimization
Ignasi Clavera, Jonas Rothfuss, John Schulman, Yasuhiro Fujita, Tamim Asfour, and Pieter Abbeel · 2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Making Gradient Descent Optimal for Strongly Convex Stochastic Optimization
Alexander Rakhlin, Ohad Shamir, and Karthik Sridharan · 2012
Cited alongside, same era.
Near-optimal PAC bounds for discounted MDPs
Tor Lattimore and Marcus Hutter · 2014
Cited alongside, same era.
High-Dimensional Continuous Control Using Generalized Advantage Estimation
John Schulman, Philipp Moritz, Sergey Levine, Michael Jordan, and Pieter Abbeel · 2016
Cited alongside, same era.
Thompson Sampling for Linear-Quadratic Control Problems
Marc Abeille and Alessandro Lazaric · 2017
Cited alongside, same era.
Optimistic posterior sampling for reinforcement learning: worst-case regret bounds
Shipra Agrawal and Randy Jia · 2017
Cited alongside, same era.
Minimax Regret Bounds for Reinforcement Learning
Mohammad Gheshlaghi Azar, Ian Osband, and Rémi Munos · 2017
Cited alongside, same era.
Regret Bounds for Robust Adaptive Control of the Linear Quadratic Regulator
Sarah Dean, Horia Mania, Nikolai Matni, Benjamin Recht, and Stephen Tu · 2018
Closest in time.
Global Convergence of Policy Gradient Methods for the Linear Quadratic Regulator
Maryam Fazel, Rong Ge, Sham Kakade, and Mehran Mesbahi · 2018
Closest in time.
Is Q-learning Provably Efficient?
Chi Jin, Zeyuan Allen-Zhu, Sebastien Bubeck, and Michael I. Jordan · 2018
Closest in time.
Derivative-Free Methods for Policy Optimization: Guarantees for Linear Quadratic Systems
Dhruv Malik, Ashwin Pananjady, Kush Bhatia, Koulik Khamaru, Peter L. Bartlett, and Martin J. Wainwright · 2018
Closest in time.
Simple random search provides a competitive approach to reinforcement learning
Horia Mania, Aurelia Guy, and Benjamin Recht · 2018
Closest in time.
Neural Network Dynamics for Model-Based Deep Reinforcement Learning with Model-Free Fine-Tuning
Anusha Nagabandi, Gregory Kahn, Ronald S. Fearing, and Sergey Levine · 2018
Closest in time.
Temporal Difference Models: Model-Free Deep RL for Model-Based Control
Vitchyr Pong, Shixiang Gu, Murtaza Dalal, and Sergey Levine · 2018
Closest in time.
A Tour of Reinforcement Learning: The View from Continuous Control
Benjamin Recht · 2018
Closest in time.
Learning Without Mixing: Towards A Sharp Analysis of Linear System Identification
Max Simchowitz, Horia Mania, Stephen Tu, Michael I. Jordan, and Benjamin Recht · 2018
Closest in time.
Model-Based Reinforcement Learning in Contextual Decision Processes
Wen Sun, Nan Jiang, Akshay Krishnamurthy, Alekh Agarwal, and John Langford · 2018
Closest in time.
Least-Squares Temporal Difference Learning for the Linear Quadratic Regulator
Stephen Tu and Benjamin Recht · 2018
Closest in time.