Fetching the paper…
Reading the bibliography…
We study derivative-free methods for policy optimization over the class of linear policies.
Contributions to the theory of optimal control
Rudolf E Kalman · 1960
Earlier work this paper cites.
A topological property of real analytic subsets
Stanislaw Lojasiewicz · 1963
Earlier work this paper cites.
Gradient methods for solving equations and inequalities
Boris T Polyak · 1964
Earlier work this paper cites.
A bound on tail probabilities for quadratic forms in independent random variables
David Lee Hanson and Farroll Tim Wright · 1971
Earlier work this paper cites.
Farrol Tim Wright · 1973
Earlier work this paper cites.
Optimal control: Basics and Beyond
Peter Whittle · 1996
Earlier work this paper cites.
PAC adaptive control of linear systems
Claude-Nicolas Fiechter · 1997
Earlier work this paper cites.
System identification
Lennart Ljung · 1998
Earlier work this paper cites.
Dynamic programming and optimal control. Vol. I
Dimitri P Bertsekas · 2005
Earlier work this paper cites.
Online convex optimization in the bandit setting: Gradient descent without a gradient
Abraham Flaxman, Adam Kalai, and Brendan McMahan · 2005
Earlier work this paper cites.
Introduction to stochastic search and optimization: estimation, simulation, and control
James C Spall · 2005
Earlier work this paper cites.
Optimal algorithms for online convex optimization with multi-point bandit feedback
Alekh Agarwal, Ofer Dekel, and Lin Xiao · 2010
Earlier work this paper cites.
Probability: theory and examples
Rick Durrett · 2010
Earlier work this paper cites.
Regret bounds for the adaptive control of linear quadratic systems
Yasin Abbasi-Yadkori and Csaba Szepesvári · 2011
Earlier work this paper cites.
Random gradient-free minimization of convex functions
Yurii Nesterov · 2011
Earlier work this paper cites.
Learning to control a low-cost manipulator using data-efficient reinforcement learning
Mark P Deisenroth, Carl E Rasmussen, and Dieter Fox · 2012
Earlier work this paper cites.
A tail inequality for quadratic forms of subgaussian random vectors
Daniel Hsu, Sham Kakade, and Tong Zhang · 2012
Earlier work this paper cites.
Efficient reinforcement learning for high dimensional linear quadratic systems
Morteza Ibrahimi, Adel Javanmard, and Benjamin V. Roy · 2012
Cited alongside, same era.
Query complexity of derivative-free optimization
Kevin G Jamieson, Robert Nowak, and Ben Recht · 2012
Cited alongside, same era.
Stochastic first- and zeroth-order methods for nonconvex stochastic programming
Saeed Ghadimi and Guanghui Lan · 2013
Cited alongside, same era.
On the complexity of bandit and derivative-free stochastic convex optimization
Ohad Shamir · 2013
Cited alongside, same era.
Optimal rates for zero-order convex optimization: The power of two function evaluations
John C Duchi, Michael I Jordan, Martin J Wainwright, and Andre Wibisono · 2015
Cited alongside, same era.
Finite time analysis of optimal adaptive policies for linear-quadratic systems
Mohamad Kazem Shirani Faradonbeh, Ambuj Tewari, and George Michailidis · 2017
Later among the works it cites.
Towards generalization and simplicity in continuous control
Aravind Rajeswaran, Kendall Lowrey, Emanuel V Todorov, and Sham M Kakade · 2017
Later among the works it cites.
An optimal algorithm for bandit and zero-order convex optimization with two-point feedback
Ohad Shamir · 2017
Later among the works it cites.
Evolution strategies as a scalable alternative to reinforcement learning
Tim Salimans, Jonathan Ho, Xi Chen, Szymon Sidor, and Ilya Sutskever · 2017
Later among the works it cites.
Domain randomization for transferring deep neural networks from simulation to the real world
Josh Tobin, Rachel Fong, Alex Ray, Jonas Schneider, Wojciech Zaremba, and Pieter Abbeel · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih et al · 2015
Cited alongside, same era.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Cited alongside, same era.
Continuous deep Q-learning with model-based acceleration
Shixiang Gu, Timothy Lillicrap, Ilya Sutskever, and Sergey Levine · 2016
Cited alongside, same era.
Linear convergence of gradient and proximal-gradient methods under the polyak-lojasiewicz condition
Hamed Karimi, Julie Nutini, and Mark Schmidt · 2016
Cited alongside, same era.
End-to-end training of deep visuomotor policies
Sergey Levine, Chelsea Finn, Trevor Darrell, and Pieter Abbeel · 2016
Cited alongside, same era.
Mastering the game of Go with deep neural networks and tree search
David Silver et al · 2016
Cited alongside, same era.
Later among the works it cites.
Improved regret bounds for Thompson sampling in linear quadratic control problems
Marc Abeille and Alessandro Lazaric · 2018
Closest in time.
Regret bounds for model-free linear quadratic control
Yasin Abbasi-Yadkori, Nevena Lazic, and Csaba Szepesvári · 2018
Closest in time.
Online linear quadratic control
Alon Cohen, Avinatan Hasidim, Tomer Koren, Nevena Lazic, Yishay Mansour, and Kunal Talwar · 2018
Closest in time.
Regret bounds for robust adaptive control of the linear quadratic regulator
Sarah Dean, Horia Mania, Nikolai Matni, Benjamin Recht, and Stephen Tu · 2018
Closest in time.
Global convergence of policy gradient methods for the linear quadratic regulator
Maryam Fazel, Rong Ge, Sham Kakade, and Mehran Mesbahi · 2018
Closest in time.
The gap between model-based and model-free methods on the linear quadratic regulator: An asymptotic viewpoint
Stephen Tu and Benjamin Recht · 2018
Closest in time.
Least-squares temporal difference learning for the linear quadratic regulator
Stephen Tu and Benjamin Recht · 2018
Closest in time.
Optimization of smooth functions with noisy observations: Local minimax rates
Yining Wang, Sivaraman Balakrishnan, and Aarti Singh · 2018
Closest in time.
Stochastic zeroth-order optimization in high dimensions
Yining Wang, Simon S Du, Sivaraman Balakrishnan, and Aarti Singh · 2018
Closest in time.
Learning linear-quadratic regulators efficiently with only T \sqrt{T} regret
Alon Cohen, Tomer Koren, and Yishay Mansour · 2019
Closest in time.
A short note on concentration inequalities for random vectors with subgaussian norm
Chi Jin, Praneeth Netrapalli, Rong Ge, Sham M. Kakade, and Michael I. Jordan · 2019
Closest in time.