Fetching the paper…
Reading the bibliography…
This work studies the problem of sequential control in an unknown, nonlinear dynamical system, where we model the underlying system dynamics as an unknown function in a known Reproducing Kernel Hilbert Space.
Online control with adversarial disturbances
Naman Agarwal, Brian Bullins, Elad Hazan, Sham M. Kakade, and Karan Singh · 1902
Earlier work this paper cites.
Optimality and approximation with policy gradient methods in Markov decision processes
Alekh Agarwal, Sham M. Kakade, Jason D. Lee, and Gaurav Mahajan · 1908
Earlier work this paper cites.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
William R. Thompson · 1933
Earlier work this paper cites.
Differential dynamic programming
David H. Jacobson and David Q. Mayne · 1970
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Michael Kearns and Satinder Singh · 2002
Earlier work this paper cites.
On the sample complexity of reinforcement learning
Sham M. Kakade · 2003
Earlier work this paper cites.
Exploration in metric state spaces
Sham M. Kakade, Michael J. Kearns, and John Langford · 2003
Earlier work this paper cites.
Exploration and apprenticeship learning in reinforcement learning
Pieter Abbeel and Andrew Y Ng · 2005
Earlier work this paper cites.
A generalized iterative LQG method for locally-optimal feedback control of constrained nonlinear stochastic systems
Emanuel Todorov and Weiwei Li · 2005
Earlier work this paper cites.
Autonomous inverted helicopter flight via reinforcement learning
Andrew Y Ng, Adam Coates, Mark Diel, Varun Ganapathi, Jamie Schulte, Ben Tse, Eric Berger, and Eric Liang · 2006
Earlier work this paper cites.
Iterative learning control: Brief survey and categorization
Hyo-Sung Ahn, YangQuan Chen, and Kevin L. Moore · 2007
Earlier work this paper cites.
Stochastic linear optimization under bandit feedback
Varsha Dani, Thomas P. Hayes, and Sham M. Kakade · 2008
Earlier work this paper cites.
Random features for large-scale kernel machines
Ali Rahimi and Benjamin Recht · 2008
Earlier work this paper cites.
Gaussian process optimization in the bandit setting: No regret and experimental design
Niranjan Srinivas, Andreas Krause, Sham M. Kakade, and Matthias Seeger · 2009
Earlier work this paper cites.
LQR-trees: Feedback motion planning on sparse randomized trees
Russ Tedrake · 2009
Earlier work this paper cites.
Regret bounds for the adaptive control of linear quadratic systems
Yasin Abbasi-Yadkori and Csaba Szepesvári · 2011
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Yasin Abbasi-Yadkori, Dávid Pál, and Csaba Szepesvári · 2011
Earlier work this paper cites.
An empirical evaluation of Thompson sampling
Olivier Chapelle and Lihong Li · 2011
Earlier work this paper cites.
PILCO: A model-based and data-efficient approach to policy search
Marc Deisenroth and Carl E. Rasmussen · 2011
Earlier work this paper cites.
Trajectory optimization for full-body movements with complex contacts
Mazen Al Borno, Martin De Lasa, and Aaron Hertzmann · 2012
Earlier work this paper cites.
Discovery of complex behaviors through contact-invariant optimization
Igor Mordatch, Emanuel Todorov, and Zoran Popović · 2012
Earlier work this paper cites.
LQR-RRT*: Optimal sampling-based motion planning with automatically derived extension heuristics
Alejandro Perez, Robert Platt, George Konidaris, Leslie Kaelbling, and Tomas Lozano-Perez · 2012
Cited alongside, same era.
Agnostic system identification for model-based reinforcement learning
Stephane Ross and J. Andrew Bagnell · 2012
Cited alongside, same era.
MuJoCo: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Cited alongside, same era.
Eluder dimension and the sample complexity of optimistic exploration
Daniel Russo and Benjamin Van Roy · 2013
Cited alongside, same era.
Learning neural network policies with guided policy search under unknown dynamics
Sergey Levine and Pieter Abbeel · 2014
Cited alongside, same era.
Model-based reinforcement learning and the Eluder dimension
Ian Osband and Benjamin Van Roy · 2014
Deep reinforcement learning in a handful of trials using probabilistic dynamics models
Kurtland Chua, Roberto Calandra, Rowan McAllister, and Sergey Levine · 2018
Later among the works it cites.
Online linear quadratic control
Alon Cohen, Avinatan Hassidim, Tomer Koren, Nevena Lazic, Yishay Mansour, and Kunal Talwar · 2018
Later among the works it cites.
Regret bounds for robust adaptive control of the linear quadratic regulator
Sarah Dean, Horia Mania, Nikolai Matni, Benjamin Recht, and Stephen Tu · 2018
Later among the works it cites.
Model-ensemble trust-region policy optimization
Thanard Kurutach, Ignasi Clavera, Yan Duan, Aviv Tamar, and Pieter Abbeel · 2018
Later among the works it cites.
Plan online, learn offline: Efficient learning and exploration via model-based control
Kendall Lowrey, Aravind Rajeswaran, Sham M. Kakade, Emanuel Todorov, and Igor Mordatch · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Learning to optimize via posterior sampling
Daniel Russo and Benjamin Van Roy · 2014
Cited alongside, same era.
Understanding machine learning: From theory to algorithms
Shai Shalev-Shwartz and Shai Ben-David · 2014
Cited alongside, same era.
Sample complexity of episodic fixed-horizon reinforcement learning
Christoph Dann and Emma Brunskill · 2015
Cited alongside, same era.
Ensemble-CIO: Full-body dynamic motion planning that transfers to physical humanoids
Igor Mordatch, Kendall Lowrey, and Emanuel Todorov · 2015
Cited alongside, same era.
Unifying count-based exploration and intrinsic motivation
Marc Bellemare, Sriram Srinivasan, Georg Ostrovski, Tom Schaul, David Saxton, and Remi Munos · 2016
Cited alongside, same era.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Cited alongside, same era.
Algorithmic framework for model-based deep reinforcement learning with theoretical guarantees
Yuping Luo, Huazhe Xu, Yuanzhi Li, Yuandong Tian, Trevor Darrell, and Tengyu Ma · 2018
Later among the works it cites.
Neural network dynamics for model-based deep reinforcement learning with model-free fine-tuning
Anusha Nagabandi, Gregory Kahn, Ronald S. Fearing, and Sergey Levine · 2018
Later among the works it cites.
Solving rubik’s cube with a robot hand
Ilge Akkaya, Marcin Andrychowicz, Maciek Chociej, Mateusz Litwin, Bob McGrew, Arthur Petron, Alex Paino, Matthias Plappert, Glenn Powell, Raphael Ribas, Jonas Schneider, Nikolas Tezak, Jerry Tworek, Peter Welinder, Lilian Weng, Qiming Yuan, Wojciech Zaremba, and Lei Zhang · 2019
Later among the works it cites.
Learning linear-quadratic regulators efficiently with only sqrtT regret
Alon Cohen, Tomer Koren, and Yishay Mansour · 2019
Later among the works it cites.
The nonstochastic control problem
Elad Hazan, Sham M. Kakade, and Karan Singh · 2019
Later among the works it cites.
Information-theoretic confidence bounds for reinforcement learning
Xiuyuan Lu and Benjamin Van Roy · 2019
Later among the works it cites.
Certainty equivalent control of LQR is efficient
Horia Mania, Stephen Tu, and Benjamin Recht · 2019
Later among the works it cites.
Model-based RL in contextual decision processes: PAC bounds and exponential improvements over model-free approaches
Wen Sun, Nan Jiang, Akshay Krishnamurthy, Alekh Agarwal, and John Langford · 2019
Later among the works it cites.
An online learning approach to model predictive control
Nolan Wagener, Ching-An Cheng, Jacob Sacks, and Byron Boots · 2019
Later among the works it cites.
Exploring model-based planning with policy networks
Tingwu Wang and Jimmy Ba · 2019
Later among the works it cites.
Benchmarking model-based reinforcement learning
Tingwu Wang, Xuchan Bao, Ignasi Clavera, Jerrick Hoang, Yeming Wen, Eric Langlois, Shunshi Zhang, Guodong Zhang, Pieter Abbeel, and Jimmy Ba · 2019
Later among the works it cites.
Model-based reinforcement learning with value-targeted regression
Alex Ayoub, Zeyu Jia, Csaba Szepesvari, Mengdi Wang, and Lin F Yang · 2020
Closest in time.
Active learning for nonlinear system identification with guarantees
Horia Mania, Michael I. Jordan, and Benjamin Recht · 2020
Closest in time.
Naive exploration is optimal for online LQR
Max Simchowitz and Dylan J Foster · 2020
Closest in time.
Lyceum: An efficient and scalable ecosystem for robot learning
Colin Summers, Kendall Lowrey, Aravind Rajeswaran, Siddhartha Srinivasa, and Emanuel Todorov · 2020
Closest in time.