Fetching the paper…
Reading the bibliography…
While reinforcement learning algorithms provide automated acquisition of optimal policies, practical application of such methods requires a number of design decisions, such as manually designing reward functions that not only define the task, but also provide sufficient shaping to accomplish it.
Numerical computation of optimal control problems with unknown final time
G.M Aly and W.C Chan · 1974
Earlier work this paper cites.
A covariance control theory
Anthony F. Hotz and Robert E. Skelton · 1987
Earlier work this paper cites.
Optimal control of two point boundary value problems
Manfred Goebel and Uldis Raitums · 1990
Earlier work this paper cites.
Q-learning
Christopher JCH Watkins and Peter Dayan · 1992
Earlier work this paper cites.
An improved covariance assignment theory for discrete systems, 1992
J.-H. Xu and R.E. Skelton · 1992
Earlier work this paper cites.
Learning to achieve goals
Leslie P Kaelbling · 1993
Earlier work this paper cites.
Towards fully probabilistic control design
Miroslav Kárnỳ · 1996
Earlier work this paper cites.
Minimum-energy covariance controllers
Karolos M. Grigoriadis and Robert E. Skelton · 1997
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Richard S Sutton and Andrew G Barto · 1998
Earlier work this paper cites.
Planning by probabilistic inference
H. Attias · 2003
Earlier work this paper cites.
Linearly-solvable Markov decision problems
E. Todorov · 2006
Earlier work this paper cites.
Probabilistic inference for solving discrete and continuous state Markov decision processes
Marc Toussaint and Amos Storkey · 2006
Earlier work this paper cites.
Probabilistic inference for solving (PO)MDPs
Marc Toussaint, Stefan Harmeling, and Amos Storkey · 2006
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
Brian D Ziebart, Andrew Maas, J Andrew Bagnell, and Anind K Dey · 2008
Earlier work this paper cites.
An expectation maximization algorithm for continuous markov decision processes with arbitrary reward
Matthew Hoffman, Nando Freitas, Arnaud Doucet, and Jan Peters · 2009
Earlier work this paper cites.
Robot trajectory optimization using approximate inference
M. Toussaint · 2009
Earlier work this paper cites.
Relative entropy policy search
Jan Peters, Katharina Mülling, and Yasemin Altün · 2010
Earlier work this paper cites.
An approximate inference approach to temporal optimization in optimal control
Konrad Rawlik, Marc Toussaint, and Sethu Vijayakumar · 2010
Earlier work this paper cites.
Algorithms for Reinforcement Learning , volume 4
Csaba Szepesvári · 2010
Cited alongside, same era.
Optimal control as a graphical model inference problem
H. J. Kappen, V. Gómez, and M. Opper · 2012
Cited alongside, same era.
On stochastic optimal control and reinforcement learning by approximate inference
K. Rawlik, M. Toussaint, and S. Vijayakumar · 2013
Cited alongside, same era.
Universal value function approximators
Tom Schaul, Daniel Horgan, Karol Gregor, and David Silver · 2015
Cited alongside, same era.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Cited alongside, same era.
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2016
Multi-goal reinforcement learning: Challenging robotics environments and request for research
Matthias Plappert, Marcin Andrychowicz, Alex Ray, Bob Mcgrew, Bowen Baker, Glenn Powell, Jonas Schneider, Josh Tobin, Maciek Chociej, Peter Welinder, Vikash Kumar, and Wojciech Zaremba · 2018
Later among the works it cites.
Temporal difference models: Model-free deep RL for model-based control
Vitchyr Pong, Shixiang Gu, Murtaza Dalal, and Sergey Levine · 2018
Later among the works it cites.
Search on the replay buffer: Bridging planning and reinforcement learning
Benjamin Eysenbach, Ruslan Salakhutdinov, and Sergey Levine · 2019
Later among the works it cites.
VIREL: A variational inference framework for reinforcement learning
Matthew Fellows, Anuj Mahajan, Tim G. J. Rudner, and Shimon Whiteson · 2019
Later among the works it cites.
Information asymmetry in KL-regularized RL
Alexandre Galashov, Siddhant M. Jayakumar, Leonard Hasenclever, Dhruva Tirumala, Jonathan Schwarz, Guillaume Desjardins, Wojciech M. Czarnecki, Yee Whye Teh, Razvan Pascanu, and Nicolas Heess · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Hindsight experience replay
Marcin Andrychowicz, Filip Wolski, Alex Ray, Jonas Schneider, Rachel Fong, Peter Welinder, Bob Mcgrew, Josh Tobin, Pieter Abbeel, and Wojciech Zaremba · 2017
Cited alongside, same era.
Self-supervised visual planning with temporal skip connections
Frederik Ebert, Chelsea Finn, Alex X Lee, and Sergey Levine · 2017
Cited alongside, same era.
Learning multi-level hierarchies with hindsight
Andrew Levy, George Konidaris, Robert Platt, and Kate Saenko · 2017
Cited alongside, same era.
Unifying task specification in reinforcement learning
Martha White · 2017
Cited alongside, same era.
Optimal steering of a linear stochastic system to a final probability distribution—part iii
Yongxin Chen, Tryphon T. Georgiou, and Michele Pavon · 2018
Cited alongside, same era.
Deep reinforcement learning in a handful of trials using probabilistic dynamics models
Kurtland Chua, Roberto Calandra, Rowan Mcallister, and Sergey Levine · 2018
Cited alongside, same era.
Later among the works it cites.
When to trust your model: Model-based policy optimization
Michael Janner, Justin Fu, Marvin Zhang, and Sergey Levine · 2019
Later among the works it cites.
Skew-Fit: State-covering self-supervised reinforcement learning
Vitchyr H. Pong, Murtaza Dalal, Steven Lin, Ashvin Nair, Shikhar Bahl, and Sergey Levine · 2019
Later among the works it cites.
Nonlinear uncertainty control with iterative covariance steering, 2019
Jack Ridderhof, Kazuhide Okamoto, and Panagiotis Tsiotras · 2019
Later among the works it cites.
End-to-end robotic reinforcement learning without reward engineering
Avi Singh, Larry Yang, Kristian Hartikainen, Chelsea Finn, and Sergey Levine · 2019
Later among the works it cites.
Self-supervised learning of distance functions for goal-conditioned reinforcement learning
Srinivas Venkattaramanujam, Eric Crawford, Thang Doan, and Doina Precup · 2019
Later among the works it cites.
Unsupervised control through non-parametric discriminative rewards
David Warde-Farley, Tom Van De Wiele, Tejas Kulkarni, Catalin Ionescu, Steven Hansen, & Volodymyr, and Mnih Deepmind · 2019
Later among the works it cites.
SOLAR: Deep structured representations for model-based reinforcement learning
Marvin Zhang, Sharad Vikram, Laura Smith, Pieter Abbeel, Matthew J. Johnson, and Sergey Levine · 2019
Later among the works it cites.
Dynamical distance learning for unsupervised and semi-supervised skill discovery
Kristian Hartikainen, Xinyang Geng, Tuomas Haarnoja, and Sergey Levine · 2020
Later among the works it cites.
Yannick Schroecker and Charles Isbell · 2020
Later among the works it cites.
Entity abstraction in visual model-based reinforcement learning
Rishi Veerapaneni, John D Co-Reyes, Michael Chang, Michael Janner, Chelsea Finn, Jiajun Wu, Joshua Tenenbaum, and Sergey Levine · 2020
Later among the works it cites.
Nonlinear covariance control via differential dynamic programming
Z. Yi, Z. Cao, E. Theodorou, and Y. Chen · 2020
Later among the works it cites.
Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning
Tianhe Yu, Deirdre Quillen, Zhanpeng He, Ryan Julian, Karol Hausman, Chelsea Finn, and Sergey Levine · 2020
Later among the works it cites.
Variational empowerment as representation learning for goal-based reinforcement learning
Jongwook Choi, Archit Sharma, Honglak Lee, Sergey Levine, and Shixiang Shane Gu · 2021
Closest in time.