Efficient high-dimensional stochastic optimal motion control using tensor-train decomposition
Alex Gorodetsky, Sertac Karaman, and Youssef Marzouk · 2015
Later among the works it cites.
Improving multi-step prediction of learned time series models
Arun Venkatraman, Martial Hebert, and J Andrew Bagnell · 2015
Later among the works it cites.
Guided policy search code implementation, 2016
C. Finn, M. Zhang, J. Fu, X. Tan, Z. McCarthy, E. Scharff, and S. Levine · 2016
Later among the works it cites.
End-to-end training of deep visuomotor policies
Sergey Levine, Chelsea Finn, Trevor Darrell, and Pieter Abbeel · 2016
Later among the works it cites.
Guided policy search via approximate mirror descent
William H Montgomery and Sergey Levine · 2016
Later among the works it cites.
Mastering the game of go with deep neural networks and tree search
David Silver et al · 2016
Later among the works it cites.
Learning to filter with predictive state inference machines
Wen Sun, Arun Venkatraman, Byron Boots, and J Andrew Bagnell · 2016
Later among the works it cites.
Thinking fast and slow with deep learning and tree search
Original
Thomas Anthony, Zheng Tian, and David Barber · 2017
Later among the works it cites.
Reset-free guided policy search: efficient deep reinforcement learning with stochastic initial states
William Montgomery, Anurag Ajay, Chelsea Finn, Pieter Abbeel, and Sergey Levine · 2017
Later among the works it cites.
Mastering the game of go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, et al · 2017
Later among the works it cites.
Deeply aggrevated: Differentiable imitation learning for sequential prediction
Wen Sun, Arun Venkatraman, Geoffrey J Gordon, Byron Boots, and J Andrew Bagnell · 2017
Later among the works it cites.
Simple random search provides a competitive approach to reinforcement learning
Original
Horia Mania, Aurelia Guy, and Benjamin Recht · 2018
Closest in time.