Fetching the paper…
Reading the bibliography…
Many continuous control tasks have easily formulated objectives, yet using them directly as a reward in reinforcement learning (RL) leads to suboptimal policies.
Reinforcement learning is direct adaptive optimal control
Richard S Sutton, Andrew G Barto, and Ronald J Williams · 1992
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
Andrew Y. Ng and Stuart J. Russell · 2000
Earlier work this paper cites.
Evolutionary algorithms and reinforcement learning: A comparative case study for architecture search
Esteban Real, Alok Aggarwal, Yanping Huang, and Quoc V. Le · 2002
Earlier work this paper cites.
Reward Shaping , pages 863–865
Eric Wiewiora · 2010
Earlier work this paper cites.
Infinite horizon model predictive control for nonlinear periodic tasks
Tom Erez, Yuval Tassa, and Emanuel Todorov · 2011
Earlier work this paper cites.
Information-theoretic regret bounds for gaussian process optimization in the bandit setting
Niranjan Srinivas, Andreas Krause, Sham Kakade, and Matthias Seeger · 2011
Earlier work this paper cites.
Synthesis and stabilization of complex behaviors through online trajectory optimization
Y. Tassa, T. Erez, and E. Todorov · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Earlier work this paper cites.
Deepdriving: Learning affordance for direct perception in autonomous driving
Chenyi Chen, Ari Seff, Alain Kornhauser, and Jianxiong Xiao · 2015
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Timothy P. Lillicrap, Jonathan J. Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Earlier work this paper cites.
High-dimensional continuous control using generalized advantage estimation
John Schulman, Philipp Moritz, Sergey Levine, Michael I. Jordan, and Pieter Abbeel · 2015
Earlier work this paper cites.
Openai gym, 2016
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Cited alongside, same era.
End-to-end training of deep visuomotor policies
Sergey Levine, Chelsea Finn, Trevor Darrell, and Pieter Abbeel · 2016
Cited alongside, same era.
Marcin Andrychowicz, Filip Wolski, Alex Ray, Jonas Schneider, Rachel Fong, Peter Welinder, Bob McGrew, Josh Tobin, Pieter Abbeel, and Wojciech Zaremba · 2017
Cited alongside, same era.
Reverse curriculum generation for reinforcement learning
Carlos Florensa, David Held, Markus Wulfmeier, Michael Zhang, and Pieter Abbeel · 2017
Cited alongside, same era.
Google vizier: A service for black-box optimization
Daniel Golovin, Benjamin Solnik, Subhodeep Moitra, Greg Kochanski, John Karro, and D. Sculley · 2017
Cited alongside, same era.
TF-Agents: A library for reinforcement learning in tensorflow
Sergio Guadarrama, Anoop Korattikara, Oscar Ramirez, Pablo Castro, Ethan Holly, Sam Fishman, Ke Wang, Ekaterina Gonina, Chris Harris, Vincent Vanhoucke, and Eugene Brevdo · 2018
Later among the works it cites.
Barc: Backward reachability curriculum for robotic reinforcement learning
Boris Ivanovic, James Harrison, Apoorva Sharma, Mo Chen, and Marco Pavone · 2018
Later among the works it cites.
Scalable deep reinforcement learning for vision-based robotic manipulation
Dmitry Kalashnikov, Alex Irpan, Peter Pastor, Julian Ibarz, Alexander Herzog, Eric Jang, Deirdre Quillen, Ethan Holly, Mrinal Kalakrishnan, Vincent Vanhoucke, and Sergey Levine · 2018
Later among the works it cites.
Evolution-guided policy gradient in reinforcement learning
Shauharda Khadka and Kagan Tumer · 2018
Later among the works it cites.
Follownet: Robot navigation by following natural language directions with deep reinforcement learning
Pararth Shah, Marek Fiser, Aleksandra Faust, Chase Kew, and Dilek Hakkani-Tur · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chenxi Liu, Barret Zoph, Jonathon Shlens, Wei Hua, Li-Jia Li, Li Fei-Fei, Alan L. Yuille, Jonathan Huang, and Kevin Murphy · 2017
Cited alongside, same era.
Large-scale evolution of image classifiers
Esteban Real, Sherry Moore, Andrew Selle, Saurabh Saxena, Yutaka Leon Suematsu, Jie Tan, Quoc V. Le, and Alexey Kurakin · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
Neural architecture search with reinforcement learning
Barret Zoph and Quoc V. Le · 2017
Cited alongside, same era.
Efficient architecture search by network transformation
Han Cai, Tianyao Chen, Weinan Zhang, Yong Yu, and Jun Wang · 2018
Cited alongside, same era.
Learning to walk via deep reinforcement learning
Tuomas Haarnoja, Aurick Zhou, Sehoon Ha, Jie Tan, George Tucker, and Sergey Levine
Cited in the paper.
Soft actor-critic algorithms and applications
Tuomas Haarnoja, Aurick Zhou, Kristian Hartikainen, George Tucker, Sehoon Ha, Jie Tan, Vikash Kumar, Henry Zhu, Abhishek Gupta, Pieter Abbeel, and Sergey Levine
Cited in the paper.
Later among the works it cites.
Tom Silver, Kelsey Allen, Josh Tenenbaum, and Leslie Pack Kaelbling · 2018
Later among the works it cites.
Learning transferable architectures for scalable image recognition
Barret Zoph, Vijay Vasudevan, Jonathon Shlens, and Quoc V. Le · 2018
Later among the works it cites.
Learning transferable architectures for scalable image recognition
Barret Zoph, Vijay Vasudevan, Jonathon Shlens, and Quoc V. Le · 2018
Later among the works it cites.
Learning navigation behaviors end-to-end with autorl
Hao-Tien Lewis Chiang, Aleksandra Faust, Marek Fiser, and Anthony Francis · 2019
Closest in time.
Learning to navigate the web
Izzeddin Gur, Ulrich Rueckert, Aleksandra Faust, and Dilek Hakkani-Tur · 2019
Closest in time.