Fetching the paper…
Reading the bibliography…
This paper investigates how to efficiently transition and update policies, trained initially with demonstrations, using off-policy actor-critic reinforcement learning.
On the theory of the Brownian motion
George E Uhlenbeck and Leonard S Ornstein. 1930 · 1930
Earlier work this paper cites.
US Military Specification MIL–F–8785C
D Moorhouse and R Woodcock. 1980 · 1980
Earlier work this paper cites.
Learning from Demonstration. In Proceedings of the 9th International Conference on Neural Information Processing Systems
Stefan Schaal. 1996 · 1996
Earlier work this paper cites.
Robot learning from demonstration. In ICML
Christopher G Atkeson and Stefan Schaal. 1997 · 1997
Earlier work this paper cites.
Learning agents for uncertain environments (extended abstract)
S Russell. 1998 · 1998
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Richard Sutton and Andrew Barto. 1998 · 1998
Earlier work this paper cites.
A Natural Policy Gradient. In Proceedings of the 14th International Conference on Neural Information Processing Systems: Natural and Synthetic
Sham Kakade. 2001 · 2001
Earlier work this paper cites.
A survey of robot learning from demonstration
Brenna D Argall, Sonia Chernova, Manuela Veloso, and Brett Browning. 2009 · 2009
Earlier work this paper cites.
Interactively shaping agents via human reinforcement: The TAMER framework. In Proceedings of the fifth international conference on Knowledge capture
W Bradley Knox and Peter Stone. 2009 · 2009
Earlier work this paper cites.
Learning from Limited Demonstrations
Beomjoon Kim, Amir-massoud Farahmand, Joelle Pineau, and Doina Precup. 2013 · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba. 2014 · 2014
Earlier work this paper cites.
Deterministic policy gradient algorithms
David Silver, Guy Lever, Nicolas Heess, Thomas Degris, Daan Wierstra, and Martin Riedmiller. 2014 · 2014
Earlier work this paper cites.
Fast and accurate deep network learning by exponential linear units (elus)
Djork-Arné Clevert, Thomas Unterthiner, and Sepp Hochreiter. 2015 · 2015
Cited alongside, same era.
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra. 2015 · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei a Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis. 2015 · 2015
Cited alongside, same era.
Prioritized Experience Replay
Tom Schaul, John Quan, Ioannis Antonoglou, and David Silver. 2015 · 2015
Cited alongside, same era.
Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor. In Proceedings of the 35th International Conference on Machine Learning
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine. 2018 · 2018
Later among the works it cites.
Deep q-learning from demonstrations. In Thirty-Second AAAI Conference on Artificial Intelligence
Todd Hester, Matej Vecerik, Olivier Pietquin, Marc Lanctot, Tom Schaul, Bilal Piot, Dan Horgan, John Quan, Andrew Sendonaris, Ian Osband, et al · 2018
Later among the works it cites.
Stable Baselines
Ashley Hill, Antonin Raffin, Maximilian Ernestus, Adam Gleave, Rene Traore, Prafulla Dhariwal, Christopher Hesse, Oleg Klimov, Alex Nichol, Matthias Plappert, Alec Radford, John Schulman, Szymon Sidor, and Yuhuai Wu. 2018 · 2018
Later among the works it cites.
Policy optimization with demonstrations. In International Conference on Machine Learning
Bingyi Kang, Zequn Jie, and Jiashi Feng. 2018 · 2018
Later among the works it cites.
Overcoming Exploration in Reinforcement Learning with Demonstrations. In 2018 IEEE International Conference on Robotics and Automation (ICRA)
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba. 2016 · 2016
Cited alongside, same era.
Generative adversarial imitation learning. In Advances in neural information processing systems
Jonathan Ho and Stefano Ermon. 2016 · 2016
Cited alongside, same era.
Interactive learning from policy-dependent human feedback. In Proceedings of the 34th International Conference on Machine Learning-Volume 70
James MacGlashan, Mark K Ho, Robert Loftin, Bei Peng, Guan Wang, David L Roberts, Matthew E Taylor, and Michael L Littman. 2017 · 2017
Cited alongside, same era.
Shital Shah, Debadeepta Dey, Chris Lovett, and Ashish Kapoor. 2017 · 2017
Cited alongside, same era.
Leveraging demonstrations for deep reinforcement learning on robotics problems with sparse rewards
Matej Večerík, Todd Hester, Jonathan Scholz, Fumin Wang, Olivier Pietquin, Bilal Piot, Nicolas Heess, Thomas Rothörl, Thomas Lampe, and Martin Riedmiller. 2017 · 2017
Cited alongside, same era.
Learning dexterous in-hand manipulation
Marcin Andrychowicz, Bowen Baker, Maciek Chociej, Rafal Jozefowicz, Bob McGrew, Jakub Pachocki, Arthur Petron, Matthias Plappert, Glenn Powell, Alex Ray, et al · 2018
Cited alongside, same era.
Reinforcement Learning from Imperfect Demonstrations
Yang Gao, Huazhe Xu, Ji Lin, Fisher Yu, Sergey Levine, and Trevor Darrell. 2018 · 2018
Cited alongside, same era.
Vinicius G. Goecks, Gregory M. Gremillion, Vernon J. Lawhern, John Valasek, and Nicholas R. Waytowich. 2018 · 2018
Cited alongside, same era.
A. Nair, B. McGrew, M. Andrychowicz, W. Zaremba, and P. Abbeel. 2018 · 2018
Later among the works it cites.
Observe and Look Further: Achieving Consistent Performance on Atari
Tobias Pohlen, Bilal Piot, Todd Hester, Mohammad Gheshlaghi Azar, Dan Horgan, David Budden, Gabriel Barth-Maron, Hado van Hasselt, John Quan, Mel Vecerík, Matteo Hessel, Rémi Munos, and Olivier Pietquin. 2018 · 2018
Later among the works it cites.
DAPG for Dexterous Hand Manipulation
Aravind Rajeswaran, Vikash Kumar, Abhishek Gupta, Giulia Vezzani, John Schulman, Emanuel Todorov, and Sergey Levine. 2018a · 2018
Later among the works it cites.
Learning Complex Dexterous Manipulation with Deep Reinforcement Learning and Demonstrations. In Proceedings of Robotics: Science and Systems
Aravind Rajeswaran, Vikash Kumar, Abhishek Gupta, Giulia Vezzani, John Schulman, Emanuel Todorov, and Sergey Levine. 2018b · 2018
Later among the works it cites.
Deep tamer: Interactive agent shaping in high-dimensional state spaces. In Thirty-Second AAAI Conference on Artificial Intelligence
Garrett Warnell, Nicholas Waytowich, Vernon Lawhern, and Peter Stone. 2018 · 2018
Later among the works it cites.
Cycle-of-Learning for Autonomous Systems from Human Interaction
Nicholas R. Waytowich, Vinicius G. Goecks, and Vernon J. Lawhern. 2018 · 2018
Later among the works it cites.
Learn What Not to Learn: Action Elimination with Deep Reinforcement Learning
Tom Zahavy, Matan Haroush, Nadav Merlis, Daniel J Mankowitz, and Shie Mannor. 2018 · 2018
Later among the works it cites.
Trial without error: Towards safe reinforcement learning via human intervention. In Proceedings of the 17th International Conference on Autonomous Agents and MultiAgent Systems
William Saunders, Girish Sastry, Andreas Stuhlmueller, and Owain Evans. 2018 · 2069
Closest in time.