Fetching the paper…
Reading the bibliography…
Much of the current work on reinforcement learning studies episodic settings, where the agent is reset between trials to an initial state distribution, often with well-shaped reward functions.
Catastrophic interference in connectionist networks: The sequential learning problem" the psychology
M. Mccloskey · 1989
Earlier work this paper cites.
Reward functions for accelerated learning
Maja J. Mataric · 1994
Earlier work this paper cites.
Child: A first step towards continual learning
Mark B. Ring · 1997
Earlier work this paper cites.
Learning to drive a bicycle using reinforcement learning and shaping
Jette Randløv and Preben Alstrøm · 1998
Earlier work this paper cites.
Reinforcement learning: An introduction
Richard S. Sutton and Andrew G. Barto · 1998
Earlier work this paper cites.
Catastrophic forgetting in connectionist networks
Robert M French · 1999
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
Andrew Y. Ng, Daishi Harada, and Stuart J. Russell · 1999
Earlier work this paper cites.
Estimation and approximation bounds for gradient-based reinforcement learning
Peter L. Bartlett and Jonathan Baxter · 2000
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Sham M. Kakade and John Langford · 2002
Earlier work this paper cites.
Reinforcement learning in pomdps without resets
Eyal Even-Dar, Sham M. Kakade, and Yishay Mansour · 2005
Earlier work this paper cites.
Curriculum learning
Yoshua Bengio, Jérôme Louradour, Ronan Collobert, and Jason Weston · 2009
Earlier work this paper cites.
Dynamic potential-based reward shaping
Sam Devlin and Daniel Kudenko · 2012
Earlier work this paper cites.
Safe exploration in markov decision processes
Teodor Mihai Moldovan and Pieter Abbeel · 2012
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Marc G. Bellemare, Yavar Naddaf, Joel Veness, and Michael H. Bowling · 2013
Earlier work this paper cites.
Policy shaping: Integrating human feedback with reinforcement learning
Shane Griffith, Kaushik Subramanian, Jonathan Scholz, Charles Lee Isbell, and Andrea Lockerd Thomaz · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Reinforcement learning from demonstration through shaping
Tim Brys, Anna Harutyunyan, Halit Bener Suay, Sonia Chernova, Matthew E. Taylor, and Ann Nowé · 2015
Earlier work this paper cites.
Learning compound multi-step controllers under unknown dynamics
Weiqiao Han, Sergey Levine, and Pieter Abbeel · 2015
Cited alongside, same era.
Deep reinforcement learning with double q-learning
Hado van Hasselt, Arthur Guez, and David Silver · 2015
Cited alongside, same era.
Concrete problems in ai safety
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul F. Christiano, John Schulman, and Dan Mané · 2016
Cited alongside, same era.
Charles Beattie, Joel Z. Leibo, Denis Teplyashin, Tom Ward, Marcus Wainwright, Heinrich Küttler, Andrew Lefrancq, Simon Green, Víctor Valdés, Amir Sadik, Julian Schrittwieser, Keith Anderson, Sarah York, Max Cant, Adam Cain, Adrian Bolton, Stephen Gaffney, Helen King, Demis Hassabis, Shane Legg, and Stig Petersen · 2016
Cited alongside, same era.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Intrinsic motivation and automatic curricula via asymmetric self-play
Sainbayar Sukhbaatar, Ilya Kostrikov, Arthur Szlam, and Rob Fergus · 2017
Later among the works it cites.
Mengdi Wang · 2017
Later among the works it cites.
Exploration by random network distillation
Yuri Burda, Harrison Edwards, Amos J. Storkey, and Oleg Klimov · 2018
Later among the works it cites.
Reset-free trial-and-error learning for robot damage recovery
Konstantinos Chatzilygeroudis, Vassilis Vassiliades, and Jean-Baptiste Mouret · 2018
Later among the works it cites.
Minimalistic gridworld environment for openai gym
Maxime Chevalier-Boisvert, Lucas Willems, and Suman Pal · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Vime: Variational information maximizing exploration
Rein Houthooft, Xi Chen, Yan Duan, John Schulman, Filip De Turck, and Pieter Abbeel · 2016
Cited alongside, same era.
Overcoming catastrophic forgetting in neural networks
James Kirkpatrick, Razvan Pascanu, Neil C. Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A. Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, Demis Hassabis, Claudia Clopath, Dharshan Kumaran, and Raia Hadsell · 2016
Cited alongside, same era.
Deep exploration via bootstrapped dqn
Ian Osband, Charles Blundell, Alexander Pritzel, and Benjamin Van Roy · 2016
Cited alongside, same era.
Andrei A. Rusu, Neil C. Rabinowitz, Guillaume Desjardins, Hubert Soyer, James Kirkpatrick, Koray Kavukcuoglu, Razvan Pascanu, and Raia Hadsell · 2016
Cited alongside, same era.
#exploration: A study of count-based exploration for deep reinforcement learning
Haoran Tang, Rein Houthooft, Davis Foote, Adam Stooke, Xi Chen, Yan Duan, John Schulman, Filip De Turck, and Pieter Abbeel · 2016
Cited alongside, same era.
Continuous adaptation via meta-learning in nonstationary and competitive environments
Maruan Al-Shedivat, Trapit Bansal, Yura Burda, Ilya Sutskever, Igor Mordatch, and Pieter Abbeel · 2017
Cited alongside, same era.
Leave no trace: Learning to reset for safe and autonomous reinforcement learning
Benjamin Eysenbach, Shixiang Gu, Julian Ibarz, and Sergey Levine · 2017
Cited alongside, same era.
Unity: A general platform for intelligent agents, 2018
Arthur Juliani, Vincent-Pierre Berges, Esh Vckay, Yuan Gao, Hunter Henry, Marwan Mattar, and Danny Lange · 2018
Later among the works it cites.
Continual reinforcement learning with complex synapses
Christos Kaplanis, Murray Shanahan, and Claudia Clopath · 2018
Later among the works it cites.
Learning to teach in cooperative multiagent reinforcement learning
Shayegan Omidshafiei, Dong-Ki Kim, Miao Liu, Gerald Tesauro, Matthew Riemer, Christopher Amato, Murray Campbell, and Jonathan P. How · 2018
Later among the works it cites.
Learning by playing solving sparse reward tasks from scratch
Martin A. Riedmiller, Roland Hafner, Thomas Lampe, Michael Neunert, Jonas Degrave, Tom Van de Wiele, Volodymyr Mnih, Nicolas Manfred Otto Heess, and Jost Tobias Springenberg · 2018
Later among the works it cites.
Progress & compress: A scalable framework for continual learning
Jonathan Schwarz, Jelena Luketina, Wojciech Marian Czarnecki, Agnieszka Grabska-Barwińska, Yee Whye Teh, Razvan Pascanu, and Raia Hadsell · 2018
Later among the works it cites.
Learning symmetric and low-energy locomotion
Wenhao Yu, Greg Turk, and Chuanjian Liu · 2018
Later among the works it cites.
Optimality and approximation with policy gradient methods in markov decision processes
Anurag Agarwal, Sham M. Kakade, Jason D. Lee, and Gaurav Mahajan · 2019
Later among the works it cites.
Go-explore: a new approach for hard-exploration problems
Adrien Ecoffet, Joost Huizinga, Joel Lehman, Kenneth O. Stanley, and Jeff Clune · 2019
Later among the works it cites.
Keeping your distance: Solving sparse reward tasks using self-balancing shaped rewards
Alexander Trott, Stephan Zheng, Caiming Xiong, and Richard Socher · 2019
Later among the works it cites.
Rui Wang, Joel Lehman, Jeff Clune, and Kenneth O. Stanley · 2019
Later among the works it cites.
The ingredients of real-world robotic reinforcement learning
Henry Zhu, Justin Yu, Abhishek Gupta, Dhruv Shah, Kristian Hartikainen, Avi Singh, Vikash Kumar, and Sergey Levine · 2020
Closest in time.