Fetching the paper…
Reading the bibliography…
Reinforcement learning methods require careful design involving a reward function to obtain the desired action policy for a given task.
On the Theory of the Brownian Motion
G. E. Uhlenbeck and L. S. Ornstein. 1930 · 1930
Earlier work this paper cites.
Efficient Training of Artificial Neural Networks for Autonomous Navigation
D. A. Pomerleau. 1991 · 1991
Earlier work this paper cites.
Long Short-Term Memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Learning from Demonstration
Stefan Schaal. 1997 · 1997
Earlier work this paper cites.
Introduction to Reinforcement Learning
Richard S. Sutton and Andrew G. Barto. 1998 · 1998
Earlier work this paper cites.
Algorithms for Inverse Reinforcement Learning. In International Conference on Machine Learning
Andrew Y. Ng and Stuart J. Russell. 2000 · 2000
Earlier work this paper cites.
Reinforcement learning with long short-term memory. In Advances in neural information processing systems
Bram Bakker. 2002 · 2002
Earlier work this paper cites.
Potential-based Shaping and Q-value Initialization Are Equivalent
Eric Wiewiora. 2003 · 2003
Earlier work this paper cites.
Apprenticeship Learning via Inverse Reinforcement Learning. In International Conference on Machine Learning
Pieter Abbeel and Andrew Y. Ng. 2004 · 2004
Earlier work this paper cites.
Maximum Entropy Inverse Reinforcement Learning. In AAAI
Brian D. Ziebart, Andrew Maas, J. Andrew (Drew) Bagnell, and Anind Dey. 2008 · 2008
Earlier work this paper cites.
Rectified Linear Units Improve Restricted Boltzmann Machines. In International Conference on Machine Learning
Vinod Nair and Geoffrey E. Hinton. 2010 · 2010
Earlier work this paper cites.
3D convolutional neural networks for human action recognition
Shuiwang Ji, Wei Xu, Ming Yang, and Kai Yu. 2013 · 2013
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization
Diederik P. Kingma and Jimmy Ba. 2014 · 2014
Cited alongside, same era.
Reinforcement Learning from Demonstration Through Shaping. In International Conference on Artificial Intelligence
Tim Brys, Anna Harutyunyan, Halit Bener Suay, Sonia Chernova, Matthew E. Taylor, and Ann Nowé. 2015 · 2015
Cited alongside, same era.
Distributed recurrent neural forward models with synaptic adaptation and CPG-based control for complex behaviors of walking robots
Sakyasingha Dasgupta, Dennis Goldschmidt, Florentin Wörgötter, and Poramate Manoonpong. 2015 · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A. Rusu, Joel Veness, Marc G. Bellemare, Alex Graves, Martin Riedmiller, Andreas K. Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis. 2015 · 2015
Cited alongside, same era.
Trust Region Policy Optimization. In International Conference on Machine Learning
End-to-End Differentiable Adversarial Imitation Learning. In International Conference on Machine Learning
Nir Baram, Oron Anschel, Itai Caspi, and Shie Mannor. 2017 · 2017
Later among the works it cites.
Yan Duan, Marcin Andrychowicz, Bradly Stadie, Jonathan Ho, Jonas Schneider, Ilya Sutskever, Pieter Abbeel, and Wojciech Zaremba. 2017 · 2017
Later among the works it cites.
Keras-FlappyBird
Ben Lau. 2017 · 2017
Later among the works it cites.
Roboschool
OpenAI. 2017 · 2017
Later among the works it cites.
gym-super-mario
Philip Paquette. 2017 · 2017
Later among the works it cites.
noreward-rl
Deepak Pathak. 2017 · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz. 2015 · 2015
Cited alongside, same era.
Maximum entropy deep inverse reinforcement learning
Markus Wulfmeier, Peter Ondruska, and Ingmar Posner. 2015 · 2015
Cited alongside, same era.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba. 2016 · 2016
Cited alongside, same era.
Generative adversarial imitation learning. In Advances in Neural Information Processing Systems
Jonathan Ho and Stefano Ermon. 2016 · 2016
Cited alongside, same era.
Timothy Lillicrap, Jonathan Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra. 2016 · 2016
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning. In International Conference on Machine Learning
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu. 2016 · 2016
Cited alongside, same era.
keras-rl
Matthias Plappert. 2016 · 2016
Cited alongside, same era.
Learning from Demonstration for Shaping Through Inverse Reinforcement Learning. In Proceedings of the 2016 International Conference on Autonomous Agents & Multiagent Systems
Halit Bener Suay, Tim Brys, Matthew E. Taylor, and Sonia Chernova. 2016 · 2016
Cited alongside, same era.
Curiosity-driven Exploration by Self-supervised Prediction. In International Conference on Machine Learning (ICML)
Deepak Pathak, Pulkit Agrawal, Alexei A. Efros, and Trevor Darrell. 2017 · 2017
Later among the works it cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. 2017 · 2017
Later among the works it cites.
Robust Imitation of Diverse Behaviors
Ziyu Wang, Josh Merel, Scott E. Reed, Greg Wayne, Nando de Freitas, and Nicolas Heess. 2017 · 2017
Later among the works it cites.
DAQN: Deep Auto-encoder and Q-Network
Daiki Kimura. 2018 · 2018
Closest in time.
Behavioral Cloning from Observation. In International Joint Conference on Artificial Intelligence
Faraz Torabi, Garrett Warnell, and Peter Stone. 2018 · 2018
Closest in time.