Fetching the paper…
Reading the bibliography…
We describe a simple scheme that allows an agent to learn about its environment in an unsupervised manner.
Some studies in machine learning using the game of checkers
Arthur L. Samuel · 1959
Earlier work this paper cites.
Curious model-building control systems
J. Schmidhuber · 1991
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J. Williams · 1992
Earlier work this paper cites.
Temporal difference learning and td-gammon
Gerald Tesauro · 1995
Earlier work this paper cites.
Intrinsically motivated reinforcement learning
Satinder P. Singh, Andrew G. Barto, and Nuttapong Chentanez · 2004
Earlier work this paper cites.
Empowerment: a universal agent-centric measure of control
Alexander S. Klyubin, Daniel Polani, and Chrystopher L. Nehaniv · 2005
Earlier work this paper cites.
An analysis of model-based interval estimation for markov decision processes
Alexander L. Strehl and Michael L. Littman · 2008
Earlier work this paper cites.
Curriculum learning
Yoshua Bengio, Jérôme Louradour, Ronan Collobert, and Jason Weston · 2009
Earlier work this paper cites.
Reinforcement learning for robot soccer
Martin Riedmiller, Thomas Gabel, Roland Hafner, and Sascha Lange · 2009
Earlier work this paper cites.
Self-paced learning for latent variable models
M. P. Kumar, Benjamin Packer, and Daphne Koller · 2010
Earlier work this paper cites.
Horde: A scalable real-time architecture for learning knowledge from unsupervised sensorimotor interaction
Richard S. Sutton, Joseph Modayil, Michael Delp, Thomas Degris, Patrick M. Pilarski, Adam White, and Doina Precup · 2011
Earlier work this paper cites.
Exploration in model-based reinforcement learning by empirically estimating learning progress
Manuel Lopes, Tobias Lang, Marc Toussaint, and Pierre-Yves Oudeyer · 2012
Cited alongside, same era.
Lecture 6.5 - rmsprop, coursera: Neural networks for machine learning, 2012
T. Tieleman and G. Hinton · 2012
Cited alongside, same era.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Cited alongside, same era.
Active learning of inverse models with intrinsically motivated goal exploration in robots
A. Baranes and P-Y. Oudeyer · 2013
Cited alongside, same era.
Intrinsic Motivation and Reinforcement Learning , pp. 17–47
Andrew G. Barto · 2013
Cited alongside, same era.
Generative adversarial nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Curiosity-driven exploration in deep reinforcement learning via bayesian neural networks
Rein Houthooft, Xi Chen, Yan Duan, John Schulman, Filip De Turck, and Pieter Abbeel · 2016
Later among the works it cites.
Mastering the game of Go with deep neural networks and tree search
David Silver, Aja Huang, Chris J. Maddison, Arthur Guez, Laurent Sifre, George van den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, Sander Dieleman, Dominik Grewe, John Nham, Nal Kalchbrenner, Ilya Sutskever, Timothy Lillicrap, Madeleine Leach, Koray Kavukcuoglu, Thore Graepel, and Demis Hassabis · 2016
Later among the works it cites.
Torchcraft: a library for machine learning research on real-time strategy games
Gabriel Synnaeve, Nantas Nardelli, Alex Auvolat, Soumith Chintala, Timothée Lacroix, Zeming Lin, Florian Richoux, and Nicolas Usunier · 2016
Later among the works it cites.
Hindsight experience replay
Marcin Andrychowicz, Dwight Crow, Alex K Ray, Jonas Schneider, Rachel Fong, Peter Welinder, Bob McGrew, Josh Tobin, Pieter Abbeel, and Wojciech Zaremba · 2017
Closest in time.
Reverse curriculum generation for reinforcement learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Universal value function approximators
Tom Schaul, Dan Horgan, Karol Gregor, and David Silver · 2015
Cited alongside, same era.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael I. Jordan, and Philipp Moritz · 2015
Cited alongside, same era.
Mazebase: A sandbox for learning from games
Sainbayar Sukhbaatar, Arthur Szlam, Gabriel Synnaeve, Soumith Chintala, and Rob Fergus · 2015
Cited alongside, same era.
Unifying count-based exploration and intrinsic motivation
Marc G. Bellemare, Sriram Srinivasan, Georg Ostrovski, Tom Schaul, David Saxton, and Rémi Munos · 2016
Cited alongside, same era.
Benchmarking deep reinforcement learning for continuous control
Yan Duan, Xi Chen, Rein Houthooft, John Schulman, and Pieter Abbeel · 2016
Cited alongside, same era.
Carlos Florensa, David Held, Markus Wulfmeier, Michael Zhang, and Pieter Abbeel · 2017
Closest in time.
Automatic goal generation for reinforcement learning agents
David Held, Xinyang Geng, Carlos Florensa, and Pieter Abbeel · 2017
Closest in time.
Adversarial learning for neural dialogue generation
Jiwei Li, Will Monroe, Tianlin Shi, Sébastien Jean, Alan Ritter, and Daniel Jurafsky · 2017
Closest in time.
Adversarial variational bayes: Unifying variational autoencoders and generative adversarial networks
Lars M. Mescheder, Sebastian Nowozin, and Andreas Geiger · 2017
Closest in time.
Curiosity-driven exploration by self-supervised prediction
Deepak Pathak, Pulkit Agrawal, Alexei A. Efros, and Trevor Darrell · 2017
Closest in time.
Robust adversarial reinforcement learning
Lerrel Pinto, James Davidson, Rahul Sukthankar, and Abhinav Gupta · 2017
Closest in time.
#exploration: A study of count-based exploration for deep reinforcement learning
Haoran Tang, Rein Houthooft, Davis Foote, Adam Stooke, Xi Chen, Yan Duan, John Schulman, Filip De Turck, and Pieter Abbeel · 2017
Closest in time.