Fetching the paper…
Reading the bibliography…
In Reinforcement Learning (RL), an agent explores the environment and collects trajectories into the memory buffer for later learning.
Pattern classification and scene analysis
Richard O Duda and Peter E Hart · 1973
Earlier work this paper cites.
Maximum likelihood from incomplete data via the em algorithm
Arthur P Dempster, Nan M Laird, and Donald B Rubin · 1977
Earlier work this paper cites.
Learning to generate focus trajectories for attentive vision
Jürgen Schmidhuber and Rudolf Huber · 1990
Earlier work this paper cites.
A possibility for implementing curiosity and boredom in model-building neural controllers
Jürgen Schmidhuber · 1991
Earlier work this paper cites.
Self-improving reactive agents based on reinforcement learning, planning and teaching
Long-Ji Lin · 1992
Earlier work this paper cites.
Acceleration of stochastic approximation by averaging
Boris T Polyak and Anatoli B Juditsky · 1992
Earlier work this paper cites.
Multitask learning
Rich Caruana · 1998
Earlier work this paper cites.
Reinforcement learning: An introduction , volume 1
Richard S Sutton and Andrew G Barto · 1998
Earlier work this paper cites.
Structure in the space of value functions
David Foster and Peter Dayan · 2002
Earlier work this paper cites.
Intrinsically motivated reinforcement learning
Nuttapong Chentanez, Andrew G Barto, and Satinder P Singh · 2005
Earlier work this paper cites.
Variational inference for dirichlet process mixtures
David M Blei, Michael I Jordan, et al · 2006
Earlier work this paper cites.
Autonomous inverted helicopter flight via reinforcement learning
Andrew Y Ng, Adam Coates, Mark Diel, Varun Ganapathi, Jamie Schulte, Ben Tse, Eric Berger, and Eric Liang · 2006
Earlier work this paper cites.
An analytic solution to discrete bayesian reinforcement learning
Pascal Poupart, Nikos Vlassis, Jesse Hoey, and Kevin Regan · 2006
Earlier work this paper cites.
To recognize shapes, first learn to generate images
Geoffrey E Hinton · 2007
Earlier work this paper cites.
Intrinsic motivation systems for autonomous mental development
Pierre-Yves Oudeyer, Frdric Kaplan, and Verena V Hafner · 2007
Earlier work this paper cites.
A discriminatively trained, multiscale, deformable part model
Pedro Felzenszwalb, David McAllester, and Deva Ramanan · 2008
Earlier work this paper cites.
Learning from imbalanced data
Haibo He and Edwardo A Garcia · 2008
Earlier work this paper cites.
Reinforcement learning of motor skills with policy gradients
Jan Peters and Stefan Schaal · 2008
Earlier work this paper cites.
Pearson correlation coefficient
Jacob Benesty, Jingdong Chen, Yiteng Huang, and Israel Cohen · 2009
Cited alongside, same era.
What is intrinsic motivation? a typology of computational approaches
Pierre-Yves Oudeyer and Frederic Kaplan · 2009
Cited alongside, same era.
Formal theory of creativity, fun, and intrinsic motivation (1990–2010)
Jürgen Schmidhuber · 2010
Cited alongside, same era.
Planning to be surprised: Optimal bayesian exploration in dynamic environments
Yi Sun, Faustino Gomez, and Jürgen Schmidhuber · 2011
Cited alongside, same era.
Horde: A scalable real-time architecture for learning knowledge from unsupervised sensorimotor interaction
Richard S Sutton, Joseph Modayil, Michael Delp, Thomas Degris, Patrick M Pilarski, Adam White, and Doina Precup · 2011
Cited alongside, same era.
Learning end-to-end goal-oriented dialog
Antoine Bordes, Y-Lan Boureau, and Jason Weston · 2016
Later among the works it cites.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Later among the works it cites.
Deep learning , volume 1
Ian Goodfellow, Yoshua Bengio, Aaron Courville, and Yoshua Bengio · 2016
Later among the works it cites.
Vime: Variational information maximizing exploration
Rein Houthooft, Xi Chen, Yan Duan, John Schulman, Filip De Turck, and Pieter Abbeel · 2016
Later among the works it cites.
End-to-end training of deep visuomotor policies
Sergey Levine, Chelsea Finn, Trevor Darrell, and Pieter Abbeel · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bruno Da Silva, George Konidaris, and Andrew Barto · 2012
Cited alongside, same era.
A review on ensembles for the class imbalance problem: bagging-, boosting-, and hybrid-based approaches
Mikel Galar, Alberto Fernandez, Edurne Barrenechea, Humberto Bustince, and Francisco Herrera · 2012
Cited alongside, same era.
Reinforcement learning to adjust parametrized motor primitives to new situations
Jens Kober, Andreas Wilhelm, Erhan Oztop, and Jan Peters · 2012
Cited alongside, same era.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Cited alongside, same era.
Exploration in model-based reinforcement learning by empirically estimating learning progress
Manuel Lopes, Tobias Lang, Marc Toussaint, and Pierre-Yves Oudeyer · 2012
Cited alongside, same era.
Machine learning: A probabilistic perspective. adaptive computation and machine learning, 2012
Kevin P Murphy · 2012
Cited alongside, same era.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Cited alongside, same era.
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2016
Later among the works it cites.
Prioritized experience replay
Tom Schaul, John Quan, Ioannis Antonoglou, and David Silver · 2016
Later among the works it cites.
Mastering the game of go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al · 2016
Later among the works it cites.
Hindsight experience replay
Marcin Andrychowicz, Filip Wolski, Alex Ray, Jonas Schneider, Rachel Fong, Peter Welinder, Bob McGrew, Josh Tobin, OpenAI Pieter Abbeel, and Wojciech Zaremba · 2017
Later among the works it cites.
Path integral guided policy search
Yevgen Chebotar, Mrinal Kalakrishnan, Ali Yahya, Adrian Li, Stefan Schaal, and Sergey Levine · 2017
Later among the works it cites.
Curiosity-driven exploration by self-supervised prediction
Deepak Pathak, Pulkit Agrawal, Alexei A Efros, and Trevor Darrell · 2017
Later among the works it cites.
Learning to push by grasping: Using multiple tasks for effective learning
Lerrel Pinto and Abhinav Gupta · 2017
Later among the works it cites.
Two-stream rnn/cnn for action recognition in 3d videos
Rui Zhao, Haider Ali, and Patrick Van der Smagt · 2017
Later among the works it cites.
A systematic study of the class imbalance problem in convolutional neural networks
Mateusz Buda, Atsuto Maki, and Maciej A Mazurowski · 2018
Later among the works it cites.
Multi-goal reinforcement learning: Challenging robotics environments and request for research
Matthias Plappert, Marcin Andrychowicz, Alex Ray, Bob McGrew, Bowen Baker, Glenn Powell, Jonas Schneider, Josh Tobin, Maciek Chociej, Peter Welinder, et al · 2018
Later among the works it cites.
Efficient dialog policy learning via positive memory retention
Rui Zhao and Volker Tresp · 2018
Later among the works it cites.
Learning goal-oriented visual dialog via tempered policy gradient
Rui Zhao and Volker Tresp · 2018
Later among the works it cites.
Maximum entropy-regularized multi-goal reinforcement learning
Rui Zhao, Xudong Sun, and Volker Tresp · 2019
Closest in time.