Fetching the paper…
Reading the bibliography…
In Hindsight Experience Replay (HER), a reinforcement learning agent is trained by treating whatever it has achieved as virtual goals.
Learning to generate focus trajectories for attentive vision
J. Schmidhuber and R. Huber · 1990
Earlier work this paper cites.
Acceleration of stochastic approximation by averaging
B. T. Polyak and A. B. Juditsky · 1992
Earlier work this paper cites.
Self-improving reactive agents based on reinforcement learning, planning and teaching
L.-J. Lin · 1992
Earlier work this paper cites.
Learning and development in neural networks: The importance of starting small
J. L. Elman · 1993
Earlier work this paper cites.
Reinforcement learning: An introduction , volume 1
R. S. Sutton and A. G. Barto · 1998
Earlier work this paper cites.
Multitask learning
R. Caruana · 1998
Earlier work this paper cites.
Structure in the space of value functions
D. Foster and P. Dayan · 2002
Earlier work this paper cites.
Optimal ordered problem solver
J. Schmidhuber · 2004
Earlier work this paper cites.
Autonomous inverted helicopter flight via reinforcement learning
A. Y. Ng, A. Coates, M. Diel, V. Ganapathi, J. Schulte, B. Tse, E. Berger, and E. Liang · 2006
Earlier work this paper cites.
Sears and Zemansky’s university physics , volume 1
H. D. Young, R. A. Freedman, and A. L. Ford · 2006
Earlier work this paper cites.
Physics for scientists and engineers
P. A. Tipler and G. Mosca · 2007
Earlier work this paper cites.
Reinforcement learning of motor skills with policy gradients
J. Peters and S. Schaal · 2008
Earlier work this paper cites.
Pearson correlation coefficient
J. Benesty, J. Chen, Y. Huang, and I. Cohen · 2009
Earlier work this paper cites.
Curriculum learning
Y. Bengio, J. Louradour, R. Collobert, and J. Weston · 2009
Earlier work this paper cites.
A tutorial on se (3) transformation parameterizations and on-manifold optimization
J.-L. Blanco · 2010
Earlier work this paper cites.
Horde: A scalable real-time architecture for learning knowledge from unsupervised sensorimotor interaction
R. S. Sutton, J. Modayil, M. Delp, T. Degris, P. M. Pilarski, A. White, and D. Precup · 2011
Cited alongside, same era.
Mujoco: A physics engine for model-based control
E. Todorov, T. Erez, and Y. Tassa · 2012
Cited alongside, same era.
Engineering mechanics: dynamics , volume 2
J. L. Meriam and L. G. Kraige · 2012
Cited alongside, same era.
B. Da Silva, G. Konidaris, and A. Barto · 2012
Cited alongside, same era.
Reinforcement learning to adjust parametrized motor primitives to new situations
J. Kober, A. Wilhelm, E. Oztop, and J. Peters · 2012
Cited alongside, same era.
Powerplay: Training an increasingly general problem solver by continually searching for the simplest still unsolvable problem
G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba · 2016
Later among the works it cites.
Continuous control with deep reinforcement learning
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra · 2016
Later among the works it cites.
Prioritized experience replay
T. Schaul, J. Quan, I. Antonoglou, and D. Silver · 2016
Later among the works it cites.
Two-stream rnn/cnn for action recognition in 3d videos
R. Zhao, H. Ali, and P. Van der Smagt · 2017
Later among the works it cites.
Path integral guided policy search
Y. Chebotar, M. Kalakrishnan, A. Yahya, A. Li, S. Schaal, and S. Levine · 2017
Later among the works it cites.
Hindsight experience replay
M. Andrychowicz, F. Wolski, A. Ray, J. Schneider, R. Fong, P. Welinder, B. McGrew, J. Tobin, O. P. Abbeel, and W. Zaremba · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Schmidhuber · 2013
Cited alongside, same era.
First experiments with powerplay
R. K. Srivastava, B. R. Steunebrink, and J. Schmidhuber · 2013
Cited alongside, same era.
Multi-task policy search for robotics
M. P. Deisenroth, P. Englert, J. Peters, and D. Fox · 2014
Cited alongside, same era.
W. Zaremba and I. Sutskever · 2014
Cited alongside, same era.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al · 2015
Cited alongside, same era.
Universal value function approximators
T. Schaul, D. Horgan, K. Gregor, and D. Silver · 2015
Cited alongside, same era.
Deep learning , volume 1
I. Goodfellow, Y. Bengio, A. Courville, and Y. Bengio · 2016
Cited alongside, same era.
Later among the works it cites.
Target-driven visual navigation in indoor scenes using deep reinforcement learning
Y. Zhu, R. Mottaghi, E. Kolve, J. J. Lim, A. Gupta, L. Fei-Fei, and A. Farhadi · 2017
Later among the works it cites.
Automatic goal generation for reinforcement learning agents
D. Held, X. Geng, C. Florensa, and P. Abbeel · 2017
Later among the works it cites.
Learning to push by grasping: Using multiple tasks for effective learning
L. Pinto and A. Gupta · 2017
Later among the works it cites.
Automated curriculum learning for neural networks
A. Graves, M. G. Bellemare, J. Menick, R. Munos, and K. Kavukcuoglu · 2017
Later among the works it cites.
Intrinsic motivation and automatic curricula via asymmetric self-play
S. Sukhbaatar, Z. Lin, I. Kostrikov, G. Synnaeve, A. Szlam, and R. Fergus · 2017
Later among the works it cites.
Reverse curriculum generation for reinforcement learning
C. Florensa, D. Held, M. Wulfmeier, and P. Abbeel · 2017
Later among the works it cites.
Learning goal-oriented visual dialog via tempered policy gradient
R. Zhao and V. Tresp · 2018
Closest in time.
Efficient dialog policy learning via positive memory retention
R. Zhao and V. Tresp · 2018
Closest in time.
Multi-goal reinforcement learning: Challenging robotics environments and request for research
M. Plappert, M. Andrychowicz, A. Ray, B. McGrew, B. Baker, G. Powell, J. Schneider, J. Tobin, M. Chociej, P. Welinder, et al · 2018
Closest in time.