Fetching the paper…
Reading the bibliography…
While recent progress in deep reinforcement learning has enabled robots to learn complex behaviors, tasks with long horizons and sparse rewards remain an ongoing challenge.
Markov decision processes: Discrete stochastic dynamic programming
M. L. Puterman · 1994
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
A. Y. Ng, D. Harada, and S. Russell · 1999
Earlier work this paper cites.
Dynamic potential-based reward shaping
S. M. Devlin and D. Kudenko · 2012
Earlier work this paper cites.
Elements of information theory
T. M. Cover and J. A. Thomas · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
E. Todorov, T. Erez, and Y. Tassa · 2012
Earlier work this paper cites.
Representation learning: A review and new perspectives
Y. Bengio, A. Courville, and P. Vincent · 2013
Earlier work this paper cites.
Auto-encoding variational bayes
D. P. Kingma and M. Welling · 2013
Earlier work this paper cites.
On the properties of neural machine translation: Encoder–decoder approaches
K. Cho, B. van Merriënboer, D. Bahdanau, and Y. Bengio · 2014
Earlier work this paper cites.
Potential based reward shaping for hierarchical reinforcement learning
Y. Gao and F. Toni · 2015
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu · 2016
Earlier work this paper cites.
Openai gym, 2016
G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba · 2016
Earlier work this paper cites.
Hindsight experience replay
M. Andrychowicz, F. Wolski, A. Ray, J. Schneider, R. Fong, P. Welinder, B. McGrew, J. Tobin, O. P. Abbeel, and W. Zaremba · 2017
Cited alongside, same era.
Deep reinforcement learning with successor features for navigation across similar environments
J. Zhang, J. T. Springenberg, J. Boedecker, and W. Burgard · 2017
Cited alongside, same era.
Neural discrete representation learning
A. van den Oord, O. Vinyals, et al · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Cited alongside, same era.
Openai baselines
P. Dhariwal, C. Hesse, O. Klimov, A. Nichol, M. Plappert, A. Radford, J. Schulman, S. Sidor, Y. Wu, and P. Zhokhov · 2017
Cited alongside, same era.
Representation learning with contrastive predictive coding
A. v. d. Oord, Y. Li, and O. Vinyals · 2018
Later among the works it cites.
Impala: Scalable distributed deep-rl with importance weighted actor-learner architectures
L. Espeholt, H. Soyer, R. Munos, K. Simonyan, V. Mnih, T. Ward, Y. Doron, V. Firoiu, T. Harley, I. Dunning, et al · 2018
Later among the works it cites.
Continual state representation learning for reinforcement learning using generative replay
H. Caselles-Dupré, M. Garcia-Ortiz, and D. Filliat · 2018
Later among the works it cites.
Minimalistic gridworld environment for openai gym
M. Chevalier-Boisvert, L. Willems, and S. Pal · 2018
Later among the works it cites.
Data-efficient hierarchical reinforcement learning
O. Nachum, S. S. Gu, H. Lee, and S. Levine · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M. Jaderberg, W. M. Czarnecki, I. Dunning, L. Marris, G. Lever, A. G. Castaneda, C. Beattie, N. C. Rabinowitz, A. S. Morcos, A. Ruderman, et al · 2018
Cited alongside, same era.
Learning to walk via deep reinforcement learning
T. Haarnoja, A. Zhou, S. Ha, J. Tan, G. Tucker, and S. Levine · 2018
Cited alongside, same era.
Learning by playing solving sparse reward tasks from scratch
M. Riedmiller, R. Hafner, T. Lampe, M. Neunert, J. Degrave, T. Wiele, V. Mnih, N. Heess, and J. T. Springenberg · 2018
Cited alongside, same era.
Near-optimal representation learning for hierarchical reinforcement learning
O. Nachum, S. Gu, H. Lee, and S. Levine · 2018
Cited alongside, same era.
Learning actionable representations with goal-conditioned policies
D. Ghosh, A. Gupta, and S. Levine · 2018
Cited alongside, same era.
Learning deep representations by mutual information estimation and maximization
R. D. Hjelm, A. Fedorov, S. Lavoie-Marchildon, K. Grewal, P. Bachman, A. Trischler, and Y. Bengio · 2018
Cited alongside, same era.
Mutual information neural estimation
M. I. Belghazi, A. Baratin, S. Rajeshwar, S. Ozair, Y. Bengio, D. Hjelm, and A. Courville · 2018
Cited alongside, same era.
Later among the works it cites.
Montezuma’s revenge solved by go-explore, a new algorithm for hard-exploration problems (sets records on pitfall, too)
A. Ecoffet, J. Huizinga, J. Lehman, K. O. Stanley, and J. Clune · 2018
Later among the works it cites.
Learning montezuma’s revenge from a single demonstration
T. Salimans and R. Chen · 2018
Later among the works it cites.
Stable baselines
A. Hill, A. Raffin, M. Ernestus, A. Gleave, R. Traore, P. Dhariwal, C. Hesse, O. Klimov, A. Nichol, M. Plappert, A. Radford, J. Schulman, S. Sidor, and Y. Wu · 2018
Later among the works it cites.
Archer: Aggressive rewards to counter bias in hindsight experience replay
S. Lanka and T. Wu · 2018
Later among the works it cites.
OpenAI Five, 2019
OpenAI · 2019
Closest in time.
Learning to generalize from sparse and underspecified rewards
R. Agarwal, C. Liang, D. Schuurmans, and M. Norouzi · 2019
Closest in time.
Reward shaping via meta-learning
H. Zou, T. Ren, D. Yan, H. Su, and J. Zhu · 2019
Closest in time.