Fetching the paper…
Reading the bibliography…
Learning with sparse rewards remains a significant challenge in reinforcement learning (RL), especially when the aim is to train a policy capable of achieving multiple different goals.
Calculus of Variations and Optimal Control Theory: A Concise Introduction
D. Liberzon · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control, 2012
E. Todorov, T. Erez, and Y. Tassa · 2012
Earlier work this paper cites.
Playing atari with deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. Graves, I. Antonoglou, D. Wierstra, and M. Riedmiller · 2013
Earlier work this paper cites.
Generative adversarial networks, 2014
I. J. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio · 2014
Earlier work this paper cites.
Continuous control with deep reinforcement learning, 2015
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, S. Petersen, C. Beattie, A. Sadik, I. Antonoglou, H. King, D. Kumaran, D. Wierstra, S. Legg, and D. Hassabis · 2015
Earlier work this paper cites.
Unifying count-based exploration and intrinsic motivation
M. Bellemare, S. Srinivasan, G. Ostrovski, T. Schaul, D. Saxton, and R. Munos · 2016
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. van den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, S. Dieleman, D. Grewe, J. Nham, N. Kalchbrenner, I. Sutskever, T. Lillicrap, M. Leach, K. Kavukcuoglu, T. Graepel, and D. Hassabis · 2016
Earlier work this paper cites.
Hindsight experience replay
M. Andrychowicz, F. Wolski, A. Ray, J. Schneider, R. Fong, P. Welinder, B. McGrew, J. Tobin, O. Pieter Abbeel, and W. Zaremba · 2017
Earlier work this paper cites.
M. Arjovsky, S. Chintala, and L. Bottou · 2017
Earlier work this paper cites.
Openai baselines
P. Dhariwal, C. Hesse, O. Klimov, A. Nichol, M. Plappert, A. Radford, J. Schulman, S. Sidor, Y. Wu, and P. Zhokhov · 2017
Earlier work this paper cites.
Automatic goal generation for reinforcement learning agents, 2017
C. Florensa, D. Held, X. Geng, and P. Abbeel · 2017
Earlier work this paper cites.
Curiosity-driven exploration by self-supervised prediction
D. Pathak, P. Agrawal, A. A. Efros, and T. Darrell · 2017
Earlier work this paper cites.
Imagination-augmented agents for deep reinforcement learning
S. Racanière, T. Weber, D. P. Reichert, L. Buesing, A. Guez, D. Rezende, A. P. Badia, O. Vinyals, N. Heess, Y. Li, R. Pascanu, P. Battaglia, D. Hassabis, D. Silver, and D. Wierstra · 2017
Earlier work this paper cites.
Maximum a posteriori policy optimisation
A. Abdolmaleki, J. T. Springenberg, Y. Tassa, R. Munos, N. Heess, and M. Riedmiller · 2018
Cited alongside, same era.
Sample-efficient deep RL with generative adversarial tree search
K. Azizzadenesheli, B. Yang, W. Liu, E. Brunskill, Z. C. Lipton, and A. Anandkumar · 2018
Cited alongside, same era.
Distributional policy gradients
G. Barth-Maron, M. W. Hoffman, D. Budden, W. Dabney, D. Horgan, D. TB, A. Muldal, N. Heess, and T. Lillicrap · 2018
Cited alongside, same era.
Sample-efficient reinforcement learning with stochastic ensemble value expansion
J. Buckman, D. Hafner, G. Tucker, E. Brevdo, and H. Lee · 2018
Cited alongside, same era.
Model-ensemble trust-region policy optimization
T. Kurutach, I. Clavera, Y. Duan, A. Tamar, and P. Abbeel · 2018
Cited alongside, same era.
Spectral normalization for generative adversarial networks
Exploration by random network distillation
Y. Burda, H. Edwards, A. Storkey, and O. Klimov · 2019
Later among the works it cites.
Curriculum-guided hindsight experience replay
M. Fang, T. Zhou, Y. Du, L. Han, and Z. Zhang · 2019
Later among the works it cites.
Learning to reach goals without reinforcement learning
D. Ghosh, A. Gupta, J. Fu, A. Reddy, C. Devin, B. Eysenbach, and S. Levine · 2019
Later among the works it cites.
Learning latent dynamics for planning from pixels
D. Hafner, T. Lillicrap, I. Fischer, R. Villegas, D. Ha, H. Lee, and J. Davidson · 2019
Later among the works it cites.
Model-based reinforcement learning for atari, 2019
L. Kaiser, M. Babaeizadeh, P. Milos, B. Osinski, R. H. Campbell, K. Czechowski, D. Erhan, C. Finn, P. Kozakowski, S. Levine, A. Mohiuddin, R. Sepassi, G. Tucker, and H. Michalewski · 2019
Later among the works it cites.
Competitive experience replay
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
T. Miyato, T. Kataoka, M. Koyama, and Y. Yoshida · 2018
Cited alongside, same era.
Neural network dynamics for model-based deep reinforcement learning with model-free fine-tuning
A. Nagabandi, G. Kahn, R. Fearing, and S. Levine · 2018
Cited alongside, same era.
Visual reinforcement learning with imagined goals, 2018
A. Nair, V. Pong, M. Dalal, S. Bahl, S. Lin, and S. Levine · 2018
Cited alongside, same era.
Temporal difference models: Model-free deep RL for model-based control
V. Pong*, S. Gu*, M. Dalal, and S. Levine · 2018
Cited alongside, same era.
Learning Complex Dexterous Manipulation with Deep Reinforcement Learning and Demonstrations
A. Rajeswaran*, V. Kumar*, A. Gupta, G. Vezzani, J. Schulman, E. Todorov, and S. Levine · 2018
Cited alongside, same era.
A general reinforcement learning algorithm that masters chess, shogi, and go through self-play
D. Silver, T. Hubert, J. Schrittwieser, I. Antonoglou, M. Lai, A. Guez, M. Lanctot, L. Sifre, D. Kumaran, T. Graepel, T. Lillicrap, K. Simonyan, and D. Hassabis · 2018
Cited alongside, same era.
Robust sampling based model predictive control with sparse objective information
G. Williams, B. Goldfain, P. Drews, K. Saigol, J. Rehg, and E. Theodorou · 2018
Cited alongside, same era.
H. Liu, A. Trott, R. Socher, and C. Xiong · 2019
Later among the works it cites.
Deep Dynamics Models for Learning Dexterous Manipulation
A. Nagabandi, K. Konoglie, S. Levine, and V. Kumar · 2019
Later among the works it cites.
Planning with goal-conditioned policies
S. Nasiriany, V. Pong, S. Lin, and S. Levine · 2019
Later among the works it cites.
Mastering atari, go, chess and shogi by planning with a learned model, 2019
J. Schrittwieser, I. Antonoglou, T. Hubert, K. Simonyan, L. Sifre, S. Schmitt, A. Guez, E. Lockhart, D. Hassabis, T. Graepel, T. Lillicrap, and D. Silver · 2019
Later among the works it cites.
Agent57: Outperforming the atari human benchmark, 2020
A. P. Badia, B. Piot, S. Kapturowski, P. Sprechmann, A. Vitvitskyi, D. Guo, and C. Blundell · 2020
Closest in time.
Dream to control: Learning behaviors by latent imagination
D. Hafner, T. Lillicrap, J. Ba, and M. Norouzi · 2020
Closest in time.
Soft hindsight experience replay, 2020
Q. He, L. Zhuang, and H. Li · 2020
Closest in time.
Curl: Contrastive unsupervised representations for reinforcement learning, 2020
A. Srinivas, M. Laskin, and P. Abbeel · 2020
Closest in time.