Fetching the paper…
Reading the bibliography…
Reinforcement learning with sparse rewards is challenging because an agent can rarely obtain non-zero rewards and hence, gradient-based optimization of parameterized policies can be incremental and slow.
Adaptive confidence and adaptive curiosity
J. Schmidhuber · 1991
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
R. S. Sutton, D. A. McAllester, S. P. Singh, and Y. Mansour · 2000
Earlier work this paper cites.
Using confidence bounds for exploitation-exploration trade-offs
P. Auer · 2002
Earlier work this paper cites.
Never give up: Learning directed exploration strategies
A. P. Badia, P. Sprechmann, A. Vitvitskyi, D. Guo, B. Piot, S. Kapturowski, O. Tieleman, M. Arjovsky, A. Pritzel, A. Bolt, et al · 2002
Earlier work this paper cites.
Agent57: Outperforming the atari human benchmark
A. P. Badia, B. Piot, S. Kapturowski, P. Sprechmann, A. Vitvitskyi, D. Guo, and C. Blundell · 2003
Earlier work this paper cites.
Intrinsically motivated reinforcement learning
N. Chentanez, A. G. Barto, and S. P. Singh · 2005
Earlier work this paper cites.
An analysis of model-based interval estimation for markov decision processes
A. L. Strehl and M. L. Littman · 2008
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
M. G. Bellemare, Y. Naddaf, J. Veness, and M. Bowling · 2013
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
D. Bahdanau, K. Cho, and Y. Bengio · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
I. Sutskever, O. Vinyals, and Q. V. Le · 2014
Earlier work this paper cites.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, S. Petersen, C. Beattie, A. Sadik, I. Antonoglou, H. King, D. Kumaran, D. Wierstra, S. Legg, and D. Hassabis · 2015
Earlier work this paper cites.
Unifying count-based exploration and intrinsic motivation
M. Bellemare, S. Srinivasan, G. Ostrovski, T. Schaul, D. Saxton, and R. Munos · 2016
Earlier work this paper cites.
Benchmarking deep reinforcement learning for continuous control
Y. Duan, X. Chen, R. Houthooft, J. Schulman, and P. Abbeel · 2016
Earlier work this paper cites.
K. Gregor, D. J. Rezende, and D. Wierstra · 2016
Earlier work this paper cites.
Learning to navigate in complex environments
P. Mirowski, R. Pascanu, F. Viola, H. Soyer, A. J. Ballard, A. Banino, M. Denil, R. Goroshin, L. Sifre, K. Kavukcuoglu, et al · 2016
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, et al · 2016
Earlier work this paper cites.
Hindsight experience replay
M. Andrychowicz, F. Wolski, A. Ray, J. Schneider, R. Fong, P. Welinder, B. McGrew, J. Tobin, O. P. Abbeel, and W. Zaremba · 2017
Earlier work this paper cites.
Variational memory addressing in generative models
J. Bornschein, A. Mnih, D. Zoran, and D. J. Rezende · 2017
Earlier work this paper cites.
Learning modular neural network policies for multi-task and multi-robot transfer
C. Devin, A. Gupta, T. Darrell, P. Abbeel, and S. Levine · 2017
Cited alongside, same era.
Revisiting the arcade learning environment: Evaluation protocols and open problems for general agents
M. C. Machado, M. G. Bellemare, E. Talvitie, J. Veness, M. Hausknecht, and M. Bowling · 2017
Cited alongside, same era.
Combining self-supervised learning and imitation for vision-based rope manipulation
A. Nair, D. Chen, P. Agrawal, P. Isola, P. Abbeel, J. Malik, and S. Levine · 2017
Cited alongside, same era.
Zero-shot task generalization with multi-task deep reinforcement learning
J. Oh, S. Singh, H. Lee, and P. Kohli · 2017
Cited alongside, same era.
Count-based exploration with neural density models
G. Ostrovski, M. G. Bellemare, A. van den Oord, and R. Munos · 2017
Cited alongside, same era.
Curiosity-driven exploration by self-supervised prediction
Recurrent experience replay in distributed reinforcement learning
S. Kapturowski, G. Ostrovski, J. Quan, R. Munos, and W. Dabney · 2018
Later among the works it cites.
Memory augmented policy optimization for program synthesis and semantic parsing
C. Liang, M. Norouzi, J. Berant, Q. V. Le, and N. Lao · 2018
Later among the works it cites.
Episodic memory deep q-networks
Z. Lin, T. Zhao, G. Yang, and L. Zhang · 2018
Later among the works it cites.
Data-efficient hierarchical reinforcement learning
O. Nachum, S. S. Gu, H. Lee, and S. Levine · 2018
Later among the works it cites.
J. Oh, Y. Guo, S. Singh, and H. Lee · 2018
Later among the works it cites.
Zero-shot visual imitation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
D. Pathak, P. Agrawal, A. A. Efros, and T. Darrell · 2017
Cited alongside, same era.
Neural episodic control
A. Pritzel, B. Uria, S. Srinivasan, A. P. Badia, O. Vinyals, D. Hassabis, D. Wierstra, and C. Blundell · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Cited alongside, same era.
# exploration: A study of count-based exploration for deep reinforcement learning
H. Tang, R. Houthooft, D. Foote, A. Stooke, O. X. Chen, Y. Duan, J. Schulman, F. DeTurck, and P. Abbeel · 2017
Cited alongside, same era.
Playing hard exploration games by watching youtube
Y. Aytar, T. Pfaff, D. Budden, T. Paine, Z. Wang, and N. de Freitas · 2018
Cited alongside, same era.
Contingency-aware exploration in reinforcement learning
J. Choi, Y. Guo, M. Moczulski, J. Oh, N. Wu, M. Norouzi, and H. Lee · 2018
Cited alongside, same era.
Diversity is all you need: Learning skills without a reward function
B. Eysenbach, A. Gupta, J. Ibarz, and S. Levine · 2018
Cited alongside, same era.
D. Pathak, P. Mahmoudieh, G. Luo, P. Agrawal, D. Chen, Y. Shentu, E. Shelhamer, J. Malik, A. A. Efros, and T. Darrell · 2018
Later among the works it cites.
Observe and look further: Achieving consistent performance on atari
T. Pohlen, B. Piot, T. Hester, M. G. Azar, D. Horgan, D. Budden, G. Barth-Maron, H. van Hasselt, J. Quan, M. Večerík, et al · 2018
Later among the works it cites.
Unsupervised control through non-parametric discriminative rewards
D. Warde-Farley, T. Van de Wiele, T. Kulkarni, C. Ionescu, S. Hansen, and V. Mnih · 2018
Later among the works it cites.
Gibson env: Real-world perception for embodied agents
F. Xia, A. R. Zamir, Z. He, A. Sax, J. Malik, and S. Savarese · 2018
Later among the works it cites.
Go-explore: a new approach for hard-exploration problems
A. Ecoffet, J. Huizinga, J. Lehman, K. O. Stanley, and J. Clune · 2019
Closest in time.
Model-based reinforcement learning for atari
L. Kaiser, M. Babaeizadeh, P. Milos, B. Osinski, R. H. Campbell, K. Czechowski, D. Erhan, C. Finn, P. Kozakowski, S. Levine, et al · 2019
Closest in time.
Unsupervised learning of object keypoints for perception and control
T. D. Kulkarni, A. Gupta, C. Ionescu, S. Borgeaud, M. Reynolds, A. Zisserman, and V. Mnih · 2019
Closest in time.
Learning abstract models for long-horizon exploration, 2019
E. Z. Liu, R. Keramati, S. Seshadri, K. Guu, P. Pasupat, E. Brunskill, and P. Liang · 2019
Closest in time.
Behaviour suite for reinforcement learning
I. Osband, Y. Doron, M. Hessel, J. Aslanides, E. Sezener, A. Saraiva, K. McKinney, T. Lattimore, C. Szepesvári, S. Singh, B. Van Roy, R. Sutton, D. Silver, and H. van Hasselt · 2019
Closest in time.
Skew-fit: State-covering self-supervised reinforcement learning
V. H. Pong, M. Dalal, S. Lin, A. Nair, S. Bahl, and S. Levine · 2019
Closest in time.
Mastering atari, go, chess and shogi by planning with a learned model
J. Schrittwieser, I. Antonoglou, T. Hubert, K. Simonyan, L. Sifre, S. Schmitt, A. Guez, E. Lockhart, D. Hassabis, T. Graepel, et al · 2019
Closest in time.
Learning novel policies for tasks
Y. Zhang, W. Yu, and G. Turk · 2019
Closest in time.