Fetching the paper…
Reading the bibliography…
Sparse rewards are double-edged training signals in reinforcement learning: easy to design but hard to optimize.
Structure and direction in thinking
D. E. Berlyne · 1965
Earlier work this paper cites.
Learning control of finite markov chains with an explicit trade-off between estimation and control
M. Sato, K. Abe, and H. Takeda · 1988
Earlier work this paper cites.
Discovering the structure of a reactive environment by exploration
M. C. Mozer and J. Bachrach · 1990
Earlier work this paper cites.
On the computational economics of reinforcement learning
A. G. Barto and S. P. Singh · 1991
Earlier work this paper cites.
A possibility for implementing curiosity and boredom in model-building neural controllers
J. Schmidhuber · 1991
Earlier work this paper cites.
Complexity and cooperation in q-learning
S. D. Whitehead · 1991
Earlier work this paper cites.
Efficient exploration in reinforcement learning
S. B. Thrun · 1992
Earlier work this paper cites.
Reinforcement learning: A survey
L. P. Kaelbling, M. L. Littman, and A. W. Moore · 1996
Earlier work this paper cites.
Learning agents for uncertain environments
S. Russell · 1998
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
A. Y. Ng, D. Harada, and S. Russell · 1999
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
M. Kearns and S. Singh · 2002
Earlier work this paper cites.
On the sample complexity of reinforcement learning
S. M. Kakade et al · 2003
Earlier work this paper cites.
Intrinsically motivated learning of hierarchical collections of skills
A. G. Barto, S. Singh, and N. Chentanez · 2004
Earlier work this paper cites.
Intrinsically motivated reinforcement learning
N. Chentanez, A. G. Barto, and S. P. Singh · 2005
Earlier work this paper cites.
An intrinsic reward mechanism for efficient exploration
Ö. Şimşek and A. G. Barto · 2006
Earlier work this paper cites.
Pac model-free reinforcement learning
A. L. Strehl, L. Li, E. Wiewiora, J. Langford, and M. L. Littman · 2006
Earlier work this paper cites.
Intrinsic motivation systems for autonomous mental development
P.-Y. Oudeyer, F. Kaplan, and V. V. Hafner · 2007
Earlier work this paper cites.
How can we define intrinsic motivation
P.-Y. Oudeyer, F. Kaplan, et al · 2008
Earlier work this paper cites.
An analysis of model-based interval estimation for markov decision processes
A. L. Strehl and M. L. Littman · 2008
Earlier work this paper cites.
Near-bayesian exploration in polynomial time
J. Z. Kolter and A. Y. Ng · 2009
Earlier work this paper cites.
Kalman temporal differences
M. Geist and O. Pietquin · 2010
Earlier work this paper cites.
Formal theory of creativity, fun, and intrinsic motivation (1990–2010)
J. Schmidhuber · 2010
Earlier work this paper cites.
Exploration in model-based reinforcement learning by empirically estimating learning progress
M. Lopes, T. Lang, M. Toussaint, and P.-Y. Oudeyer · 2012
Earlier work this paper cites.
Deep reinforcement learning in large discrete action spaces
G. Dulac-Arnold, R. Evans, H. van Hasselt, P. Sunehag, Lillicrap, et al · 2015
Earlier work this paper cites.
Continuous control with deep reinforcement learning
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra · 2015
Earlier work this paper cites.
Variational information maximisation for intrinsically motivated reinforcement learning
S. Mohamed and D. Jimenez Rezende · 2015
Cited alongside, same era.
Universal value function approximators
T. Schaul, D. Horgan, K. Gregor, and D. Silver · 2015
Cited alongside, same era.
Incentivizing exploration in reinforcement learning with deep predictive models
B. C. Stadie, S. Levine, and P. Abbeel · 2015
Cited alongside, same era.
Concrete problems in ai safety
D. Amodei, C. Olah, J. Steinhardt, P. Christiano, J. Schulman, and D. Mané · 2016
Cited alongside, same era.
Unifying count-based exploration and intrinsic motivation
M. Bellemare, S. Srinivasan, G. Ostrovski, T. Schaul, D. Saxton, and R. Munos · 2016
Cited alongside, same era.
Variational intrinsic control
K. Gregor, D. J. Rezende, and D. Wierstra · 2016
Computational theories of curiosity-driven learning
P.-Y. Oudeyer · 2018
Later among the works it cites.
The uncertainty bellman equation and exploration
B. O’Donoghue, I. Osband, R. Munos, and V. Mnih · 2018
Later among the works it cites.
Parameter space noise for exploration
M. Plappert, R. Houthooft, P. Dhariwal, S. Sidor, R. Y. Chen, X. Chen, T. Asfour, et al · 2018
Later among the works it cites.
Episodic curiosity through reachability
N. Savinov, A. Raichuk, D. Vincent, R. Marinier, M. Pollefeys, T. Lillicrap, and S. Gelly · 2018
Later among the works it cites.
Intrinsic motivation and automatic curricula via asymmetric self-play
S. Sukhbaatar, Z. Lin, I. Kostrikov, G. Synnaeve, A. Szlam, and R. Fergus · 2018
Later among the works it cites.
Reinforcement learning: An introduction
R. S. Sutton and A. G. Barto · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Vime: Variational information maximizing exploration
R. Houthooft, X. Chen, Y. Duan, J. Schulman, F. De Turck, and P. Abbeel · 2016
Cited alongside, same era.
Hierarchical deep reinforcement learning: Integrating temporal abstraction and intrinsic motivation
T. D. Kulkarni, K. Narasimhan, A. Saeedi, and J. Tenenbaum · 2016
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu · 2016
Cited alongside, same era.
Deep exploration via bootstrapped dqn
I. Osband, C. Blundell, A. Pritzel, and B. Van Roy · 2016
Cited alongside, same era.
Surprise-based intrinsic motivation for deep reinforcement learning
J. Achiam and S. Sastry · 2017
Cited alongside, same era.
Hindsight experience replay
M. Andrychowicz, F. Wolski, A. Ray, J. Schneider, R. Fong, P. Welinder, B. McGrew, J. Tobin, O. P. Abbeel, and W. Zaremba · 2017
Cited alongside, same era.
Later among the works it cites.
M. G. Azar, B. Piot, B. A. Pires, J.-B. Grill, F. Altché, and R. Munos · 2019
Later among the works it cites.
Never give up: Learning directed exploration strategies
A. P. Badia, P. Sprechmann, A. Vitvitskyi, D. Guo, B. Piot, S. Kapturowski, O. Tieleman, M. Arjovsky, A. Pritzel, A. Bolt, et al · 2019
Later among the works it cites.
Mulex: Disentangling exploitation from exploration in deep rl
L. Beyer, D. Vincent, O. Teboul, S. Gelly, M. Geist, and O. Pietquin · 2019
Later among the works it cites.
Large-scale study of curiosity-driven learning
Y. Burda, H. Edwards, D. Pathak, A. Storkey, T. Darrell, and A. A. Efros · 2019
Later among the works it cites.
Feature control as intrinsic motivation for hierarchical reinforcement learning
N. Dilokthanakul, C. Kaplanis, N. Pawlowski, and M. Shanahan · 2019
Later among the works it cites.
Go-explore: a new approach for hard-exploration problems
A. Ecoffet, J. Huizinga, J. Lehman, K. O. Stanley, and J. Clune · 2019
Later among the works it cites.
TorchBeast: A PyTorch Platform for Distributed RL
H. Küttler, N. Nardelli, T. Lavril, et al · 2019
Later among the works it cites.
Ride: Rewarding impact-driven exploration for procedurally-generated environments
R. Raileanu and T. Rocktäschel · 2019
Later among the works it cites.
The Bitter Lesson
R. Sutton · 2019
Later among the works it cites.
Scheduled intrinsic drive: A hierarchical take on intrinsically motivated exploration
J. Zhang, N. Wetzel, N. Dorka, J. Boedecker, and W. Burgard · 2019
Later among the works it cites.
Agent57: Outperforming the atari human benchmark
A. P. Badia, B. Piot, S. Kapturowski, et al · 2020
Later among the works it cites.
Learning with amigo: Adversarially motivated intrinsic goals
A. Campero, R. Raileanu, H. Küttler, et al · 2020
Later among the works it cites.
Higher: Improving instruction following with hindsight generation for experience replay
G. Cideron, M. Seurin, F. Strub, and O. Pietquin · 2020
Later among the works it cites.
Intrinsically motivated goal-conditioned reinforcement learning: a short survey
C. Colas, T. Karch, O. Sigaud, and P.-Y. Oudeyer · 2020
Later among the works it cites.
Show me the way: Intrinsic motivation from demonstrations
L. Hussenot, R. Dadashi, M. Geist, and O. Pietquin · 2020
Later among the works it cites.
Discretizing continuous action space for on-policy optimization
Y. Tang and S. Agrawal · 2020
Later among the works it cites.
Bebold: Exploration beyond the boundary of explored regions
T. Zhang, H. Xu, X. Wang, Y. Wu, K. Keutzer, J. E. Gonzalez, and Y. Tian · 2020
Later among the works it cites.
Adversarially guided actor-critic
Y. Flet-Berliac, J. Ferret, O. Pietquin, P. Preux, and M. Geist · 2021
Closest in time.