Fetching the paper…
Reading the bibliography…
Common approaches for task-agnostic exploration learn tabula-rasa --the agent assumes isolated environments and no prior knowledge or experience.
Tackling climate change with machine learning
D. Rolnick, P. L. Donti, L. H. Kaack, K. Kochanski, A. Lacoste, K. Sankaran, A. S. Ross, N. Milojevic-Dupont, N. Jaques, A. Waldman-Brown, et al · 1906
Earlier work this paper cites.
Asymptotically efficient adaptive allocation rules
T. Lai and H. Robbins · 1985
Earlier work this paper cites.
A possibility for lmplementing curiosity and boredom in Model-Building neural controllers
J. Schmidhuber · 1991
Earlier work this paper cites.
Continual learning in reinforcement environments
M. B. Ring · 1994
Earlier work this paper cites.
Lifelong robot learning
S. Thrun and T. M. Mitchell · 1995
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
CHILD: A first step towards continual learning
M. B. Ring · 1998
Earlier work this paper cites.
On amount and quality of bias in reinforcement learning
G. Hailu and G. Sommer · 1999
Earlier work this paper cites.
Intrinsic and extrinsic motivations: Classic definitions and new directions
R. M. Ryan and E. L. Deci · 2000
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
P. Auer, N. Cesa-Bianchi, and P. Fischer · 2002
Earlier work this paper cites.
R-MAX - A general polynomial time algorithm for near-optimal reinforcement learning
R. I. Brafman and M. Tennenholtz · 2002
Earlier work this paper cites.
Similarity estimation techniques from rounding algorithms
M. S. Charikar · 2002
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
M. Kearns and S. Singh · 2002
Earlier work this paper cites.
All else being equal be empowered
A. S. Klyubin, D. Polani, and C. L. Nehaniv · 2005
Earlier work this paper cites.
Probabilistic policy reuse in a reinforcement learning agent
F. Fernández and M. Veloso · 2006
Earlier work this paper cites.
Developmental robotics, optimal artificial curiosity, creativity, music, and the fine arts
J. Schmidhuber · 2006
Earlier work this paper cites.
Logarithmic online regret bounds for undiscounted reinforcement learning
P. Auer and R. Ortner · 2007
Earlier work this paper cites.
An analysis of model-based interval estimation for Markov decision processes
A. L. Strehl and M. L. Littman · 2008
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
T. Jaksch, R. Ortner, and P. Auer · 2010
Earlier work this paper cites.
Information-seeking, curiosity, and attention: computational and neural mechanisms
J. Gottlieb, P. Oudeyer, M. Lopes, and A. Baranes · 2013
Earlier work this paper cites.
Fast and accurate deep network learning by exponential linear units (ELUs), 2015
D.-A. Clevert, T. Unterthiner, and S. Hochreiter · 2015
Cited alongside, same era.
Variational inference with normalizing flows
D. Rezende and S. Mohamed · 2015
Cited alongside, same era.
Policy distillation
A. A. Rusu, S. G. Colmenarejo, C. Gulcehre, G. Desjardins, J. Kirkpatrick, R. Pascanu, V. Mnih, K. Kavukcuoglu, and R. Hadsell · 2015
Cited alongside, same era.
Incentivizing exploration in reinforcement learning with deep predictive models
B. C. Stadie, S. Levine, and P. Abbeel · 2015
Cited alongside, same era.
Unifying count-based exploration and intrinsic motivation
M. G. Bellemare, S. Srinivasan, G. Ostrovski, T. Schaul, D. Saxton, and R. Munos · 2016
Cited alongside, same era.
VIME: Variational information maximizing exploration
R. Houthooft, X. Chen, Y. Duan, J. Schulman, F. De Turck, and P. Abbeel · 2016
Is Q-learning provably efficient?
C. Jin, Z. Allen-Zhu, S. Bubeck, and M. I. Jordan · 2018
Later among the works it cites.
Progress & compress: A scalable framework for continual learning
J. Schwarz, J. Luketina, W. M. Czarnecki, A. Grabska-Barwinska, Y. W. Teh, R. Pascanu, and R. Hadsell · 2018
Later among the works it cites.
Exploration by random network distillation
Y. Burda, H. Edwards, A. Storkey, and O. Klimov · 2019
Later among the works it cites.
Exploiting Action-Value uncertainty to drive exploration in reinforcement learning
C. D’Eramo, A. Cini, and M. Restelli · 2019
Later among the works it cites.
Guidelines for reinforcement learning in healthcare
O. Gottesman, F. Johansson, M. Komorowski, A. Faisal, D. Sontag, F. Doshi-Velez, and L. A. Celi · 2019
Later among the works it cites.
TorchBeast: A PyTorch Platform for Distributed RL, 2019
H. Küttler, N. Nardelli, T. Lavril, M. Selvatici, V. Sivakumar, T. Rocktäschel, and E. Grefenstette · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
A survey of transfer learning
K. Weiss, T. M. Khoshgoftaar, and D. Wang · 2016
Cited alongside, same era.
Successor features for transfer in reinforcement learning
A. Barreto, W. Dabney, R. Munos, J. J. Hunt, T. Schaul, H. Van Hasselt, and D. Silver · 2017
Cited alongside, same era.
What can machine learning do? workforce implications
E. Brynjolfsson and T. Mitchell · 2017
Cited alongside, same era.
Model-agnostic meta-learning for fast adaptation of deep networks
C. Finn, P. Abbeel, and S. Levine · 2017
Cited alongside, same era.
Overcoming catastrophic forgetting in neural networks
J. Kirkpatrick, R. Pascanu, N. Rabinowitz, J. Veness, G. Desjardins, A. A. Rusu, K. Milan, J. Quan, T. Ramalho, A. Grabska-Barwinska, D. Hassabis, C. Clopath, D. Kumaran, and R. Hadsell · 2017
Cited alongside, same era.
Count-based exploration with neural density models
G. Ostrovski, M. G. Bellemare, A. van den Oord, and R. Munos · 2017
Cited alongside, same era.
Later among the works it cites.
Deep exploration via randomized value functions
I. Osband, B. V. Roy, D. J. Russo, and Z. Wen · 2019
Later among the works it cites.
Efficient off-policy meta-reinforcement learning via probabilistic context variables
K. Rakelly, A. Zhou, C. Finn, S. Levine, and D. Quillen · 2019
Later among the works it cites.
Habitat: A Platform for Embodied AI Research
M. Savva, A. Kadian, O. Maksymets, Y. Zhao, E. Wijmans, B. Jain, J. Straub, J. Liu, V. Koltun, J. Malik, D. Parikh, and D. Batra · 2019
Later among the works it cites.
Receding horizon curiosity
M. Schultheis, B. Belousov, H. Abdulsamad, and J. Peters · 2019
Later among the works it cites.
The Replica dataset: A digital replica of indoor spaces, 2019
J. Straub, T. Whelan, L. Ma, Y. Chen, E. Wijmans, S. Green, J. J. Engel, R. Mur-Artal, C. Ren, S. Verma, A. Clarkson, M. Yan, B. Budge, Y. Yan, X. Pan, J. Yon, Y. Zou, K. Leon, N. Carter, J. Briales, T. Gillingham, E. Mueggler, L. Pesqueira, M. Savva, D. Batra, H. M. Strasdat, R. D. Nardi, M. Goesele, S. Lovegrove, and R. Newcombe · 2019
Later among the works it cites.
Neural topological slam for visual navigation
D. S. Chaplot, R. Salakhutdinov, A. Gupta, and S. Gupta · 2020
Later among the works it cites.
Q-learning with UCB exploration is sample efficient for Infinite-Horizon MDP
K. Dong, Y. Wang, X. Chen, and L. Wang · 2020
Later among the works it cites.
Assistive gym: A physics simulation framework for assistive robotics
Z. Erickson, V. Gangaram, A. Kapusta, C. K. Liu, and C. C. Kemp · 2020
Later among the works it cites.
Fast task inference with variational intrinsic successor features
S. Hansen, W. Dabney, A. Barreto, T. Van de Wiele, D. Warde-Farley, and V. Mnih · 2020
Later among the works it cites.
Curriculum learning for reinforcement learning domains: A framework and survey
S. Narvekar, B. Peng, M. Leonetti, J. Sinapov, M. E. Taylor, and P. Stone · 2020
Later among the works it cites.
RIDE: Rewarding Impact-Driven Exploration for Procedurally-Generated Environments
R. Raileanu and T. Rocktäschel · 2020
Later among the works it cites.
How should an agent practice?
J. Rajendran, R. Lewis, V. Veeriah, H. Lee, and S. Singh · 2020
Later among the works it cites.
Reinforcement learning with prototypical representations
D. Yarats, R. Fergus, A. Lazaric, and L. Pinto · 2021
Closest in time.