Fetching the paper…
Reading the bibliography…
Exploration is essential in reinforcement learning, particularly in environments where external rewards are sparse.
Judgment of contingency between responses and outcomes
Herbert M Jenkins and William C Ward · 1965
Earlier work this paper cites.
Dynamic programming
Richard Bellman · 1966
Earlier work this paper cites.
Integrated modeling and control based on reinforcement learning and dynamic programming
Richard S Sutton · 1990
Earlier work this paper cites.
Curious model-building control systems
Jürgen Schmidhuber · 1991
Earlier work this paper cites.
Improving generalization for temporal difference learning: The successor representation
Peter Dayan · 1993
Earlier work this paper cites.
Temporal difference learning and td-gammon
Gerald Tesauro et al · 1995
Earlier work this paper cites.
Exploration bonuses and dual control
Peter Dayan and Terrence J Sejnowski · 1996
Earlier work this paper cites.
Explorations in efficient reinforcement learning
Marco Alexander Wiering · 1999
Earlier work this paper cites.
Logarithmic online regret bounds for undiscounted reinforcement learning
Peter Auer and Ronald Ortner · 2006
Earlier work this paper cites.
Decision-theoretic planning with non-markovian rewards
Sylvie Thiébaux, Charles Gretton, John Slaney, David Price, and Froduald Kabanza · 2006
Earlier work this paper cites.
An analysis of model-based interval estimation for markov decision processes
Alexander L Strehl and Michael L Littman · 2008
Earlier work this paper cites.
Structure learning in human sequential decision-making
Daniel Acuna and Paul R Schrater · 2008
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Marc G Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling · 2013
Earlier work this paper cites.
Temporal contingency
Charles R Gallistel, Andrew R Craig, and Timothy A Shahan · 2014
Cited alongside, same era.
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Cited alongside, same era.
Action-conditional video prediction using deep networks in atari games
Junhyuk Oh, Xiaoxiao Guo, Honglak Lee, Richard L Lewis, and Satinder Singh · 2015
Cited alongside, same era.
Deep exploration via bootstrapped dqn
Ian Osband, Charles Blundell, Alexander Pritzel, and Benjamin Van Roy · 2016
Cited alongside, same era.
Karol Gregor, Danilo Jimenez Rezende, and Daan Wierstra · 2016
Cited alongside, same era.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Later among the works it cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Later among the works it cites.
The successor representation: its computational logic and neural substrates
Samuel J Gershman · 2018
Later among the works it cites.
Exploration by random network distillation
Yuri Burda, Harrison Edwards, Amos Storkey, and Oleg Klimov · 2018
Later among the works it cites.
Revisiting the arcade learning environment: Evaluation protocols and open problems for general agents
Marlos C Machado, Marc G Bellemare, Erik Talvitie, Joel Veness, Matthew Hausknecht, and Michael Bowling · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Marc Bellemare, Sriram Srinivasan, Georg Ostrovski, Tom Schaul, David Saxton, and Remi Munos · 2016
Cited alongside, same era.
Curiosity-driven exploration by self-supervised prediction
Deepak Pathak, Pulkit Agrawal, Alexei A Efros, and Trevor Darrell · 2017
Cited alongside, same era.
A unified view of entropy-regularized markov decision processes
Gergely Neu, Anders Jonsson, and Vicenç Gómez · 2017
Cited alongside, same era.
Noisy networks for exploration
Meire Fortunato, Mohammad Gheshlaghi Azar, Bilal Piot, Jacob Menick, Ian Osband, Alex Graves, Vlad Mnih, Remi Munos, Demis Hassabis, Olivier Pietquin, et al · 2017
Cited alongside, same era.
Model-agnostic meta-learning for fast adaptation of deep networks
Chelsea Finn, Pieter Abbeel, and Sergey Levine · 2017
Cited alongside, same era.
Meta learning shared hierarchies
Kevin Frans, Jonathan Ho, Xi Chen, Pieter Abbeel, and John Schulman · 2017
Cited alongside, same era.
Ex2: Exploration with exemplar models for deep reinforcement learning
Justin Fu, John Co-Reyes, and Sergey Levine · 2017
Cited alongside, same era.
Will Dabney, Georg Ostrovski, and André Barreto · 2020
Later among the works it cites.
Count-based exploration with the successor representation
Marlos C Machado, Marc G Bellemare, and Michael Bowling · 2020
Later among the works it cites.
A survey of exploration methods in reinforcement learning
Susan Amin, Maziar Gomrokchi, Harsh Satija, Herke van Hoof, and Doina Precup · 2021
Later among the works it cites.
A first-occupancy representation for reinforcement learning
Ted Moskovitz, Spencer R Wilson, and Maneesh Sahani · 2021
Later among the works it cites.
The learning of prospective and retrospective cognitive maps within neural circuits
Vijay Mohan K Namboodiri and Garret D Stuber · 2021
Later among the works it cites.
Cic: Contrastive intrinsic control for unsupervised skill discovery
Michael Laskin, Hao Liu, Xue Bin Peng, Denis Yarats, Aravind Rajeswaran, and Pieter Abbeel · 2022
Later among the works it cites.
Duncan Bailey and Marcelo Mattar · 2022
Later among the works it cites.
Intrinsically motivated exploration as empowerment
Franziska Brändle, Lena J Stocks, Joshua Tenenbaum, Samuel J Gershman, and Eric Schulz · 2022
Later among the works it cites.