Fetching the paper…
Reading the bibliography…
Exploration remains a central challenge for reinforcement learning (RL).
Guided search: an alternative to the feature integration model for visual search
J. M. Wolfe, K. R. Cave, and S. L. Franzel · 1989
Earlier work this paper cites.
Curious model-building control systems
Jürgen Schmidhuber · 1991
Earlier work this paper cites.
Efficient exploration in reinforcement learning, 1992
Sebastian B Thrun · 1992
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Play and exploration in children and animals
Thomas G Power · 1999
Earlier work this paper cites.
Between MDPs and semi-MDPs: A framework for temporal abstraction in reinforcement learning
Richard S Sutton, Doina Precup, and Satinder Singh · 1999
Earlier work this paper cites.
A cortical network sensitive to stimulus salience in a neutral behavioral context across multiple sensory modalities
Jonathan Downar, Adrian P Crawley, David J Mikulis, and Karen D Davis · 2002
Earlier work this paper cites.
Homeostatic plasticity in the developing nervous system
Gina G Turrigiano and Sacha B Nelson · 2004
Earlier work this paper cites.
Empowerment: A universal agent-centric measure of control
Alexander S Klyubin, Daniel Polani, and Chrystopher L Nehaniv · 2005
Earlier work this paper cites.
Should I stay or should I go? How the human brain manages the trade-off between exploitation and exploration
Jonathan D Cohen, Samuel M McClure, and Angela J Yu · 2007
Earlier work this paper cites.
Experiments in socially guided exploration: Lessons learned in building robots that learn with and without human teachers
Andrea L Thomaz and Cynthia Breazeal · 2008
Earlier work this paper cites.
Ensemble algorithms in reinforcement learning
Marco A Wiering and Hado Van Hasselt · 2008
Earlier work this paper cites.
What is intrinsic motivation? a typology of computational approaches
Pierre-Yves Oudeyer and Frederic Kaplan · 2009
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Thomas Jaksch, Ronald Ortner, and Peter Auer · 2010
Earlier work this paper cites.
Formal theory of creativity, fun, and intrinsic motivation (1990–2010)
Jürgen Schmidhuber · 2010
Earlier work this paper cites.
Adaptive ε \varepsilon -greedy exploration in reinforcement learning based on value differences
Michel Tokic · 2010
Earlier work this paper cites.
Build order optimization in starcraft
David Churchill and Michael Buro · 2011
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Marc G Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling · 2013
Earlier work this paper cites.
Kullback–leibler upper confidence bounds for optimal sequential allocation
Olivier Cappé, Aurélien Garivier, Odalric-Ambrym Maillard, Rémi Munos, Gilles Stoltz, et al · 2013
Earlier work this paper cites.
Rank the episodes: A simple approach for exploration in procedurally-generated environments
Daochen Zha, Wenye Ma, Lei Yuan, Xia Hu, and Ji Liu · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Exploration versus exploitation in space, mind, and society
T. T. Hills, P. M. Todd, D. Lazer, A. D. Redish, and I. D. Couzin · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A. Rusu, Joel Veness, Marc G. Bellemare, Alex Graves, Martin Riedmiller, Andreas K. Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis · 2015
Cited alongside, same era.
Unifying count-based exploration and intrinsic motivation
Marc G Bellemare, Sriram Srinivasan, Georg Ostrovski, Tom Schaul, David Saxton, and Rémi Munos · 2016
Cited alongside, same era.
Karol Gregor, Danilo Jimenez Rezende, and Daan Wierstra · 2016
Cited alongside, same era.
Vime: Variational information maximizing exploration
Rein Houthooft, Xi Chen, Yan Duan, John Schulman, Filip De Turck, and Pieter Abbeel · 2016
Cited alongside, same era.
Recurrent experience replay in distributed reinforcement learning
Steven Kapturowski, Georg Ostrovski, Will Dabney, John Quan, and Remi Munos · 2019
Later among the works it cites.
Bumblebees learn foraging routes through exploitation-exploration cycles
J. M. Kembro, M. Lihoreau, J. Garriga, E. P. Raposo, and F. Bartumeus · 2019
Later among the works it cites.
Adapting behaviour via intrinsic reward: A survey and empirical study, 2019
Cam Linke, Nadia M. Ady, Martha White, Thomas Degris, and Adam White · 2019
Later among the works it cites.
Adapting behaviour for learning progress, 2019
Tom Schaul, Diana Borsa, David Ding, David Szepesvari, Georg Ostrovski, Will Dabney, and Simon Osindero · 2019
Later among the works it cites.
Structured, uncertainty-driven exploration in real-world consumer choice
Eric Schulz, Rahul Bhui, Bradley C. Love, Bastien Brier, Michael T. Todd, and Samuel J. Gershman · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Max Jaderberg, Volodymyr Mnih, Wojciech Marian Czarnecki, Tom Schaul, Joel Z Leibo, David Silver, and Koray Kavukcuoglu · 2016
Cited alongside, same era.
Adaptive skills adaptive partitions (ASAP)
Daniel J Mankowitz, Timothy Arthur Mann, and Shie Mannor · 2016
Cited alongside, same era.
Prioritized experience replay
Tom Schaul, John Quan, Ioannis Antonoglou, and David Silver · 2016
Cited alongside, same era.
Dueling network architectures for deep reinforcement learning
Ziyu Wang, Tom Schaul, Matteo Hessel, Hado Hasselt, Marc Lanctot, and Nando Freitas · 2016
Cited alongside, same era.
A Primer on Foraging and the Explore/Exploit Trade-Off for Psychiatry Research
M. A. Addicott, J. M. Pearson, M. M. Sweitzer, D. L. Barack, and M. L. Platt · 2017
Cited alongside, same era.
Count-based exploration with neural density models
Georg Ostrovski, Marc G. Bellemare, Aäron van den Oord, and Rémi Munos · 2017
Cited alongside, same era.
JAX: composable transformations of Python+NumPy programs, 2018
James Bradbury, Roy Frostig, Peter Hawkins, Matthew James Johnson, Chris Leary, Dougal Maclaurin, George Necula, Adam Paszke, Jake VanderPlas, Skye Wanderman-Milne, and Qiao Zhang · 2018
Cited alongside, same era.
Sample-efficient reinforcement learning with stochastic ensemble value expansion
Jacob Buckman, Danijar Hafner, George Tucker, Eugene Brevdo, and Honglak Lee · 2018
Cited alongside, same era.
Louis Bagot, Kevin Mets, and Steven Latré · 2020
Later among the works it cites.
Reverb: An efficient data storage and transport system for ml research, 2020
Albin Cassirer, Gabriel Barth-Maron, Thibault Sottiaux, Manuel Kroiss, and Eugene Brevdo · 2020
Later among the works it cites.
Dopaminergic modulation of the exploration/exploitation trade-off in human decision-making
K. Chakroun, D. Mathar, A. Wiehler, F. Ganzer, and J. Peters · 2020
Later among the works it cites.
Temporally-extended ϵ \epsilon -greedy exploration, 2020
Will Dabney, Georg Ostrovski, and André Barreto · 2020
Later among the works it cites.
A bayesian approach to robust reinforcement learning
Esther Derman, Daniel Mankowitz, Timothy Mann, and Shie Mannor · 2020
Later among the works it cites.
Temporal difference uncertainties as a signal for exploration
Sebastian Flennerhag, Jane X Wang, Pablo Sprechmann, Francesco Visin, Alexandre Galashov, Steven Kapturowski, Diana L Borsa, Nicolas Heess, Andre Barreto, and Razvan Pascanu · 2020
Later among the works it cites.
Haiku: Sonnet for JAX, 2020
Tom Hennigan, Trevor Cai, Tamara Norman, and Igor Babuschkin · 2020
Later among the works it cites.
Optax: Composable gradient transformation and optimisation, in JAX!, 2020
Matteo Hessel, David Budden, Fabio Viola, Mihaela Rosca, Eren Sezener, and Tom Hennigan · 2020
Later among the works it cites.
Bandit algorithms
Tor Lattimore and Csaba Szepesvári · 2020
Later among the works it cites.
Differential effects of psychotic illness on directed and random exploration
James A. Waltz, Robert C. Wilson, Matthew A. Albrecht, Michael J. Frank, and James M. Gold · 2020
Later among the works it cites.
Fast and slow curiosity for high-level exploration in reinforcement learning
Nicolas Bougie and Ryutaro Ichise · 2021
Closest in time.
Coverage as a principle for discovering transferable behavior in reinforcement learning
Víctor Campos, Pablo Sprechmann, Steven Hansen, Andre Barreto, Steven Kapturowski, Alex Vitvitskyi, Adrià Puigdomènech Badia, and Charles Blundell · 2021
Closest in time.
Increased random exploration in schizophrenia is associated with inflammation
Flurin Cathomas, Federica Klaus, Karoline Guetter, Hui-Kuan Chung, Anjali Raja Beharelle, Tobias R. Spiller, Rebecca Schlegel, Erich Seifritz, Matthias N. Hartmann-Riemer, Philippe N. Tobler, and Stefan Kaiser · 2021
Closest in time.
Go-explore: a new approach for hard-exploration problems, 2021
Adrien Ecoffet, Joost Huizinga, Joel Lehman, Kenneth O. Stanley, and Jeff Clune · 2021
Closest in time.
Deep reinforcement learning with dynamic optimism
Ted Moskovitz, Jack Parker-Holder, Aldo Pacchiano, and Michael Arbel · 2021
Closest in time.
Return-based scaling: Yet another normalisation trick for deep RL
Tom Schaul, Georg Ostrovski, Iurii Kemaev, and Diana Borsa · 2021
Closest in time.