Fetching the paper…
Reading the bibliography…
Despite the close connection between exploration and sample efficiency, most state of the art reinforcement learning algorithms include no considerations for exploration beyond maximizing the entropy of the policy.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
William R Thompson · 1933
Earlier work this paper cites.
Prioritized sweeping: Reinforcement learning with less data and less time
Andrew W. Moore and Christopher G. Atkeson · 1993
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Michael Kearns and Satinder P. Singh · 1998
Earlier work this paper cites.
R-max - a general polynomial time algorithm for near-optimal reinforcement learning
R. Brafman and Moshe Tennenholtz · 2002
Earlier work this paper cites.
Reinforcement learning on explicitly specified time scales
Ralf Schoknecht and Martin A. Riedmiller · 2003
Earlier work this paper cites.
Pac model-free reinforcement learning
Alexander L. Strehl, L. Li, Eric Wiewiora, J. Langford, and M. Littman · 2006
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
T. Jaksch, R. Ortner, and P. Auer · 2008
Earlier work this paper cites.
An analysis of model-based interval estimation for markov decision processes
Alexander L Strehl and Michael L Littman · 2008
Earlier work this paper cites.
Near-bayesian exploration in polynomial time
J. Z. Kolter and A. Ng · 2009
Earlier work this paper cites.
Normal reference bandwidths for the general order, multivariate kernel density derivative estimator
D. Henderson and Christopher F. Parmeter · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents (extended abstract)
Marc G. Bellemare, Yavar Naddaf, J. Veness, and Michael Bowling · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Earlier work this paper cites.
Incentivizing exploration in reinforcement learning with deep predictive models
Bradly C. Stadie, Sergey Levine, and Pieter Abbeel · 2015
Earlier work this paper cites.
Pareto smoothed importance sampling
A. Vehtari, A. Gelman, and J. Gabry · 2015
Earlier work this paper cites.
Unifying count-based exploration and intrinsic motivation
Marc Bellemare, Sriram Srinivasan, Georg Ostrovski, Tom Schaul, David Saxton, and Remi Munos · 2016
Cited alongside, same era.
Deep reinforcement learning with double q-learning
H. V. Hasselt, A. Guez, and D. Silver · 2016
Cited alongside, same era.
Vime: Variational information maximizing exploration
Rein Houthooft, Xi Chen, Yan Duan, John Schulman, Filip De Turck, and Pieter Abbeel · 2016
Cited alongside, same era.
Deep exploration via bootstrapped dqn
Ian Osband, Charles Blundell, Alexander Pritzel, and Benjamin Van Roy · 2016
Cited alongside, same era.
Tom Schaul, John Quan, Ioannis Antonoglou, and D. Silver · 2016
Cited alongside, same era.
Count-based exploration with neural density models
Georg Ostrovski, Marc G Bellemare, Aäron van den Oord, and Rémi Munos · 2017
Parameter space noise for exploration
Matthias Plappert, Rein Houthooft, Prafulla Dhariwal, S. Sidor, Richard Y. Chen, Xi Chen, T. Asfour, P. Abbeel, and Marcin Andrychowicz · 2018
Later among the works it cites.
Learning by playing - solving sparse reward tasks from scratch
Martin A. Riedmiller, Roland Hafner, T. Lampe, Michael Neunert, J. Degrave, T. Wiele, V. Mnih, N. Heess, and Jost Tobias Springenberg · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Later among the works it cites.
Yuval Tassa, Yotam Doron, Alistair Muldal, Tom Erez, Yazhe Li, Diego de Las Casas, David Budden, Abbas Abdolmaleki, Josh Merel, Andrew Lefrancq, et al · 2018
Later among the works it cites.
Deep exploration via randomized value functions
Ian Osband, Benjamin Van Roy, Daniel J. Russo, and Zheng Wen · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Curiosity-driven exploration by self-supervised prediction
Deepak Pathak, Pulkit Agrawal, Alexei A Efros, and Trevor Darrell · 2017
Cited alongside, same era.
Data-efficient deep reinforcement learning for dexterous manipulation
I. Popov, N. Heess, T. Lillicrap, Roland Hafner, Gabriel Barth-Maron, Matej Vecerík, T. Lampe, Y. Tassa, T. Erez, and Martin A. Riedmiller · 2017
Cited alongside, same era.
#exploration: A study of count-based exploration for deep reinforcement learning
Haoran Tang, Rein Houthooft, Davis Foote, Adam Stooke, OpenAI Xi Chen, Yan Duan, John Schulman, Filip DeTurck, and Pieter Abbeel · 2017
Cited alongside, same era.
Maximum a posteriori policy optimisation
Abbas Abdolmaleki, Jost Tobias Springenberg, Yuval Tassa, Rémi Munos, Nicolas Manfred Otto Heess, and Martin A. Riedmiller · 2018
Cited alongside, same era.
JAX: composable transformations of Python+NumPy programs, 2018
James Bradbury, Roy Frostig, Peter Hawkins, Matthew James Johnson, Chris Leary, Dougal Maclaurin, George Necula, Adam Paszke, Jake VanderPlas, Skye Wanderman-Milne, and Qiao Zhang · 2018
Cited alongside, same era.
Exploration by random network distillation
Yuri Burda, Harrison Edwards, Amos Storkey, and Oleg Klimov · 2018
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, S. Gross, Francisco Massa, A. Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Z. Lin, N. Gimelshein, L. Antiga, Alban Desmaison, Andreas Köpf, Edward Yang, Zach DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala · 2019
Later among the works it cites.
Never give up: Learning directed exploration strategies
Adrià Puigdomènech Badia, P. Sprechmann, Alex Vitvitskyi, Daniel Guo, B. Piot, Steven Kapturowski, O. Tieleman, Martín Arjovsky, A. Pritzel, Andew Bolt, and Charles Blundell · 2020
Later among the works it cites.
Temporally-extended ϵ \epsilon -greedy exploration
Will Dabney, Georg Ostrovski, and A. Barreto · 2020
Later among the works it cites.
Flax: A neural network library and ecosystem for JAX, 2020
Jonathan Heek, Anselm Levskaya, Avital Oliver, Marvin Ritter, Bertrand Rondepierre, Andreas Steiner, and Marc van Zee · 2020
Later among the works it cites.
Count-based exploration with the successor representation
Marlos C. Machado, Marc G. Bellemare, and Michael H. Bowling · 2020
Later among the works it cites.
Continuous-discrete reinforcement learning for hybrid control in robotics, 2020
Michael Neunert, Abbas Abdolmaleki, Markus Wulfmeier, Thomas Lampe, Jost Tobias Springenberg, Roland Hafner, Francesco Romano, Jonas Buchli, Nicolas Heess, and Martin Riedmiller · 2020
Later among the works it cites.
Optimistic exploration even with a pessimistic initialisation
Tabish Rashid, B. Peng, Wendelin Böhmer, and S. Whiteson · 2020
Later among the works it cites.
On bonus based exploration methods in the arcade learning environment
Adrien Ali Taiga, William Fedus, Marlos C. Machado, Aaron Courville, and Marc G. Bellemare · 2020
Later among the works it cites.
Dynamics-aware embeddings
William F. Whitney, Rajat Agarwal, Kyunghyun Cho, and Abhinav Gupta · 2020
Later among the works it cites.
Soft actor-critic (sac) implementation in pytorch
Denis Yarats and Ilya Kostrikov · 2020
Later among the works it cites.
Kernel operations on the gpu, with autodiff, without memory overflows
Benjamin Charlier, Jean Feydy, Joan Alexis Glaunès, François-David Collin, and Ghislain Durif · 2021
Closest in time.