Fetching the paper…
Reading the bibliography…
Exploration is essential for solving complex Reinforcement Learning (RL) tasks.
GAIT: a geometric approach to information theory
Jose Gallego-Posada, Ankit Vani, Max Schwarzer, and Simon Lacoste-Julien · 1906
Earlier work this paper cites.
Dynamic programming and modern control theory
Richard Bellman and Robert Kalaba · 1965
Earlier work this paper cites.
Possible generalization of Boltzmann-Gibbs statistics
Constantino Tsallis · 1988
Earlier work this paper cites.
A possibility for implementing curiosity and boredom in model-building neural controllers
Jürgen Schmidhuber · 1991
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
Martin Puterman · 1994
Earlier work this paper cites.
Dynamic programming and optimal control , volume 1
Dimitri Bertsekas · 1995
Earlier work this paper cites.
Nonparametric entropy estimation: An overview
Jan Beirlant, Edward J Dudewicz, László Györfi, and Edward C Van der Meulen · 1997
Earlier work this paper cites.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 1998
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S. Sutton, David A. McAllester, Satinder P. Singh, and Yishay Mansour · 1999
Earlier work this paper cites.
R-max-a general polynomial time algorithm for near-optimal reinforcement learning
Ronen I Brafman and Moshe Tennenholtz · 2002
Earlier work this paper cites.
A Sparse Sampling Algorithm for Near-Optimal Planning in Large Markov Decision Processes
Michael Kearns, Yishay Mansour, and Andrew Ng · 2002
Earlier work this paper cites.
Slow feature analysis: Unsupervised learning of invariances
Laurenz Wiskott and Terrence J Sejnowski · 2002
Earlier work this paper cites.
The linear programming approach to approximate dynamic programming
D.P. de Farias and B. Van Roy · 2003
Earlier work this paper cites.
On the sample complexity of reinforcement learning
Sham Machandranath Kakade et al · 2003
Earlier work this paper cites.
Prox-method with rate of convergence o (1/t) for variational inequalities with lipschitz continuous monotone operators and smooth convex-concave saddle point problems
Arkadi Nemirovski · 2004
Earlier work this paper cites.
Pac model-free reinforcement learning
Alexander L Strehl, Lihong Li, Eric Wiewiora, John Langford, and Michael L Littman · 2006
Earlier work this paper cites.
Noise-contrastive estimation: A new estimation principle for unnormalized statistical models
Michael Gutmann and Aapo Hyvärinen · 2010
Earlier work this paper cites.
Estimating divergence functionals and the likelihood ratio by convex risk minimization
XuanLong Nguyen, Martin J Wainwright, and Michael I Jordan · 2010
Cited alongside, same era.
Autonomous exploration for navigating in MDPs
Shiau Hong Lim and Peter Auer · 2012
Cited alongside, same era.
Near-optimal pac bounds for discounted mdps
Tor Lattimore and Marcus Hutter · 2014
Cited alongside, same era.
Autonomous learning of state representations for control: An emerging field aims to autonomously learn state representations for reinforcement learning agents from their real-world sensor observations
Wendelin Böhmer, Jost Tobias Springenberg, Joschka Boedecker, Martin Riedmiller, and Klaus Obermayer · 2015
Cited alongside, same era.
Sample complexity of episodic fixed-horizon reinforcement learning
Christoph Dann and Emma Brunskill · 2015
Cited alongside, same era.
Impala: Scalable distributed deep-rl with importance weighted actor-learner architectures
Lasse Espeholt, Hubert Soyer, Remi Munos, Karen Simonyan, Volodymir Mnih, Tom Ward, Yotam Doron, Vlad Firoiu, Tim Harley, Iain Dunning, et al · 2018
Later among the works it cites.
Provably efficient maximum entropy exploration
Elad Hazan, Sham M. Kakade, Karan Singh, and Abby Van Soest · 2018
Later among the works it cites.
State representation learning for control: An overview
Timothée Lesort, Natalia Díaz-Rodríguez, Jean-Franois Goudou, and David Filliat · 2018
Later among the works it cites.
Randomized prior functions for deep reinforcement learning
Ian Osband, John Aslanides, and Albin Cassirer · 2018
Later among the works it cites.
Episodic curiosity through reachability
Nikolay Savinov, Anton Raichuk, Raphaël Marinier, Damien Vincent, Marc Pollefeys, Timothy Lillicrap, and Sylvain Gelly · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Cited alongside, same era.
Unifying count-based exploration and intrinsic motivation
Marc Bellemare, Sriram Srinivasan, Georg Ostrovski, Tom Schaul, David Saxton, and Remi Munos · 2016
Cited alongside, same era.
Nips 2016 tutorial: Generative adversarial networks
Ian Goodfellow · 2016
Cited alongside, same era.
Karol Gregor, Danilo Jimenez Rezende, and Daan Wierstra · 2016
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Cited alongside, same era.
Deep exploration via bootstrapped dqn
Ian Osband, Charles Blundell, Alexander Pritzel, and Benjamin Van Roy · 2016
Cited alongside, same era.
Later among the works it cites.
Unsupervised feature learning via non-parametric instance-level discrimination
Zhirong Wu, Yuanjun Xiong, Stella Yu, and Dahua Lin · 2018
Later among the works it cites.
Autonomous exploration for navigating in non-stationary cmps, 2019
Pratik Gajane, Ronald Ortner, Peter Auer, and Csaba Szepesvari · 2019
Later among the works it cites.
Marginalized state distribution entropy regularization in policy optimization
Riashat Islam, Zafarali Ahmed, and Doina Precup · 2019
Later among the works it cites.
Efficient exploration via state marginal matching
Lisa Lee, Benjamin Eysenbach, Emilio Parisotto, Eric Xing, Sergey Levine, and Ruslan Salakhutdinov · 2019
Later among the works it cites.
Geometric losses for distributional learning
Arthur Mensch, Mathieu Blondel, and Gabriel Peyré · 2019
Later among the works it cites.
Wasserstein dependency measure for representation learning
Sherjil Ozair, Corey Lynch, Yoshua Bengio, Aaron Van den Oord, Sergey Levine, and Pierre Sermanet · 2019
Later among the works it cites.
Skew-fit: State-covering self-supervised reinforcement learning
Vitchyr H Pong, Murtaza Dalal, Steven Lin, Ashvin Nair, Shikhar Bahl, and Sergey Levine · 2019
Later among the works it cites.
No-regret exploration in goal-oriented reinforcement learning
Jean Tarbouriech, Evrard Garcelon, Michal Valko, Matteo Pirotta, and Alessandro Lazaric · 2019
Later among the works it cites.
Near-optimal regret bounds for stochastic shortest path
Alon Cohen, Haim Kaplan, Yishay Mansour, and Aviv Rosenberg · 2020
Later among the works it cites.
Reward-free exploration for reinforcement learning
Chi Jin, Akshay Krishnamurthy, Max Simchowitz, and Tiancheng Yu · 2020
Later among the works it cites.
An intrinsically-motivated approach for learning highly exploring and fast mixing policies
Mirco Mutti and Marcello Restelli · 2020
Later among the works it cites.
Behaviour suite for reinforcement learning
Ian Osband, Yotam Doron, Matteo Hessel, John Aslanides, Eren Sezener, Andre Saraiva, Katrina McKinney, Tor Lattimore, Csaba Szepesvári, Satinder Singh, Benjamin Van Roy, Richard Sutton, David Silver, and Hado van Hasselt · 2020
Later among the works it cites.