Fetching the paper…
Reading the bibliography…
We investigate the classical active pure exploration problem in Markov Decision Processes, where the agent sequentially selects actions and, from the resulting system trajectory, aims at identifying the best policy as fast as possible.
Perturbation theory and finite markov chains
P. Schweitzer · 1968
Earlier work this paper cites.
2 - inequalities and laws of large numbers
P. Hall and C.C. Heyde · 1980
Earlier work this paper cites.
Asymptotically efficient adaptive allocation rules
T.L. Lai and H. Robbins · 1985
Earlier work this paper cites.
Efficient reinforcement learning
C. Fiechter · 1994
Earlier work this paper cites.
Optimal adaptive policies for markov decision processes
Apostolos N. Burnetas and Michael N. Katehakis · 1997
Earlier work this paper cites.
Finite-sample convergence rates for q-learning and indirect algorithms
Michael Kearns and Satinder Singh · 1998
Earlier work this paper cites.
A sparse sampling algorithm for near-optimal planning in large markov decision processes
Michael Kearns, Yishay Mansour, and Andrew Y. Ng · 1999
Earlier work this paper cites.
On the sample complexity of reinforcement learning
Sham Machandranath Kakade · 2003
Earlier work this paper cites.
Elements of Information Theory (Wiley Series in Telecommunications and Signal Processing)
Thomas M. Cover and Joy A. Thomas · 2006
Earlier work this paper cites.
Action elimination and stopping conditions for the multi-armed bandit and reinforcement learning problems
Eyal Even-Dar, Shie Mannor, and Yishay Mansour · 2006
Earlier work this paper cites.
Markov chains and mixing times
David A. Levin, Yuval Peres, and Elizabeth L. Wilmer · 2006
Earlier work this paper cites.
An analysis of model-based interval estimation for markov decision processes
Alexander L Strehl and Michael L Littman · 2008
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Peter Auer, Thomas Jaksch, and Ronald Ortner · 2009
Earlier work this paper cites.
Optimism in reinforcement learning and kullback-leibler divergence
Sarah Filippi, Olivier Cappé, and Aurélien Garivier · 2010
Earlier work this paper cites.
Algorithms for Reinforcement Learning
Csaba Szepesvari · 2010
Cited alongside, same era.
Convergence of adaptive and interacting markov chain monte carlo algorithms
G. Fort, É. Moulines, and P. Priouret · 2011
Cited alongside, same era.
Minimax pac bounds on the sample complexity of reinforcement learning with a generative model
Mohammad Gheshlaghi Azar, Rémi Munos, and Hilbert J Kappen · 2013
Cited alongside, same era.
Pac optimal planning for invasive species management : Improved exploration for reinforcement learning from simulator-defined mdps
Thomas Dietterich, Majid Taleghan, and Mark Crowley · 2013
Cited alongside, same era.
Optimal best arm identification with fixed confidence
Aurélien Garivier and Emilie Kaufmann · 2016
Cited alongside, same era.
On the complexity of best-arm identification in multi-armed bandit models
Pc-pg: Policy cover directed exploration for provable policy gradient learning
Alekh Agarwal, Mikael Henaff, Sham Kakade, and Wen Sun · 2020
Later among the works it cites.
Model-based reinforcement learning with a generative model is minimax optimal
Alekh Agarwal, Sham Kakade, and Lin F. Yang · 2020
Later among the works it cites.
Planning in markov decision processes with gap-dependent sample complexity
Anders Jonsson, Emilie Kaufmann, Pierre Menard, Omar Darwiche Domingues, Edouard Leurent, and Michal Valko · 2020
Later among the works it cites.
Bandit Algorithms
Tor Lattimore and Csaba Szepesvári · 2020
Later among the works it cites.
Breaking the sample size barrier in model-based reinforcement learning with a generative model
Gen Li, Yuting Wei, Yuejie Chi, Yuantao Gu, and Yuxin Chen · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Emilie Kaufmann, Olivier Cappé, and Aurélien Garivier · 2016
Cited alongside, same era.
Simple bayesian algorithms for best arm identification
D. Russo · 2016
Cited alongside, same era.
Mengdi Wang · 2017
Cited alongside, same era.
Mixture martingales revisited with applications to sequential tests and confidence intervals
Emilie Kaufmann and Wouter M. Koolen · 2018
Cited alongside, same era.
Reinforcement Learning: An Introduction
Richard S. Sutton and Andrew G. Barto · 2018
Cited alongside, same era.
Near-optimal time and sample complexities for solving markov decision processes with a generative model
Aaron Sidford, Mengdi Wang, Xian Wu, Lin Yang, and Yinyu Ye · 2018
Cited alongside, same era.
Non-asymptotic pure exploration by solving games
Rémy Degenne, Wouter M. Koolen, and Pierre Ménard · 2019
Cited alongside, same era.
Sample complexity of asynchronous q-learning: Sharper analysis and variance reduction, 2020
Gen Li, Yuting Wei, Yuejie Chi, Yuantao Gu, and Yuxin Chen · 2020
Later among the works it cites.
Fast active learning for pure exploration in reinforcement learning
Pierre M’enard, O. D. Domingues, Anders Jonsson, E. Kaufmann, Edouard Leurent, and Michal Valko · 2020
Later among the works it cites.
Episodic reinforcement learning in finite mdps: Minimax lower bounds revisited
O. D. Domingues, Pierre M’enard, E. Kaufmann, and Michal Valko · 2021
Closest in time.
Provably correct optimization and exploration with non-linear policies
Fei Feng, W. Yin, Alekh Agarwal, and Lin F. Yang · 2021
Closest in time.
Adaptive reward-free exploration
Emilie Kaufmann, Pierre Ménard, Omar Darwiche Domingues, Anders Jonsson, Edouard Leurent, and Michal Valko · 2021
Closest in time.
Is q-learning minimax optimal? a tight sample complexity analysis, 2021
Gen Li, Changxiao Cai, Yuxin Chen, Yuantao Gu, Yuting Wei, and Yuejie Chi · 2021
Closest in time.
Adaptive sampling for best policy identification in markov decision processes
Aymen Al Marjani and Alexandre Proutiere · 2021
Closest in time.
Beyond No Regret: Instance-Dependent PAC Reinforcement Learning
Andrew Wagenmaker, Max Simchowitz, and Kevin Jamieson · 2021
Closest in time.
Cautiously optimistic policy optimization and exploration with linear function approximation
Andrea Zanette, Ching-An Cheng, and Alekh Agarwal · 2021
Closest in time.