Fetching the paper…
Reading the bibliography…
In probably approximately correct (PAC) reinforcement learning (RL), an agent is required to identify an $\epsilon$-optimal policy with probability $1-\delta$.
An analysis of approximations for maximizing submodular set functions—i
George L Nemhauser, Laurence A Wolsey, and Marshall L Fisher · 1978
Earlier work this paper cites.
Algorithms for solving for the minimal flow in a network
Yu V Voitishin · 1980
Earlier work this paper cites.
Asymptotically efficient adaptive allocation rules
Tze Leung Lai and Herbert Robbins · 1985
Earlier work this paper cites.
Minimum flows in (s, t) planar networks
Veena Adlakha, Barbara Gladysz, and Jerzy Kamburowski · 1991
Earlier work this paper cites.
Efficient reinforcement learning
Claude-Nicolas Fiechter · 1994
Earlier work this paper cites.
Optimal adaptive policies for markov decision processes
Apostolos N Burnetas and Michael N Katehakis · 1997
Earlier work this paper cites.
An alternate linear algorithm for the minimum flow problem
Veena G Adlakha · 1999
Earlier work this paper cites.
A sparse sampling algorithm for near-optimal planning in large markov decision processes
Michael J. Kearns, Yishay Mansour, and Andrew Y. Ng · 2002
Earlier work this paper cites.
On the Sample Complexity of Reinforcement Learning
Sham Kakade · 2003
Earlier work this paper cites.
Sequential and parallel algorithms for minimum flows
Eleonor Ciurea and Laura Ciupala · 2004
Earlier work this paper cites.
The Sample Complexity of Exploration in the Multi-Armed Bandit Problem
S. Mannor and J. Tsitsiklis · 2004
Earlier work this paper cites.
Action Elimination and Stopping Conditions for the Multi-Armed Bandit and Reinforcement Learning Problems
E. Even-Dar, S. Mannor, and Y. Mansour · 2006
Earlier work this paper cites.
Optimistic linear programming gives logarithmic regret for irreducible mdps
Ambuj Tewari and Peter L. Bartlett · 2007
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Peter Auer, Thomas Jaksch, and Ronald Ortner · 2008
Earlier work this paper cites.
Reinforcement learning in finite mdps: Pac analysis
Alexander L Strehl, Lihong Li, and Michael L Littman · 2009
Earlier work this paper cites.
Open loop optimistic planning
S Bubeck and R Munos · 2010
Earlier work this paper cites.
graph2tab, a library to convert experimental workflow graphs into tabular formats
Marco Brandizi, Natalja Kurbatova, Ugis Sarkans, and Philippe Rocca-Serra · 2012
Earlier work this paper cites.
Submodular function maximization
Andreas Krause and Daniel Golovin · 2014
Earlier work this paper cites.
Simple regret optimization in online planning for markov decision processes
Zohar Feldman and Carmel Domshlak · 2014
Cited alongside, same era.
Sample complexity of episodic fixed-horizon reinforcement learning
Christoph Dann and Emma Brunskill · 2015
Cited alongside, same era.
Optimal best arm identification with fixed confidence
Aurélien Garivier and Emilie Kaufmann · 2016
Cited alongside, same era.
On the complexity of best-arm identification in multi-armed bandit models
Emilie Kaufmann, Olivier Cappé, and Aurélien Garivier · 2016
Cited alongside, same era.
Blazing the trails before beating the path: Sample-efficient monte-carlo planning
J.-B. Grill, M. Valko, and R. Munos · 2016
Cited alongside, same era.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Cited alongside, same era.
An asymptotically optimal primal-dual incremental algorithm for contextual linear bandits
Andrea Tirinzoni, Matteo Pirotta, Marcello Restelli, and Alessandro Lazaric · 2020
Later among the works it cites.
Pc-pg: Policy cover directed exploration for provable policy gradient learning
Alekh Agarwal, Mikael Henaff, Sham Kakade, and Wen Sun · 2020
Later among the works it cites.
Fast active learning for pure exploration in reinforcement learning
Pierre Ménard, Omar Darwiche Domingues, Anders Jonsson, Emilie Kaufmann, Edouard Leurent, and Michal Valko · 2021
Later among the works it cites.
Episodic reinforcement learning in finite mdps: Minimax lower bounds revisited
Omar Darwiche Domingues, Pierre Ménard, Emilie Kaufmann, and Michal Valko · 2021
Later among the works it cites.
Adaptive sampling for best policy identification in markov decision processes
Aymen Al Marjani and Alexandre Proutiere · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Exploration in structured reinforcement learning
Jungseul Ok, Alexandre Proutiere, and Damianos Tranos · 2018
Cited alongside, same era.
Policy certificates: Towards accountable reinforcement learning
Christoph Dann, Lihong Li, Wei Wei, and Emma Brunskill · 2019
Cited alongside, same era.
Almost horizon-free structure-aware best policy identification with a generative model
Andrea Zanette, Mykel J. Kochenderfer, and Emma Brunskill · 2019
Cited alongside, same era.
Non-asymptotic gap-dependent regret bounds for tabular mdps
Max Simchowitz and Kevin G. Jamieson · 2019
Cited alongside, same era.
Pure exploration with multiple correct answers
Rémy Degenne and Wouter M. Koolen · 2019
Cited alongside, same era.
Non-asymptotic pure exploration by solving games
Rémy Degenne, Wouter M Koolen, and Pierre Ménard · 2019
Cited alongside, same era.
Navigating to the best policy in markov decision processes
Aymen Al Marjani, Aurélien Garivier, and Alexandre Proutiere · 2021
Later among the works it cites.
A fully problem-dependent regret lower bound for finite-horizon mdps
Andrea Tirinzoni, Matteo Pirotta, and Alessandro Lazaric · 2021
Later among the works it cites.
Christoph Dann, Teodor V. Marinov, Mehryar Mohri, and Julian Zimmert · 2021
Later among the works it cites.
Fine-grained gap-dependent bounds for tabular mdps via adaptive multi-step bootstrap
Haike Xu, Tengyu Ma, and Simon S Du · 2021
Later among the works it cites.
Nonasymptotic sequential tests for overlapping hypotheses applied to near-optimal arm identification in bandit models
Aurélien Garivier and Emilie Kaufmann · 2021
Later among the works it cites.
Towards instance-optimal offline reinforcement learning with pessimism
Ming Yin and Yu-Xiang Wang · 2021
Later among the works it cites.
Adaptive reward-free exploration
Emilie Kaufmann, Pierre Ménard, Omar Darwiche Domingues, Anders Jonsson, Edouard Leurent, and Michal Valko · 2021
Later among the works it cites.
Near optimal reward-free reinforcement learning
Zihan Zhang, Simon Du, and Xiangyang Ji · 2021
Later among the works it cites.
Regret analysis in deterministic reinforcement learning
Damianos Tranos and Alexandre Proutière · 2021
Later among the works it cites.
Fast pure exploration via frank-wolfe
Po-An Wang, Ruo-Chun Tzeng, and Alexandre Proutiere · 2021
Later among the works it cites.
rlberry - A Reinforcement Learning Library for Research and Education
Omar Darwiche Domingues, Yannis Flet-Berliac, Edouard Leurent, Pierre Ménard, Xuedong Shang, and Michal Valko · 2021
Later among the works it cites.
Beyond no regret: Instance-dependent PAC reinforcement learning
Andrew Wagenmaker, Max Simchowitz, and Kevin G. Jamieson · 2022
Closest in time.