Fetching the paper…
Reading the bibliography…
Past research on interactive decision making problems (bandits, reinforcement learning, etc.) mostly focuses on the minimax regret that measures the algorithm's performance on the hardest instance.
Sequential design of experiments
Herman Chernoff · 1959
Earlier work this paper cites.
Asymptotically efficient adaptive allocation rules
Tze Leung Lai, Herbert Robbins, et al · 1985
Earlier work this paper cites.
Asymptotically efficient adaptive allocation schemes for controlled iid processes: Finite parameter space
A Rajeev, Demosthenis Teneketzis, and Venkatachalam Anantharam · 1989
Earlier work this paper cites.
Nonparametric estimation of nonstationary spatial covariance structure
Paul D Sampson and Peter Guttorp · 1992
Earlier work this paper cites.
Optimal adaptive policies for Markov decision processes
Apostolos N Burnetas and Michael N Katehakis · 1997
Earlier work this paper cites.
Asymptotically efficient adaptive choice of control laws incontrolled markov chains
Todd L Graves and Tze Leung Lai · 1997
Earlier work this paper cites.
Elements of information theory
Thomas M Cover · 1999
Earlier work this paper cites.
Empirical Processes in M-estimation
Sara Van de Geer · 2000
Earlier work this paper cites.
Rényi divergence measures for commonly used univariate continuous distributions
Manuel Gil, Fady Alajaji, and Tamas Linder · 2013
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Earlier work this paper cites.
Taming the monster: A fast and simple algorithm for contextual bandits
Alekh Agarwal, Daniel Hsu, Satyen Kale, John Langford, Lihong Li, and Robert Schapire · 2014
Earlier work this paper cites.
Lipschitz bandits: Regret lower bound and optimal algorithms
Stefan Magureanu, Richard Combes, and Alexandre Proutiere · 2014
Earlier work this paper cites.
Rényi divergence and kullback-leibler divergence
Tim Van Erven and Peter Harremos · 2014
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Earlier work this paper cites.
Mastering the game of Go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al · 2016
Earlier work this paper cites.
Minimax regret bounds for reinforcement learning
Mohammad Gheshlaghi Azar, Ian Osband, and Rémi Munos · 2017
Cited alongside, same era.
The end of optimism? an asymptotic analysis of finite-armed linear bandits
Tor Lattimore and Csaba Szepesvari · 2017
Cited alongside, same era.
Mastering the game of Go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, et al · 2017
Cited alongside, same era.
Is q-learning provably efficient?
Chi Jin, Zeyuan Allen-Zhu, Sebastien Bubeck, and Michael I Jordan · 2018
Cited alongside, same era.
Exploration in structured reinforcement learning
Jungseul Ok, Alexandre Proutiere, and Damianos Tranos · 2018
Cited alongside, same era.
Bandit algorithms
Tor Lattimore and Csaba Szepesvári · 2020
Later among the works it cites.
Finding all ϵ \epsilon -good arms in stochastic bandits
Blake Mason, Lalit Jain, Ardhendu Tripathy, and Robert Nowak · 2020
Later among the works it cites.
A novel confidence-based algorithm for structured bandits
Andrea Tirinzoni, Alessandro Lazaric, and Marcello Restelli · 2020
Later among the works it cites.
Kefan Dong, Jiaqi Yang, and Tengyu Ma · 2021
Later among the works it cites.
The statistical complexity of interactive decision making
Dylan J Foster, Sham M Kakade, Jian Qian, and Alexander Rakhlin · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Christopher Berner, Greg Brockman, Brooke Chan, Vicki Cheung, Przemyslaw Debiak, Christy Dennison, David Farhi, Quirin Fischer, Shariq Hashme, Chris Hesse, et al · 2019
Cited alongside, same era.
Nearly minimax-optimal regret for linearly parameterized bandits
Yingkai Li, Yining Wang, and Yuan Zhou · 2019
Cited alongside, same era.
Non-asymptotic gap-dependent regret bounds for tabular MDPs
Max Simchowitz and Kevin G Jamieson · 2019
Cited alongside, same era.
Grandmaster level in starcraft ii using multi-agent reinforcement learning
Oriol Vinyals, Igor Babuschkin, Wojciech M Czarnecki, Michaël Mathieu, Andrew Dudzik, Junyoung Chung, David H Choi, Richard Powell, Timo Ewalds, Petko Georgiev, et al · 2019
Cited alongside, same era.
Andrea Zanette and Emma Brunskill · 2019
Cited alongside, same era.
Structure adaptive algorithms for stochastic bandits
Rémy Degenne, Han Shao, and Wouter Koolen · 2020
Cited alongside, same era.
Beyond UCB: Optimal and efficient contextual bandits with regression oracles
Dylan Foster and Alexander Rakhlin · 2020
Cited alongside, same era.
Samarth Gupta, Shreyas Chaudhari, Gauri Joshi, and Osman Yağan · 2021
Later among the works it cites.
Asymptotically optimal information-directed sampling
Johannes Kirschner, Tor Lattimore, Claire Vernade, and Csaba Szepesvári · 2021
Later among the works it cites.
Hypothesis testing: Lecture notes for math 6263 at georgia tech
Cheng Mao · 2021
Later among the works it cites.
A fully problem-dependent regret lower bound for finite-horizon mdps
Andrea Tirinzoni, Matteo Pirotta, and Alessandro Lazaric · 2021
Later among the works it cites.
Fine-grained gap-dependent bounds for tabular mdps via adaptive multi-step bootstrap
Haike Xu, Tengyu Ma, and Simon S Du · 2021
Later among the works it cites.
Q-learning with logarithmic regret
Kunhe Yang, Lin Yang, and Simon Du · 2021
Later among the works it cites.
Magnetic control of tokamak plasmas through deep reinforcement learning
Jonas Degrave, Federico Felici, Jonas Buchli, Michael Neunert, Brendan Tracey, Francesco Carpanese, Timo Ewalds, Roland Hafner, Abbas Abdolmaleki, Diego de Las Casas, et al · 2022
Closest in time.
On the complexity of all ε \varepsilon -best arms identification
Aymen Al Marjani, Tomáš Kocák, and Aurélien Garivier · 2022
Closest in time.
Near instance-optimal pac reinforcement learning for deterministic mdps
Andrea Tirinzoni, Aymen Al-Marjani, and Emilie Kaufmann · 2022
Closest in time.