Fetching the paper…
Reading the bibliography…
Existing metrics for reinforcement learning (RL) such as regret, PAC bounds, or uniform-PAC (Dann et al., 2017), typically evaluate the cumulative performance, while allowing the agent to play an arbitrarily bad policy at any finite time t.
The equivalence of two extremum problems
J. Kiefer and J. Wolfowitz · 1960
Earlier work this paper cites.
Action elimination and stopping conditions for the multi-armed bandit and reinforcement learning problems
E. Even-Dar, S. Mannor, Y. Mansour, and S. Mahadevan · 2006
Earlier work this paper cites.
Online linear optimization and adaptive routing
B. Awerbuch and R. Kleinberg · 2008
Earlier work this paper cites.
High-probability regret bounds for bandit online linear optimization
P. L. Bartlett, V. Dani, T. P. Hayes, S. M. Kakade, A. Rakhlin, and A. Tewari · 2008
Earlier work this paper cites.
Stochastic linear optimization under bandit feedback
V. Dani, T. P. Hayes, and S. M. Kakade · 2008
Earlier work this paper cites.
Regret bounds and minimax policies under partial monitoring
J.-Y. Audibert and S. Bubeck · 2010
Earlier work this paper cites.
Best arm identification in multi-armed bandits
J.-Y. Audibert, S. Bubeck, and R. Munos · 2010
Earlier work this paper cites.
A contextual-bandit approach to personalized news article recommendation
L. Li, W. Chu, J. Langford, and R. E. Schapire · 2010
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Y. Abbasi-Yadkori, D. Pál, and C. Szepesvári · 2011
Earlier work this paper cites.
Pac subset selection in stochastic multi-armed bandits
S. Kalyanakrishnan, A. Tewari, P. Auer, and P. Stone · 2012
Earlier work this paper cites.
Thompson sampling for contextual bandits with linear payoffs
S. Agrawal and N. Goyal · 2013
Earlier work this paper cites.
Taming the monster: A fast and simple algorithm for contextual bandits
A. Agarwal, D. Hsu, S. Kale, J. Langford, L. Li, and R. Schapire · 2014
Earlier work this paper cites.
Combinatorial pure exploration of multi-armed bandits
S. Chen, T. Lin, I. King, M. R. Lyu, and W. Chen · 2014
Cited alongside, same era.
lil’ucb: An optimal exploration algorithm for multi-armed bandits
K. Jamieson, M. Malloy, R. Nowak, and S. Bubeck · 2014
Cited alongside, same era.
Multi-armed bandit models for the optimal design of clinical trials: benefits and challenges
S. S. Villar, J. Bowden, and J. Wason · 2015
Cited alongside, same era.
Anytime optimal algorithms in stochastic multi-armed bandits
R. Degenne and V. Perchet · 2016
Cited alongside, same era.
Minimax regret bounds for reinforcement learning
M. G. Azar, I. Osband, and R. Munos · 2017
Cited alongside, same era.
Unifying pac and regret: Uniform pac bounds for episodic reinforcement learning
C. Dann, T. Lattimore, and E. Brunskill · 2017
Bias no more: high-probability data-dependent regret bounds for adversarial bandits and mdps
C.-W. Lee, H. Luo, C.-Y. Wei, and M. Zhang · 2020
Later among the works it cites.
Geometric exploration for online control
O. Plevrakis and E. Hazan · 2020
Later among the works it cites.
On reward-free reinforcement learning with linear function approximation
R. Wang, S. S. Du, L. Yang, and R. R. Salakhutdinov · 2020
Later among the works it cites.
Uniform-pac bounds for reinforcement learning with linear function approximation
J. He, D. Zhou, and Q. Gu · 2021
Later among the works it cites.
Achieving near instance-optimality and minimax-optimality in stochastic and adversarial linear bandits simultaneously
C.-W. Lee, H. Luo, C.-Y. Wei, M. Zhang, and X. Zhang · 2021
Later among the works it cites.
Instance-optimal pac algorithms for contextual bandits
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Is q-learning provably efficient?
C. Jin, Z. Allen-Zhu, S. Bubeck, and M. I. Jordan · 2018
Cited alongside, same era.
Efficient pure exploration in adaptive round model
T. Jin, J. Shi, X. Xiao, and E. Chen · 2019
Cited alongside, same era.
Generalized policy elimination: an efficient algorithm for nonparametric contextual bandits
A. Bibaut, A. Chambaz, and M. Laan · 2020
Cited alongside, same era.
Learning adversarial markov decision processes with bandit feedback and unknown transition
C. Jin, T. Jin, H. Luo, S. Sra, and T. Yu · 2020
Cited alongside, same era.
Bandit algorithms
T. Lattimore and C. Szepesvári · 2020
Cited alongside, same era.
Learning with good feature representations in bandits and in rl with a generative model
T. Lattimore, C. Szepesvari, and G. Weisz · 2020
Cited alongside, same era.
Z. Li, L. Ratliff, K. G. Jamieson, L. Jain, et al · 2022
Later among the works it cites.
Beyond no regret: Instance-dependent pac reinforcement learning
A. J. Wagenmaker, M. Simchowitz, and K. Jamieson · 2022
Later among the works it cites.
Contextual bandits with large action spaces: Made practical
Y. Zhu, D. J. Foster, J. Langford, and P. Mineiro · 2022
Later among the works it cites.
Last-iterate convergent policy gradient primal-dual methods for constrained mdps
D. Ding, C.-Y. Wei, K. Zhang, and A. Ribeiro · 2023
Later among the works it cites.
Efficient batched algorithm for contextual linear bandits with large action space via soft elimination
O. Hanna, L. Yang, and C. Fragouli · 2023
Later among the works it cites.
Reload: Reinforcement learning with optimistic ascent-descent for last-iterate convergence in constrained mdps
T. Moskovitz, B. O’Donoghue, V. Veeriah, S. Flennerhag, S. Singh, and T. Zahavy · 2023
Later among the works it cites.
Uniform-pac guarantees for model-based rl with bounded eluder dimension
Y. Wu, J. He, and Q. Gu · 2023
Later among the works it cites.