Fetching the paper…
Reading the bibliography…
We take initial steps in studying PAC-MDP algorithms with limited adaptivity, that is, algorithms that change its exploration policy as infrequently as possible during regret minimization.
Finite-time analysis of the multiarmed bandit problem
P. Auer, N. Cesa-Bianchi, and P. Fischer · 2002
Earlier work this paper cites.
R-max-a general polynomial time algorithm for near-optimal reinforcement learning
R. I. Brafman and M. Tennenholtz · 2002
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
M. Kearns and S. Singh · 2002
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
T. Jaksch, R. Ortner, and P. Auer · 2010
Earlier work this paper cites.
Designing a pilot sequential multiple assignment randomized trial for developing an adaptive treatment strategy
D. Almirall, S. N. Compton, M. Gunlicks-Stoessel, N. Duan, and S. A. Murphy · 2012
Earlier work this paper cites.
A ”smart” design for building individualized treatment sequences
H. Lei, I. Nahum-Shani, K. Lynch, D. Oslin, and S. A. Murphy · 2012
Earlier work this paper cites.
Online learning with switching costs and other adaptive adversaries
N. Cesa-Bianchi, O. Dekel, and O. Shamir · 2013
Earlier work this paper cites.
(more) efficient reinforcement learning via posterior sampling
I. Osband, D. Russo, and B. Van Roy · 2013
Earlier work this paper cites.
Introduction to smart designs for the development of adaptive interventions: with application to weight loss research
D. Almirall, I. Nahum-Shani, N. E. Sherwood, and S. A. Murphy · 2014
Earlier work this paper cites.
Concurrent pac rl
Z. Guo and E. Brunskill · 2015
Cited alongside, same era.
Personalized ad recommendation systems for life-time value optimization with guarantees
G. Theocharous, P. S. Thomas, and M. Ghavamzadeh · 2015
Cited alongside, same era.
Batched bandit problems
V. Perchet, P. Rigollet, S. Chassang, E. Snowberg, et al · 2016
Cited alongside, same era.
Machine-learning-assisted materials discovery using failed experiments
P. Raccuglia, K. C. Elbert, P. D. Adler, C. Falk, M. B. Wenny, A. Mollo, M. Zeller, S. A. Friedler, J. Schrier, and A. J. Norquist · 2016
Cited alongside, same era.
Optimistic posterior sampling for reinforcement learning: worst-case regret bounds
S. Agrawal and R. Jia · 2017
Cited alongside, same era.
Minimax regret bounds for reinforcement learning
M. G. Azar, I. Osband, and R. Munos · 2017
Cited alongside, same era.
A lyapunov-based approach to safe reinforcement learning
Y. Chow, O. Nachum, E. Duenez-Guzman, and M. Ghavamzadeh · 2018
Later among the works it cites.
Policy certificates: Towards accountable reinforcement learning
C. Dann, L. Li, W. Wei, and E. Brunskill · 2018
Later among the works it cites.
Minimax bounds on stochastic batched convex optimization
J. Duchi, F. Ruan, and C. Yun · 2018
Later among the works it cites.
Is Q-learning provably efficient?
C. Jin, Z. Allen-Zhu, S. Bubeck, and M. I. Jordan · 2018
Later among the works it cites.
Learning to optimize join queries with deep reinforcement learning
S. Krishnan, Z. Yang, K. Goldberg, J. Hellerstein, and I. Stoica · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Data-efficient policy evaluation through behavior policy search
J. P. Hanna, P. S. Thomas, P. Stone, and S. Niekum · 2017
Cited alongside, same era.
Device placement optimization with reinforcement learning
A. Mirhoseini, H. Pham, Q. V. Le, B. Steiner, R. Larsen, Y. Zhou, N. Kumar, M. Norouzi, S. Bengio, and J. Dean · 2017
Cited alongside, same era.
A survey on compiler autotuning using machine learning
A. H. Ashouri, W. Killian, J. Cavazos, G. Palermo, and C. Silvano · 2018
Cited alongside, same era.
Z. Gao, Y. Han, Z. Ren, and Z. Zhou · 2019
Closest in time.
Incomplete conditional density estimation for fast materials discovery
P. Nguyen, T. Tran, S. Gupta, S. Rana, M. Barnett, and S. Venkatesh · 2019
Closest in time.
Convergent policy optimization for safe reinforcement learning
M. Yu, Z. Yang, M. Kolar, and Z. Wang · 2019
Closest in time.