Fetching the paper…
Reading the bibliography…
The performance of a reinforcement learning algorithm can vary drastically during learning because of exploration.
Optimal composition of real-time systems
Zilberstein, S. and Russell, S · 1996
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Kakade, S. M. and Langford, J · 2002
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Kearns, M. and Singh, S · 2002
Earlier work this paper cites.
On the sample complexity of reinforcement learning
Kakade, S · 2003
Earlier work this paper cites.
PAC model-free reinforcement learning
Strehl, A. L., Li, L., Wiewiora, E., Langford, J., and Littman, M. L · 2006
Earlier work this paper cites.
Knows what it knows: a framework for self-aware learning
Li, L., Littman, M. L., and Walsh, T. J · 2008
Earlier work this paper cites.
An analysis of model-based interval estimation for markov decision processes
Strehl, A. L. and Littman, M. L · 2008
Earlier work this paper cites.
Reinforcement learning in finite MDPs: PAC analysis
Strehl, A. L., Li, L., and Littman, M. L · 2009
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Jaksch, T., Ortner, R., and Auer, P · 2010
Earlier work this paper cites.
Model-based reinforcement learning with nearly tight exploration complexity bounds
Szita, I. and Szepesvári, C · 2010
Earlier work this paper cites.
Pac bounds for discounted mdps
Lattimore, T. and Hutter, M · 2012
Earlier work this paper cites.
(more) efficient reinforcement learning via posterior sampling
Osband, I., Russo, D., and Van Roy, B · 2013
Earlier work this paper cites.
Safe policy iteration
Pirotta, M., Restelli, M., Pecorino, A., and Calandriello, D · 2013
Cited alongside, same era.
Online learning in mdps with side information
Abbasi-Yadkori, Y. and Neu, G · 2014
Cited alongside, same era.
Sample complexity of episodic fixed-horizon reinforcement learning
Dann, C. and Brunskill, E · 2015
Cited alongside, same era.
Contextual Markov decision processes
Hallak, A., Di Castro, D., and Mannor, S · 2015
Cited alongside, same era.
Safe policy improvement by minimizing robust baseline regret
Ghavamzadeh, M., Petrik, M., and Chow, Y · 2016
Cited alongside, same era.
Contextual decision processes with low bellman rank are pac-learnable
Jiang, N., Krishnamurthy, A., Agarwal, A., Langford, J., and Schapire, R. E · 2017
Later among the works it cites.
Fairness incentives for myopic agents
Kannan, S., Kearns, M., Morgenstern, J., Pai, M., Roth, A., Vohra, R., and Wu, Z. S · 2017
Later among the works it cites.
Multi-step off-policy learning without importance sampling ratios
Mahmood, A. R., Yu, H., and Sutton, R. S · 2017
Later among the works it cites.
On oracle-efficient pac reinforcement learning with rich observations
Dann, C., Jiang, N., Krishnamurthy, A., Agarwal, A., Langford, J., and Schapire, R. E · 2018
Closest in time.
Uniform, nonparametric, non-asymptotic confidence sequences
Howard, S. R., Ramdas, A., Mc Auliffe, J., and Sekhon, J · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jabbari, S., Joseph, M., Kearns, M., Morgenstern, J., and Roth, A · 2016
Cited alongside, same era.
Doubly robust off-policy value evaluation for reinforcement learning
Jiang, N. and Li, L · 2016
Cited alongside, same era.
Fairness in learning: Classic and contextual bandits
Joseph, M., Kearns, M., Morgenstern, J. H., and Roth, A · 2016
Cited alongside, same era.
Generalization and exploration via randomized value functions
Osband, I., Van Roy, B., and Wen, Z · 2016
Cited alongside, same era.
Data-efficient off-policy policy evaluation for reinforcement learning
Thomas, P. and Brunskill, E · 2016
Cited alongside, same era.
Minimax regret bounds for reinforcement learning
Azar, M. G., Osband, I., and Munos, R · 2017
Cited alongside, same era.
Unifying pac and regret: Uniform pac bounds for episodic reinforcement learning
Dann, C., Lattimore, T., and Brunskill, E · 2017
Cited alongside, same era.
Closest in time.
Is Q-learning provably efficient?
Jin, C., Allen-Zhu, Z., Bubeck, S., and Jordan, M. I · 2018
Closest in time.
Bandit Algorithms
Lattimore, T. and Czepesvari, C · 2018
Closest in time.
Markov decision processes with continuous side information
Modi, A., Jiang, N., Singh, S., and Tewari, A · 2018
Closest in time.
The externalities of exploration and how data diversity helps exploitation
Raghavan, M., Slivkins, A., Vaughan, J. W., and Wu, Z. S · 2018
Closest in time.
High-confidence error estimates for learned value functions
Sajed, T., Chung, W., and White, M · 2018
Closest in time.
Zanette, A. and Brunskill, E · 2019
Closest in time.