Fetching the paper…
Reading the bibliography…
Exploration in reinforcement learning (RL) suffers from the curse of dimensionality when the state-action space is large.
Sample-optimal parametric q-learning with linear transition models
Yang, L. F. and Wang, M. (2019) · 1902
Earlier work this paper cites.
Contextual markov decision processes using generalized linear models
Modi, A. and Tewari, A. (2019) · 1903
Earlier work this paper cites.
Dynamic programming
Bellman, R. (1966) · 1966
Earlier work this paper cites.
On Tail Probability for Martigales
Freedman, D. A. (1975) · 1975
Earlier work this paper cites.
Residual algorithms: Reinforcement learning with function approximation
Baird, L. (1995) · 1995
Earlier work this paper cites.
Analysis of temporal-diffference learning with function approximation
Tsitsiklis, J. N. and Van Roy, B. (1997) · 1997
Earlier work this paper cites.
On the sample complexity of reinforcement learning
Kakade, S. M. (2003) · 2003
Earlier work this paper cites.
Kernel methods for pattern analysis
Shawe-Taylor, J., Cristianini, N., et al. (2004) · 2004
Earlier work this paper cites.
Pac model-free reinforcement learning
Strehl, A. L., Li, L., Wiewiora, E., Langford, J., and Littman, M. L. (2006) · 2006
Earlier work this paper cites.
Stochastic linear optimization under bandit feedback
Dani, V., Hayes, T. P., and Kakade, S. M. (2008) · 2008
Earlier work this paper cites.
An analysis of linear models, linear value-function approximation, and feature selection for reinforcement learning
Parr, R., Li, L., Taylor, G., Painter-Wakefield, C., and Littman, M. L. (2008) · 2008
Earlier work this paper cites.
Random features for large-scale kernel machines
Rahimi, A. and Recht, B. (2008) · 2008
Earlier work this paper cites.
Reinforcement learning in finite MDPs: PAC analysis
Strehl, A. L., Li, L., and Littman, M. L. (2009) · 2009
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Jaksch, T., Ortner, R., and Auer, P. (2010) · 2010
Earlier work this paper cites.
A contextual-bandit approach to personalized news article recommendation
Li, L., Chu, W., Langford, J., and Schapire, R. E. (2010) · 2010
Cited alongside, same era.
Linearly parameterized bandits
Rusmevichientong, P. and Tsitsiklis, J. N. (2010) · 2010
Cited alongside, same era.
Model-based reinforcement learning with nearly tight exploration complexity bounds
Szita, I. and Szepesvári, C. (2010) · 2010
Cited alongside, same era.
Improved algorithms for linear stochastic bandits
Abbasi-Yadkori, Y., Pál, D., and Szepesvári, C. (2011) · 2011
Cited alongside, same era.
Contextual Bandits with Linear Payoff Functions
Chu, W., Li, L., Reyzin, L., and Schapire, R. E. (2011) · 2011
Cited alongside, same era.
Regret analysis of stochastic and nonstochastic multi-armed bandit problems
Bubeck, S., Cesa-Bianchi, N., et al. (2012) · 2012
Cited alongside, same era.
On lower bounds for regret in reinforcement learning
Osband, I. and Van Roy, B. (2016) · 2016
Later among the works it cites.
Safe, multi-agent, reinforcement learning for autonomous driving
Shalev-Shwartz, S., Shammah, S., and Shashua, A. (2016) · 2016
Later among the works it cites.
Optimistic posterior sampling for reinforcement learning: worst-case regret bounds
Agrawal, S. and Jia, R. (2017) · 2017
Later among the works it cites.
Minimax regret bounds for reinforcement learning
Azar, M. G., Osband, I., and Munos, R. (2017) · 2017
Later among the works it cites.
On kernelized multi-armed bandits
Chowdhury, S. R. and Gopalan, A. (2017) · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Minimax pac bounds on the sample complexity of reinforcement learning with a generative model
Azar, M. G., Munos, R., and Kappen, H. J. (2013) · 2013
Cited alongside, same era.
Reinforcement learning in robotics: A survey
Kober, J., Bagnell, J. A., and Peters, J. (2013) · 2013
Cited alongside, same era.
Playing atari with deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M. (2013) · 2013
Cited alongside, same era.
Finite-time analysis of kernelised contextual bandits
Valko, M., Korda, N., Munos, R., Flaounas, I., and Cristianini, N. (2013) · 2013
Cited alongside, same era.
Near-optimal pac bounds for discounted mdps
Lattimore, T. and Hutter, M. (2014) · 2014
Cited alongside, same era.
Sample complexity of episodic fixed-horizon reinforcement learning
Dann, C. and Brunskill, E. (2015) · 2015
Cited alongside, same era.
Osband, I., Van Roy, B., Russo, D., and Wen, Z. (2017) · 2017
Later among the works it cites.
Mastering the game of go without human knowledge
Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., Hubert, T., Baker, L., Lai, M., Bolton, A., et al. (2017) · 2017
Later among the works it cites.
Randomized sketches for kernels: Fast and optimal nonparametric regression
Yang, Y., Pilanci, M., Wainwright, M. J., et al. (2017) · 2017
Later among the works it cites.
Efficient exploration through bayesian deep q-networks
Azizzadenesheli, K., Brunskill, E., and Anandkumar, A. (2018) · 2018
Later among the works it cites.
What doubling tricks can and can’t do for multi-armed bandits
Besson, L. and Kaufmann, E. (2018) · 2018
Later among the works it cites.
Policy certificates: Towards accountable reinforcement learning
Dann, C., Li, L., Wei, W., and Brunskill, E. (2018) · 2018
Later among the works it cites.
Is q-learning provably efficient?
Jin, C., Allen-Zhu, Z., Bubeck, S., and Jordan, M. I. (2018) · 2018
Later among the works it cites.
Sidford, A., Wang, M., Wu, X., Yang, L. F., and Ye, Y. (2018) · 2018
Later among the works it cites.
Online learning in kernelized markov decision processes
Chowdhury, S. R. and Gopalan, A. (2019) · 2019
Closest in time.