Fetching the paper…
Reading the bibliography…
We consider the recently proposed reinforcement learning (RL) framework of Contextual Markov Decision Processes (CMDP), where the agent interacts with a (potentially adversarial) sequence of episodic tabular MDPs.
Reinforcement leaning in feature space: Matrix bandit, kernels, and regret bound
Yang, L. F. and Wang, M. (2019b) · 1905
Earlier work this paper cites.
Provably efficient reinforcement learning with linear function approximation
Jin, C., Yang, Z., Wang, Z., and Jordan, M. I. (2019) · 1907
Earlier work this paper cites.
Graphical models
Lauritzen, S. L. (1996) · 1996
Earlier work this paper cites.
Using confidence bounds for exploitation-exploration trade-offs
Auer, P. (2002) · 2002
Earlier work this paper cites.
Reinforcement learning with immediate rewards and linear hypotheses
Abe, N., Biermann, A. W., and Long, P. M. (2003) · 2003
Earlier work this paper cites.
Logarithmic online regret bounds for undiscounted reinforcement learning
Auer, P. and Ortner, R. (2007) · 2007
Earlier work this paper cites.
Logarithmic regret algorithms for online convex optimization
Hazan, E., Agarwal, A., and Kale, S. (2007) · 2007
Earlier work this paper cites.
Parametric bandits: The generalized linear case
Filippi, S., Cappe, O., Garivier, A., and Szepesvári, C. (2010) · 2010
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Jaksch, T., Ortner, R., and Auer, P. (2010) · 2010
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Abbasi-Yadkori, Y., Pál, D., and Szepesvári, C. (2011) · 2011
Earlier work this paper cites.
Contextual bandits with linear payoff functions
Chu, W., Li, L., Reyzin, L., and Schapire, R. (2011) · 2011
Cited alongside, same era.
Online-to-confidence-set conversions and application to sparse stochastic bandits
Abbasi-Yadkori, Y., Pal, D., and Szepesvari, C. (2012) · 2012
Cited alongside, same era.
Selective sampling algorithms for cost-sensitive multiclass prediction
Agarwal, A. (2013) · 2013
Cited alongside, same era.
(more) efficient reinforcement learning via posterior sampling
Osband, I., Russo, D., and Van Roy, B. (2013) · 2013
Cited alongside, same era.
Online learning in MDPs with side information
Abbasi-Yadkori, Y. and Neu, G. (2014) · 2014
Cited alongside, same era.
Minimax regret bounds for reinforcement learning
Azar, M. G., Osband, I., and Munos, R. (2017) · 2017
Later among the works it cites.
Scalable generalized linear bandits: Online computation and hashing
Jun, K.-S., Bhargava, A., Nowak, R., and Willett, R. (2017) · 2017
Later among the works it cites.
Logistic regression: The importance of being improper
Foster, D. J., Kale, S., Luo, H., Mohri, M., and Sridharan, K. (2018) · 2018
Later among the works it cites.
Markov decision processes with continuous side information
Modi, A., Jiang, N., Singh, S., and Tewari, A. (2018) · 2018
Later among the works it cites.
A tutorial on thompson sampling
Russo, D. J., Van Roy, B., Kazerouni, A., Osband, I., Wen, Z., et al. (2018) · 2018
Later among the works it cites.
Policy certificates: Towards accountable reinforcement learning
Dann, C., Li, L., Wei, W., and Brunskill, E. (2019) · 2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hallak, A., Di Castro, D., and Mannor, S. (2015) · 2015
Cited alongside, same era.
On lower bounds for regret in reinforcement learning
Osband, I. and Van Roy, B. (2016) · 2016
Cited alongside, same era.
Generalization and exploration via randomized value functions
Osband, I., Van Roy, B., and Wen, Z. (2016) · 2016
Cited alongside, same era.
Online stochastic linear optimization under one-bit feedback
Zhang, L., Yang, T., Jin, R., Xiao, Y., and Zhou, Z.-h. (2016) · 2016
Cited alongside, same era.
Sample-optimal parametric q-learning using linearly additive features
Yang, L. and Wang, M. (2019a)
Cited in the paper.
Worst-case regret bounds for exploration via randomized value functions
Russo, D. (2019) · 2019
Closest in time.
Bandit Algorithms
Lattimore, T. and Szepesvári, C. (2020) · 2020
Closest in time.
Provably optimal algorithms for generalized linear contextual bandits
Li, L., Lu, Y., and Zhou, D. (2017) · 2080
Closest in time.