Fetching the paper…
Reading the bibliography…
We study online learning of finite Markov decision process (MDP) problems when a side information vector is available.
Asymptotically efficient adaptive allocation rules
Tze Leung Lai and Herbert Robbins · 1985
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Richard S. Sutton · 1988
Earlier work this paper cites.
An analysis of temporal-difference learning with function approximation
John N. Tsitsiklis and Benjamin Van Roy · 1997
Earlier work this paper cites.
A design for testing clinical strategies: biased individually tailored within-subject randomization
P. W. Lavori and R. Dawson · 2000
Earlier work this paper cites.
Marginal mean models for dynamic regimes
S. A. Murphy, M. J. van der Laan, and J. M. Robins · 2001
Earlier work this paper cites.
Using confidence bounds for exploitation-exploration trade-offs
Peter Auer · 2002
Earlier work this paper cites.
Finite time analysis of the multiarmed bandit problem
Peter Auer, Nicolò Cesa-Bianchi, and Paul Fischer · 2002
Cited alongside, same era.
The on-line shortest path problem under partial monitoring
András György, Tamás Linder, Gábor Lugosi, and Gy · 2007
Cited alongside, same era.
Stochastic linear optimization under bandit feedback
Varsha Dani, Thomas P. Hayes, and Sham M. Kakade · 2008
Cited alongside, same era.
Convergent temporal-difference learning with arbitrary smooth function approximation
H. R. Maei, Cs. Szepesvári, S. Bhatnagar, D. Precup, D. Silver, and R. S. Sutton · 2009
Cited alongside, same era.
Parametric bandits: The generalized linear case
Sarah Filippi, Olivier Cappé, Aurélien Garivier, and Csaba Szepesvári · 2010
Cited alongside, same era.
Near-optimal regret bounds for reinforcement learning
Thomas Jaksch, Ronald Ortner, and Peter Auer · 2010
A contextual-bandit approach to personalized news article recommendation
Lihong Li, Wei Chu, John Langford, and Robert E. Schapire · 2010
Later among the works it cites.
Toward off-policy learning control with function approximation
H. R. Maei, Cs. Szepesvári, S. Bhatnagar, and R. S. Sutton · 2010
Later among the works it cites.
Regret bounds for the adaptive control of linear quadratic systems
Y. Abbasi-Yadkori and Cs. Szepesvári · 2011
Later among the works it cites.
Online Learning for Linearly Parametrized Control Problems
Y. Abbasi-Yadkori · 2012
Later among the works it cites.
The adversarial stochastic shortest path problem with unknown transition probabilities
Gergely Neu, András György, and Csaba Szepesvári · 2012
Later among the works it cites.
Online regret bounds for undiscounted continuous reinforcement learning
R. Ortner and D. Ryabko · 2012
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Fast gradient-descent methods for temporal-difference learning with linear function approximation
R. S. Sutton, H. R. Maei, D. Precup, S. Bhatnagar, D. Silver, Cs. Szepesvári, and E. Wiewiora
Cited in the paper.
A convergent O(n) algorithm for off-policy temporal-difference learning with linear function approximation
R. S. Sutton, Cs. Szepesvári, and H. R. Maei
Cited in the paper.
Online learning in Markov decision processes with arbitrarily changing rewards and transitions
Jia Yuan Yu and Shie Mannor
Cited in the paper.
Arbitrarily modulated Markov decision processes
Jia Yuan Yu and Shie Mannor
Cited in the paper.
Later among the works it cites.