Improved algorithms for linear stochastic bandits
Abbasi-Yadkori, Y.; Pál, D.; and Szepesvári, C. 2011 · 2011
Cited alongside, same era.
Contextual bandit algorithms with supervised learning guarantees
Beygelzimer, A.; Langford, J.; Li, L.; Reyzin, L.; and Schapire, R. 2011 · 2011
Cited alongside, same era.
Contextual bandits with linear payoff functions
Chu, W.; Li, L.; Reyzin, L.; and Schapire, R. 2011 · 2011
Cited alongside, same era.
Thompson sampling for contextual bandits with linear payoffs
Agrawal, S.; and Goyal, N. 2013 · 2013
Cited alongside, same era.
(More) efficient reinforcement learning via posterior sampling
Osband, I.; Russo, D.; and Van Roy, B. 2013 · 2013
Cited alongside, same era.
Online learning in episodic Markovian decision processes by relative entropy policy search
Zimin, A.; and Neu, G. 2013 · 2013
Cited alongside, same era.
Sample complexity of episodic fixed-horizon reinforcement learning
Dann, C.; and Brunskill, E. 2015 · 2015
Cited alongside, same era.
Thompson sampling for learning parameterized Markov decision processes
Gopalan, A.; and Mannor, S. 2015 · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, V.; Kavukcuoglu, K.; Silver, D.; Rusu, A. A.; Veness, J.; Bellemare, M.; Graves, A.; Riedmiller, M.; Fidjeland, A.; Ostrovski, G.; et al. 2015 · 2015
Cited alongside, same era.
End-to-end training of deep visuomotor policies
Levine, S.; Finn, C.; Darrell, T.; and Abbeel, P. 2016 · 2016
Cited alongside, same era.
On lower bounds for regret in reinforcement learning
Original
Osband, I.; and Van Roy, B. 2016 · 2016
Cited alongside, same era.
Linear thompson sampling revisited
Abeille, M.; Lazaric, A.; et al. 2017 · 2017
Cited alongside, same era.