Fetching the paper…
Reading the bibliography…
We consider a team of reinforcement learning agents that concurrently learn to operate in a common environment.
A Bayesian framework for reinforcement learning
Malcolm J. A. Strens · 2000
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Michael J. Kearns and Satinder P. Singh · 2002
Earlier work this paper cites.
Computational Statistics
James E. Gentle · 2009
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Thomas Jaksch, Ronald Ortner, and Peter Auer · 2010
Earlier work this paper cites.
Random Number Generation and Monte Carlo Methods
James E. Gentle · 2013
Earlier work this paper cites.
(More) efficient reinforcement learning via posterior sampling
Ian Osband, Daniel Russo, and Benjamin Van Roy · 2013
Cited alongside, same era.
PAC optimal exploration in continuous space Markov decision processes
Jason Pazis and Ronald Parr · 2013
Cited alongside, same era.
Concurrent reinforcement learning from customer interactions
D. Silver, Barker Newnham, L, S. Weller, and J. McFall · 2013
Cited alongside, same era.
Concurrent PAC RL
Z. Guo and E. Brunskill · 2015
Cited alongside, same era.
Model-based reinforcement learning and the eluder dimension
Ian Osband and Benjamin Van Roy
Cited in the paper.
Near-optimal reinforcement learning in factored MDPs
Ian Osband and Benjamin Van Roy
Cited in the paper.
On optimistic versus randomized exploration in reinforcement learning
Ian Osband and Benjamin Van Roy
Cited in the paper.
Why is posterior sampling better than optimism for reinforcement learning
Ian Osband and Benjamin Van Roy
Cited in the paper.
Efficient pac-optimal exploration in concurrent, continuous state mdps with delayed updates
Jason Pazis and Ronald Parr · 2016
Later among the works it cites.
Thompson sampling for stochastic control: The finite parameter case
Michael Jong Kim · 2017
Later among the works it cites.
Ensemble sampling
Xiuyuan Lu and Benjamin Van Roy · 2017
Later among the works it cites.
A tutorial on Thompson sampling
Daniel Russo, Benjamin Van Roy, Abbas Kazerouni, Ian Osband, and Zheng Wen · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…