Fetching the paper…
Reading the bibliography…
Ensemble sampling serves as a practical approximation to Thompson sampling when maintaining an exact posterior distribution over model parameters is computationally intractable.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
William R. Thompson · 1933
Earlier work this paper cites.
Explaining the Gibbs sampler
George Casella and Edward I George · 1992
Earlier work this paper cites.
Exponential convergence of Langevin distributions and their discrete approximations
Gareth O. Roberts and Richard L. Tweedie · 1996
Earlier work this paper cites.
Elements of Information Theory
Thomas M. Cover and Joy A. Thomas · 2006
Earlier work this paper cites.
Introduction to Nonparametric Estimation
Alexandre B. Tsybakov · 2009
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Yasin Abbasi-Yadkori, Dávid Pál, and Csaba Szepesvári · 2011
Earlier work this paper cites.
An empirical evaluation of Thompson sampling
Olivier Chapelle and Lihong Li · 2011
Earlier work this paper cites.
Tight bounds for the expected risk of linear classifiers and PAC-Bayes finite-sample guarantees
Jean Honorio and Tommi Jaakkola · 2014
Earlier work this paper cites.
Deep exploration via bootstrapped DQN
Ian Osband, Charles Blundell, Alexander Pritzel, and Benjamin Van Roy · 2016
Earlier work this paper cites.
An information-theoretic analysis of Thompson sampling
Daniel Russo and Benjamin Van Roy · 2016
Earlier work this paper cites.
Deep neural networks as gaussian processes
Jaehoon Lee, Yasaman Bahri, Roman Novak, Samuel S Schoenholz, Jeffrey Pennington, and Jascha Sohl-Dickstein · 2017
Earlier work this paper cites.
Ensemble sampling
Xiuyuan Lu and Benjamin Van Roy · 2017
Earlier work this paper cites.
Coordinated exploration in concurrent reinforcement learning
Maria Dimakopoulou and Benjamin Van Roy · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Cited alongside, same era.
Efficient online recommendation via low-rank ensemble sampling
Xiuyuan Lu, Zheng Wen, and Branislav Kveton · 2018
Cited alongside, same era.
Randomized prior functions for deep reinforcement learning
Ian Osband, John Aslanides, and Albin Cassirer · 2018
Cited alongside, same era.
A tutorial on Thompson sampling
Daniel J. Russo, Benjamin Van Roy, Abbas Kazerouni, Ian Osband, and Zheng Wen · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
Richard S. Sutton and Andrew G. Barto · 2018
Cited alongside, same era.
Hypermodels for exploration
Vikranth Dwaracherla, Xiuyuan Lu, Morteza Ibrahimi, Ian Osband, Zheng Wen, and Benjamin Van Roy · 2020
Later among the works it cites.
Botao Hao, Jie Zhou, Zheng Wen, and Will Wei Sun · 2020
Later among the works it cites.
Perturbed-history exploration in stochastic linear bandits
Branislav Kveton, Csaba Szepesvári, Mohammad Ghavamzadeh, and Craig Boutilier · 2020
Later among the works it cites.
Bandit algorithms
Tor Lattimore and Csaba Szepesvári · 2020
Later among the works it cites.
Targeting for long-term outcomes
Jeremy Yang, Dean Eckles, Paramveer Dhillon, and Sinan Aral · 2020
Later among the works it cites.
Graphical models meet bandits: A variational Thompson sampling approach
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Variational inference for the multi-armed contextual bandit
Iñigo Urteaga and Chris Wiggins · 2018
Cited alongside, same era.
Bootstrap Thompson sampling and sequential decision problems in the behavioral sciences
Dean Eckles and Maurits Kaptein · 2019
Cited alongside, same era.
Garbage in, reward out: Bootstrapping exploration in multi-armed bandits
Branislav Kveton, Csaba Szepesvári, Sharan Vaswani, Zheng Wen, Tor Lattimore, and Mohammad Ghavamzadeh · 2019
Cited alongside, same era.
Deep exploration via randomized value functions
Ian Osband, Benjamin Van Roy, Daniel J. Russo, and Zheng Wen · 2019
Cited alongside, same era.
Thompson sampling and approximate inference
My Phan, Yasin Abbasi-Yadkori, and Justin Domke · 2019
Cited alongside, same era.
Scalable Thompson sampling via optimal transport
Ruiyi Zhang, Zheng Wen, Changyou Chen, Chen Fang, Tong Yu, and Lawrence Carin · 2019
Cited alongside, same era.
Tong Yu, Branislav Kveton, Zheng Wen, Ruiyi Zhang, and Ole J. Mengshoel · 2020
Later among the works it cites.
Anti-concentrated confidence bonuses for scalable exploration
Jordan T. Ash, Cyril Zhang, Surbhi Goel, Akshay Krishnamurthy, and Sham Kakade · 2021
Later among the works it cites.
Deep exploration for recommendation systems
Zheqing Zhu and Benjamin Van Roy · 2021
Later among the works it cites.
Generalized Bayesian upper confidence bound with approximate inference for bandit problems
Ziyi Huang, Henry Lam, Amirhossein Meisami, and Haofeng Zhang · 2022
Closest in time.
The neural testbed: Evaluating joint predictions
Ian Osband, Zheng Wen, Seyed Mohammad Asghari, Vikranth Dwaracherla, Xiuyuan Lu, Morteza Ibrahimi, Dieterich Lawson, Botao Hao, Brendan O’Donoghue, and Benjamin Van Roy · 2022
Closest in time.
From predictions to decisions: The importance of joint predictive distributions
Zheng Wen, Ian Osband, Chao Qin, Xiuyuan Lu, Morteza Ibrahimi, Vikranth Dwaracherla, Mohammad Asghari, and Benjamin Van Roy · 2022
Closest in time.