Fetching the paper…
Reading the bibliography…
Thompson sampling and other Bayesian sequential decision-making algorithms are among the most popular approaches to tackle explore/exploit trade-offs in (contextual) bandits.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
William R Thompson · 1933
Earlier work this paper cites.
Fundamentals of Statistical Exponential Families with Applications in Statistical Decision Theory
Lawrence D Brown · 1986
Earlier work this paper cites.
An overview of robust bayesian analysis
James O Berger, Elías Moreno, Luis Raul Pericchi, M Jesús Bayarri, José M Bernardo, Juan A Cano, Julián De la Horra, Jacinto Martín, David Ríos-Insúa, Bruno Betrò, A. Dasgupta, Paul Gustafson, Larry Wasserman, Joseph B. Kadane, Cid Srinivasan, Michael Lavine, Anthony O’Hagan, Wolfgang Polasek, Christian P. Robert, Constantinos Goutis, Fabrizio Ruggeri, Gabriella Salinetti, and Siva Sivaganesan · 1994
Earlier work this paper cites.
Estimation of parameters in the beta binomial model
Ram C. Tripathi, Ramesh C. Gupta, and John Gurland · 1994
Earlier work this paper cites.
Explanation-based neural network learning: A lifelong learning approach
Sebastian Thrun · 1996
Earlier work this paper cites.
Theoretical models of learning to learn
Jonathan Baxter · 1998
Earlier work this paper cites.
Lifelong learning algorithms
Sebastian Thrun · 1998
Earlier work this paper cites.
A model of inductive bias learning
Jonathan Baxter · 2000
Earlier work this paper cites.
A bayesian framework for reinforcement learning
Malcolm Strens · 2000
Earlier work this paper cites.
Learning to learn using gradient descent
Sepp Hochreiter, A Steven Younger, and Peter R Conwell · 2001
Earlier work this paper cites.
Lectures on the Coupling Method
Torgny Lindvall · 2002
Earlier work this paper cites.
On the sample complexity of reinforcement learning
Sham Machandranath Kakade · 2003
Earlier work this paper cites.
Clustering with bregman divergences
Arindam Banerjee, Srujana Merugu, Inderjit S Dhillon, Joydeep Ghosh, and John Lafferty · 2005
Cited alongside, same era.
Stochastic dominance and applications to finance, risk and economics
Songsak Sriboonchita, Wing-Keung Wong, Sompong Dhompongsa, and Hung T Nguyen · 2009
Cited alongside, same era.
Analysis of thompson sampling for the multi-armed bandit problem
Shipra Agrawal and Navin Goyal · 2012
Cited alongside, same era.
Thompson sampling: An asymptotically optimal finite-time analysis
Emilie Kaufmann, Nathaniel Korda, and Rémi Munos · 2012
Cited alongside, same era.
The knowledge gradient algorithm for a general class of online learning problems
Ilya O Ryzhov, Warren B Powell, and Peter I Frazier · 2012
Cited alongside, same era.
Sequential transfer in multi-armed bandit with finite set of models
Mohammad Gheshlaghi Azar, Alessandro Lazaric, and Emma Brunskill · 2013
Linear thompson sampling revisited
Marc Abeille and Alessandro Lazaric · 2017
Later among the works it cites.
Why is posterior sampling better than optimism for reinforcement learning?
Ian Osband and Benjamin Van Roy · 2017
Later among the works it cites.
Learning to reinforcement learn
Jane X Wang, Zeb Kurth-Nelson, Dhruva Tirumala, Hubert Soyer, Joel Z Leibo, Remi Munos, Charles Blundell, Dharshan Kumaran, and Matt Botvinick · 2017
Later among the works it cites.
Improved regret bounds for thompson sampling in linear quadratic control problems
Marc Abeille and Alessandro Lazaric · 2018
Later among the works it cites.
High-dimensional probability: An introduction with applications in data science
Roman Vershynin · 2018
Later among the works it cites.
Thompson sampling and approximate inference
My Phan, Yasin Abbasi Yadkori, and Justin Domke · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Thompson sampling for complex online problems
Aditya Gopalan, Shie Mannor, and Yishay Mansour · 2014
Cited alongside, same era.
Bayesian optimal control of smoothly parameterized systems
Yasin Abbasi-Yadkori and Csaba Szepesvári · 2015
Cited alongside, same era.
Bayesian reinforcement learning: A survey
Mohammad Ghavamzadeh, Shie Mannor, Joelle Pineau, and Aviv Tamar · 2015
Cited alongside, same era.
On the prior sensitivity of thompson sampling
Che-Yu Liu and Lihong Li · 2016
Cited alongside, same era.
Simple bayesian algorithms for best arm identification
Daniel Russo · 2016
Cited alongside, same era.
An information-theoretic analysis of thompson sampling
Daniel Russo and Benjamin Van Roy · 2016
Cited alongside, same era.
Meta-learning with stochastic linear bandits
Leonardo Cella, Alessandro Lazaric, and Massimiliano Pontil · 2020
Later among the works it cites.
Meta reinforcement learning as task inference
Jan Humplik, Alexandre Galashov, Leonard Hasenclever, Pedro A Ortega, Yee Whye Teh, and Nicolas Heess · 2020
Later among the works it cites.
Near-optimal representation learning for linear bandits and linear RL
Jiachen Hu, Xiaoyu Chen, Chi Jin, Lihong Li, and Liwei Wang · 2021
Closest in time.
Branislav Kveton, Mikhail Konobeev, Manzil Zaheer, Chih-wei Hsu, Martin Mladenov, Craig Boutilier, and Csaba Szepesvari · 2021
Closest in time.
Impact of representation learning in linear bandits
Jiaqi Yang, Wei Hu, Jason D Lee, and Simon S Du · 2021
Closest in time.