Fetching the paper…
Reading the bibliography…
We propose ${\tt AdaTS}$, a Thompson sampling algorithm that adapts sequentially to bandit tasks that it interacts with.
Meta dynamic pricing: Transfer learning across experiments
Hamsa Bastani, David Simchi-Levi, and Ruihao Zhu · 1902
Earlier work this paper cites.
Empirical Bayes regret minimization
Chih-Wei Hsu, Branislav Kveton, Ofer Meshi, Martin Mladenov, and Csaba Szepesvari · 1904
Earlier work this paper cites.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
William R. Thompson · 1933
Earlier work this paper cites.
Asymptotically efficient adaptive allocation rules
T. L. Lai and Herbert Robbins · 1985
Earlier work this paper cites.
Explanation-Based Neural Network Learning - A Lifelong Learning Approach
Sebastian Thrun · 1996
Earlier work this paper cites.
Theoretical models of learning to learn
Jonathan Baxter · 1998
Earlier work this paper cites.
Lifelong learning algorithms
Sebastian Thrun · 1998
Earlier work this paper cites.
A model of inductive bias learning
Jonathan Baxter · 2000
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
Peter Auer, Nicolo Cesa-Bianchi, and Paul Fischer · 2002
Earlier work this paper cites.
Multi-armed bandit algorithms and empirical evaluation
Joannes Vermorel and Mehryar Mohri · 2005
Earlier work this paper cites.
Differentiable meta-learning in contextual bandits
Branislav Kveton, Martin Mladenov, Chih-Wei Hsu, Manzil Zaheer, Csaba Szepesvari, and Craig Boutilier · 2006
Earlier work this paper cites.
Policy gradient optimization of Thompson sampling policies
Seungki Min, Ciamac Moallemi, and Daniel Russo · 2006
Earlier work this paper cites.
Differentiable linear bandit algorithm
Kaige Yang and Laura Toni · 2006
Earlier work this paper cites.
Stochastic linear optimization under bandit feedback
Varsha Dani, Thomas Hayes, and Sham Kakade · 2008
Earlier work this paper cites.
Mnist handwritten digit database
Yann LeCun, Corinna Cortes, and CJ Burges · 2010
Earlier work this paper cites.
Provable benefits of representation learning in linear bandits
Jiaqi Yang, Wei Hu, Jason Lee, and Simon Du · 2010
Cited alongside, same era.
Improved algorithms for linear stochastic bandits
Yasin Abbasi-Yadkori, David Pal, and Csaba Szepesvari · 2011
Cited alongside, same era.
Analysis of Thompson sampling for the multi-armed bandit problem
Shipra Agrawal and Navin Goyal · 2012
Cited alongside, same era.
An empirical evaluation of Thompson sampling
Olivier Chapelle and Lihong Li · 2012
Cited alongside, same era.
Combinatorial network optimization with unknown variables: Multi-armed bandits with linear rewards and individual observations
Yi Gai, Bhaskar Krishnamachari, and Rahul Jain · 2012
Cited alongside, same era.
Meta-learning of exploration/exploitation strategies: The multi-armed bandit case
Tight regret bounds for stochastic combinatorial semi-bandits
Branislav Kveton, Zheng Wen, Azin Ashkan, and Csaba Szepesvari · 2015
Later among the works it cites.
Human-level concept learning through probabilistic program induction
Brenden M Lake, Ruslan Salakhutdinov, and Joshua B Tenenbaum · 2015
Later among the works it cites.
Efficient learning in large-scale combinatorial semi-bandits
Zheng Wen, Branislav Kveton, and Azin Ashkan · 2015
Later among the works it cites.
RL 2 : Fast reinforcement learning via slow reinforcement learning
Yan Duan, John Schulman, Xi Chen, Peter Bartlett, Ilya Sutskever, and Pieter Abbeel · 2016
Later among the works it cites.
Multi-task learning for contextual bandits
Aniket Anand Deshmukh, Urun Dogan, and Clayton Scott · 2017
Later among the works it cites.
Model-agnostic meta-learning for fast adaptation of deep networks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Francis Maes, Louis Wehenkel, and Damien Ernst · 2012
Cited alongside, same era.
Thompson sampling for contextual bandits with linear payoffs
Shipra Agrawal and Navin Goyal · 2013
Cited alongside, same era.
Sequential transfer in multi-armed bandit with finite set of models
Mohammad Gheshlaghi Azar, Alessandro Lazaric, and Emma Brunskill · 2013
Cited alongside, same era.
Combinatorial multi-armed bandit: General framework, results and applications
Wei Chen, Yajun Wang, and Yang Yuan · 2013
Cited alongside, same era.
Bayesian Data Analysis
Andrew Gelman, John Carlin, Hal Stern, David Dunson, Aki Vehtari, and Donald Rubin · 2013
Cited alongside, same era.
Combinatorial multi-armed bandit and its extension to probabilistically triggered arms
Wei Chen, Yajun Wang, and Yang Yuan · 2014
Cited alongside, same era.
Online clustering of bandits
Claudio Gentile, Shuai Li, and Giovanni Zappella · 2014
Cited alongside, same era.
Chelsea Finn, Pieter Abbeel, and Sergey Levine · 2017
Later among the works it cites.
Probabilistic model-agnostic meta-learning
Chelsea Finn, Kelvin Xu, and Sergey Levine · 2018
Later among the works it cites.
A tutorial on Thompson sampling
Daniel Russo, Benjamin Van Roy, Abbas Kazerouni, Ian Osband, and Zheng Wen · 2018
Later among the works it cites.
Bandit Algorithms
Tor Lattimore and Csaba Szepesvari · 2019
Later among the works it cites.
Information-theoretic confidence bounds for reinforcement learning
Xiuyuan Lu and Benjamin Van Roy · 2019
Later among the works it cites.
Differentiable meta-learning of bandit policies
Craig Boutilier, Chih-Wei Hsu, Branislav Kveton, Martin Mladenov, Csaba Szepesvari, and Manzil Zaheer · 2020
Later among the works it cites.
Meta-learning with stochastic linear bandits
Leonardo Cella, Alessandro Lazaric, and Massimiliano Pontil · 2020
Later among the works it cites.
Latent bandits revisited
Joey Hong, Branislav Kveton, Manzil Zaheer, Yinlam Chow, Amr Ahmed, and Craig Boutilier · 2020
Later among the works it cites.
Meta-Thompson sampling
Branislav Kveton, Mikhail Konobeev, Manzil Zaheer, Chih-Wei Hsu, Martin Mladenov, Craig Boutilier, and Csaba Szepesvari · 2021
Closest in time.