Fetching the paper…
Reading the bibliography…
We study a sequential decision problem where the learner faces a sequence of $K$-armed bandit tasks.
The weighted majority algorithm
N. Littlestone and M.K. Warmuth · 1994
Earlier work this paper cites.
Gambling in a rigged casino: The adversarial multi-armed bandit problem
P. Auer, N. Cesa-Bianchi, Y. Freund, and R. E. Schapire · 1995
Earlier work this paper cites.
Is learning the n-th thing any easier than learning the first?
Sebastian Thrun · 1996
Earlier work this paper cites.
Bandit problems with infinitely many arms
Donald A. Berry, Robert W. Chen, Alan Zame, David C. Heath, and Larry A. Shepp · 1997
Earlier work this paper cites.
Multitask learning
Rich Caruana · 1997
Earlier work this paper cites.
A model of inductive bias learning
Jonathan Baxter · 2000
Earlier work this paper cites.
The nonstochastic multiarmed bandit problem
P. Auer, N. Cesa-Bianchi, Y. Freund, and R. E. Schapire · 2002
Earlier work this paper cites.
Tracking a small set of experts by mixing past posteriors
Olivier Bousquet and Manfred K. Warmuth · 2002
Earlier work this paper cites.
Approximating min sum set cover
Uriel Feige, László Lovász, and Prasad Tetali · 2004
Earlier work this paper cites.
Prediction, learning, and games
Nicolo Cesa-Bianchi and Gábor Lugosi · 2006
Earlier work this paper cites.
An online algorithm for maximizing submodular functions
Matthew Streeter and Daniel Golovin · 2007
Earlier work this paper cites.
Learning diverse rankings with multi-armed bandits
Filip Radlinski, Robert Kleinberg, and Thorsten Joachims · 2008
Earlier work this paper cites.
Algorithms for infinitely many-armed bandits
Yizao Wang, Jean-Yves Audibert, and Rémi Munos · 2008
Earlier work this paper cites.
Minimax policies for adversarial and stochastic bandits
Jean-Yves Audibert and Sébastien Bubeck · 2009
Earlier work this paper cites.
Ucb revisited: Improved regret bounds for the stochastic multi-armed bandit problem
P. Auer and R. Ortner · 2010
Earlier work this paper cites.
Non-stochastic bandit slate problems
S. Kale, L. Reyzin, and R. E. Schapire · 2010
Cited alongside, same era.
Sequential transfer in multi-armed bandit with finite set of models
Mohammad Gheshlaghi Azar, Alessandro Lazaric, and Emma Brunskill · 2013
Cited alongside, same era.
Two-target algorithms for infinite-armed bandits with Bernoulli rewards
T. Bonald and A. Proutiere · 2013
Cited alongside, same era.
Online clustering of bandits
Claudio Gentile, Shuai Li, and Giovanni Zappella · 2014
Cited alongside, same era.
Latent bandits
Odalric-Ambrym Maillard and Shie Mannor · 2014
Cited alongside, same era.
Simple regret for infinitely many armed bandits
Alexandra Carpentier and Michal Valko · 2015
Cited alongside, same era.
Meta-learning with stochastic linear bandits
Leonardo Cella, Alessandro Lazaric, and Massimiliano Pontil · 2020
Later among the works it cites.
Infinite arms bandit: Optimality via confidence bounds, 2020
Hock Peng Chan and Shouri Hu · 2020
Later among the works it cites.
From finite to countable-armed bandits
Anand Kalvit and Assaf Zeevi · 2020
Later among the works it cites.
Meta-learning for mixed linear regression
Weihao Kong, Raghav Somani, Zhao Song, Sham Kakade, and Sewoong Oh · 2020
Later among the works it cites.
Differentiable meta-learning in contextual bandits
Branislav Kveton, Martin Mladenov, Chih-Wei Hsu, Manzil Zaheer, Csaba Szepesvári, and Craig Boutilier · 2020
Later among the works it cites.
Bandit Algorithms
T. Lattimore and C. Szepesvári · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Joon Kwon, Vianney Perchet, and Claire Vernade · 2017
Cited alongside, same era.
Bilevel programming for hyperparameter optimization and meta-learning, 2018
Luca Franceschi, Paolo Frasconi, Saverio Salzo, Riccardo Grazzi, and Massimilano Pontil · 2018
Cited alongside, same era.
A new algorithm for non-stationary contextual bandits: Efficient, optimal and parameter-free
Yifang Chen, Chung-Wei Lee, Haipeng Luo, and Chen-Yu Wei · 2019
Cited alongside, same era.
Learning-to-learn stochastic gradient descent with biased regularization
Giulia Denevi, Carlo Ciliberto, Riccardo Grazzi, and Massimiliano Pontil · 2019
Cited alongside, same era.
Marginal posterior sampling for slate bandits
M. Dimakopoulou, N. Vlassis, and T. Jebara · 2019
Cited alongside, same era.
Provable guarantees for gradient-based meta-learning
Mikhail Khodak, Maria-Florina Balcan, and Ameet Talwalkar · 2019
Cited alongside, same era.
Bandit Algorithms
Tor Lattimore and Csaba Szepesvari · 2020
Later among the works it cites.
Algorithms for slate bandits with non-separable reward functions
Jason Rhuggenaath, Alp Akcay, Yingqian Zhang, and Uzay Kayma · 2020
Later among the works it cites.
Meta-thompson sampling
Branislav Kveton, Mikhail Konobeev, Manzil Zaheer, Chih wei Hsu, Martin Mladenov, Craig Boutilier, and Csaba Szepesvari · 2021
Later among the works it cites.
Transfer learning in bandits with latent continuity
Hyejin Park, Seiyun Shin, Kwang-Sung Jun, and Jungseul Ok · 2021
Later among the works it cites.
Provable meta-learning of linear representations
Nilesh Tripuraneni, Chi Jin, and Michael I. Jordan · 2021
Later among the works it cites.
Non-stationary reinforcement learning without prior knowledge: An optimal black-box approach
Chen-Yu Wei and Haipeng Luo · 2021
Later among the works it cites.
Meta-learning adversarial bandits
Maria-Florina Balcan, Keegan Harris, Mikhail Khodak, and Zhiwei Steven Wu · 2022
Closest in time.
Online meta-learning in adversarial multi-armed bandits
Ilya Osadchiy, Kfir Y Levy, and Ron Meir · 2022
Closest in time.