Fetching the paper…
Reading the bibliography…
We investigate meta-learning procedures in the setting of stochastic linear bandits tasks.
Some aspects of the sequential design of experiments
Herbert Robbins · 1952
Earlier work this paper cites.
Least squares estimates in stochastic regression models with applications to identification and control of dynamic systems
Tze Leung Lai and Ching Zong Wei · 1982
Earlier work this paper cites.
A model of inductive bias learning
Jonathan Baxter · 2000
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
Peter Auer, Nicolò Cesa-Bianchi, and Paul Fischer · 2002
Earlier work this paper cites.
Using confidence bounds for exploitation-exploration trade-offs
Peter Auer · 2003
Earlier work this paper cites.
Herbert robbins and sequential analysis
David Siegmund · 2003
Earlier work this paper cites.
A framework for learning predictive structures from multiple tasks and unlabeled data
Rie Kubota Ando and Tong Zhang · 2005
Earlier work this paper cites.
Learning theory estimates via integral operators and their approximations
Steve Smale and Ding-Xuan Zhou · 2007
Earlier work this paper cites.
Online linear optimization and adaptive routing
Baruch Awerbuch and Robert Kleinberg · 2008
Earlier work this paper cites.
Movie recommendation using random walks over the contextual graph
Toine Bogers · 2010
Earlier work this paper cites.
Linear algorithms for online multitask classification
Giovanni Cavallanti, Nicolò Cesa-Bianchi, and Claudio Gentile · 2010
Earlier work this paper cites.
A contextual-bandit approach to personalized news article recommendation
Lihong Li, Wei Chu, John Langford, and Robert E Schapire · 2010
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Yasin Abbasi-Yadkori, Dávid Pál, and Csaba Szepesvári · 2011
Earlier work this paper cites.
2nd international workshop on information heterogeneity and fusion in recommender systems (hetrec 2011)
Iván Cantador, Peter Brusilovsky, and Tsvi Kuflik · 2011
Cited alongside, same era.
Contextual bandits with linear payoff functions
Wei Chu, Lihong Li, Lev Reyzin, and Robert Schapire · 2011
Cited alongside, same era.
Regret analysis of stochastic and nonstochastic multi-armed bandit problems
Sébastien Bubeck, Nicolo Cesa-Bianchi, et al · 2012
Cited alongside, same era.
Sequential transfer in multi-armed bandit with finite set of models
Mohammad Gheshlaghi Azar, Alessandro Lazaric, and Emma Brunskill · 2013
Cited alongside, same era.
A gang of bandits
Nicolò Cesa-Bianchi, Claudio Gentile, and Giovanni Zappella · 2013
Cited alongside, same era.
Excess risk bounds for multitask learning with trace norm regularization
Andreas Maurer and Massimiliano Pontil · 2013
The benefit of multitask representation learning
Andreas Maurer, Massimiliano Pontil, and Bernardino Romera-Paredes · 2016
Later among the works it cites.
Lifelong learning with weighted majority votes
Anastasia Pentina and Ruth Urner · 2016
Later among the works it cites.
Regret Bounds for Lifelong Learning
Pierre Alquier, The Tien Mai, and Massimiliano Pontil · 2017
Later among the works it cites.
Multi-task learning for contextual bandits
Aniket An Deshmukh, Urun Dogan, and Clay Scott · 2017
Later among the works it cites.
On context-dependent clustering of bandits
Claudio Gentile, Shuai Li, Purushottam Kar, Alexandros Karatzoglou, Giovanni Zappella, and Evans Etrue · 2017
Later among the works it cites.
Transfer learning in multi-armed bandits: A causal approach
Junzhe Zhang and Elias Bareinboim · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Sparse coding for multitask and transfer learning
Andreas Maurer, Massimiliano Pontil, and Bernardino Romera-Paredes · 2013
Cited alongside, same era.
Sparse Multi-task Reinforcement Learning
Daniele Calandriello, Alessandro Lazaric, and Marcello Restelli · 2014
Cited alongside, same era.
Online clustering of bandits
Claudio Gentile, Shuai Li, and Giovanni Zappella · 2014
Cited alongside, same era.
Multi-task linear bandits
Marta Soare, Ouais Alsharif, Alessandro Lazaric, and Joelle Pineau · 2014
Cited alongside, same era.
The movielens datasets: History and context
F. Maxwell Harper and Joseph A. Konstan · 2015
Cited alongside, same era.
Multi-armed bandit models for the optimal design of clinical trials: benefits and challenges
Sofía S Villar, Jack Bowden, and James Wason · 2015
Cited alongside, same era.
Later among the works it cites.
Transferable contextual bandit for cross-domain recommendation
B. Liu, Y. Wei, Zhang Y., Z. Yan, and Q. Yang · 2018
Later among the works it cites.
Provable guarantees for gradient-based meta-learning
Maria-Florina Balcan, Mikhail Khodak, and Ameet Talwalkar · 2019
Later among the works it cites.
Stochastic bandits with delay-dependent payoffs
Leonardo Cella and Nicolò Cesa-Bianchi · 2019
Later among the works it cites.
Learning-to-learn stochastic gradient descent with biased regularization
Giulia Denevi, Carlo Ciliberto, Riccardo Grazzi, and Massimiliano Pontil · 2019
Later among the works it cites.
Efficient linear bandits through matrix sketching
Ilja Kuzborskij, Leonardo Cella, and Nicolò Cesa-Bianchi · 2019
Later among the works it cites.
Bandit algorithms
Tor Lattimore and Csaba Szepesvári · 2020
Closest in time.