Fetching the paper…
Reading the bibliography…
We study online learning with bandit feedback across multiple tasks, with the goal of improving average performance across tasks if they are similar according to some natural task-similarity measure.
Possible generalization of Boltzmann-Gibbs statistics
Constantino Tsallis · 1988
Earlier work this paper cites.
Learning to Learn
Sebastian Thrun and Lorien Pratt · 1998
Earlier work this paper cites.
The nonstochastic multiarmed bandit problem
Peter Auer, Nicolò Cesa-Bianchi, Yoav Freund, and Robert E. Schapire · 2002
Earlier work this paper cites.
Some properties of q-logarithm and q-exponential functions in Tsallis statistics
Takuya Yamano · 2002
Earlier work this paper cites.
Mirror descent and nonlinear projected subgradient methods for convex optimization
Amir Beck and Marc Teboulle · 2003
Earlier work this paper cites.
Path kernels and multiplicative updates
Eiji Takimoto and Manfred K. Warmuth · 2003
Earlier work this paper cites.
Adaptive routing with end-to-end feedback: Distributed learning and geometric approaches
Baruch Awerbuch and Robert D. Kleinberg · 2004
Earlier work this paper cites.
Efficient algorithms for online decision problems
Adam Kalai and Santosh Vempala · 2005
Earlier work this paper cites.
Prediction, Learning, and Games
Nicolò Cesa-Bianchi and Gábor Lugosi · 2006
Earlier work this paper cites.
Competing in the dark: An efficient algorithm for bandit linear optimization
Jacob Abernethy, Elad Hazan, and Alexander Rakhlin · 2008
Earlier work this paper cites.
The price of bandit information for online optimization
Varsha Dani, Thomas Hayes, and Sham Kakade · 2008
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Yasin Abbasi-Yadkori, Dávid Pál, and Csaba Szepesvári · 2011
Cited alongside, same era.
Minimax policies for combinatorial prediction games
Jean-Yves Audibert, Sébastien Bubeck, and Gábor Lugosi · 2011
Cited alongside, same era.
Online learning and online convex optimization
Shai Shalev-Shwartz · 2011
Cited alongside, same era.
Sequential transfer in multi-armed bandit with finite set of models
Alessandro Lazaric, Emma Brunskill, et al · 2013
Cited alongside, same era.
Fighting bandits with a new kind of smoothness
Jacob Abernethy, Chansoo Lee, and Ambuj Tewari · 2015
Cited alongside, same era.
Introduction to online convex optimization
Elad Hazan · 2015
Cited alongside, same era.
Meta-learning bandit policies by gradient ascent
Branislav Kveton, Martin Mladenov, Chih-Wei Hsu, Manzil Zaheer, Csaba Szepesvari, and Craig Boutilier · 2020
Later among the works it cites.
Learning-to-learn non-convex piecewise-Lipschitz functions
Maria-Florina Balcan, Mikhail Khodak, Dravyansh Sharma, and Ameet Talwalkar · 2021
Later among the works it cites.
No regrets for learning the prior in bandits
Soumya Basu, Branislav Kveton, Manzil Zaheer, and Csaba Szepesvári · 2021
Later among the works it cites.
Federated hyperparameter tuning: Challenges, baselines, and connections to weight-sharing
Mikhail Khodak, Renbo Tu, Tian Li, Liam Li, Maria-Florina Balcan, Virginia Smith, and Ameet Talwalkar · 2021
Later among the works it cites.
Meta-thompson sampling
Branislav Kveton, Mikhail Konobeev, Manzil Zaheer, Chih-wei Hsu, Martin Mladenov, Craig Boutilier, and Csaba Szepesvari · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Explore no more: Improved high-probability regret bounds for non-stochastic bandits
Gergely Neu · 2015
Cited alongside, same era.
Model-agnostic meta-learning for fast adaptation of deep networks
Chelsea Finn, Pieter Abbeel, and Sergey Levine · 2017
Cited alongside, same era.
Lecture 13
Haipeng Luo · 2017
Cited alongside, same era.
On first-order meta-learning algorithms
Alex Nichol, Joshua Achiam, and John Schulman · 2018
Cited alongside, same era.
Meta-learning with stochastic linear bandits
Leonardo Cella, Alessandro Lazaric, and Massimiliano Pontil · 2020
Cited alongside, same era.
Learning-to-learn stochastic gradient descent with biased regularization
Giulia Denevi, Carlo Ciliberto, Riccardo Grazzi, and Massimiliano Pontil
Cited in the paper.
Michael Mitzenmacher and Sergei Vassilvitskii · 2021
Later among the works it cites.
Meta-learning effective exploration strategies for contextual bandits
Amr Sharaf and III Hal Daumé · 2021
Later among the works it cites.
Bayesian decision-making under misspecified priors with applications to meta-learning
Max Simchowitz, Christopher Tosh, Akshay Krishnamurthy, Daniel J Hsu, Thodoris Lykouris, Miro Dudik, and Robert E Schapire · 2021
Later among the works it cites.
Non-stationary bandits and meta-learning with a small set of optimal arms
MohammadJavad Azizi, Thang Duong, Yasin Abbasi-Yadkori, András György, Claire Vernade, and Mohammad Ghavamzadeh · 2022
Closest in time.
Learning predictions for algorithms with predictions
Mikhail Khodak, Maria-Florina Balcan, Ameet Talwalkar, and Sergei Vassilvitskii · 2022
Closest in time.
Multi-environment meta-learning in stochastic linear bandits
Ahmadreza Moradipari, Mohammad Ghavamzadeh, Taha Rajabzadeh, Christos Thrampoulidis, and Mahnoosh Alizadeh · 2022
Closest in time.