Fetching the paper…
Reading the bibliography…
We study meta-learning for adversarial multi-armed bandits.
Robust and efficient estimation by minimising a density power divergence
Ayanendranath Basu, Ian R Harris, Nils L Hjort, and MC Jones · 1998
Earlier work this paper cites.
The nonstochastic multiarmed bandit problem
Peter Auer, Nicolo Cesa-Bianchi, Yoav Freund, and Robert E Schapire · 2002
Earlier work this paper cites.
Logarithmic regret algorithms for online convex optimization
Elad Hazan, Amit Agarwal, and Satyen Kale · 2007
Earlier work this paper cites.
Minimax policies for adversarial and stochastic bandits
Jean-Yves Audibert, Sébastien Bubeck, et al · 2009
Earlier work this paper cites.
Families of alpha-beta-and gamma-divergences: Flexible and robust measures of similarities
Andrzej Cichocki and Shun-ichi Amari · 2010
Earlier work this paper cites.
Minimax policies for combinatorial prediction games
Jean-Yves Audibert, Sébastien Bubeck, and Gábor Lugosi · 2011
Earlier work this paper cites.
Regret analysis of stochastic and nonstochastic multi-armed bandit problems
Sébastien Bubeck and Nicolo Cesa-Bianchi · 2012
Earlier work this paper cites.
The best of both worlds: Stochastic and adversarial bandits
Sébastien Bubeck and Aleksandrs Slivkins · 2012
Earlier work this paper cites.
Hoeffding’s inequality for supermartingales
Xiequan Fan, Ion Grama, and Quansheng Liu · 2012
Earlier work this paper cites.
Meta-learning of exploration/exploitation strategies: The multi-armed bandit case
Francis Maes, Louis Wehenkel, and Damien Ernst · 2012
Earlier work this paper cites.
Fighting bandits with a new kind of smoothness
Jacob Abernethy, Chansoo Lee, and Ambuj Tewari · 2015
Cited alongside, same era.
From ads to interventions: Contextual bandits in mobile health
Ambuj Tewari and Susan A Murphy · 2017
Cited alongside, same era.
Best of both worlds: Stochastic & adversarial best-arm identification
Yasin Abbasi-Yadkori, Peter Bartlett, Victor Gabillon, Alan Malek, and Michal Valko · 2018
Cited alongside, same era.
Joaquin Vanschoren · 2018
Cited alongside, same era.
Online-within-online meta-learning
Giulia Denevi, Dimitris Stamos, Carlo Ciliberto, and Massimiliano Pontil · 2019
Cited alongside, same era.
Meta-learning in neural networks: A survey
Timothy Hospedales, Antreas Antoniou, Paul Micaelli, and Amos Storkey · 2020
Later among the works it cites.
Bandit algorithms
Tor Lattimore and Csaba Szepesvári · 2020
Later among the works it cites.
Provable benefits of representation learning in linear bandits
Jiaqi Yang, Wei Hu, Jason D Lee, and Simon S Du · 2020
Later among the works it cites.
Branislav Kveton, Mikhail Konobeev, Manzil Zaheer, Chih-wei Hsu, Martin Mladenov, Craig Boutilier, and Csaba Szepesvari · 2021
Later among the works it cites.
Meta-strategy for learning tuning parameters with guarantees
Dimitri Meunier and Pierre Alquier · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Elad Hazan · 2019
Cited alongside, same era.
Adaptive gradient-based meta-learning methods
Mikhail Khodak, Maria-Florina F Balcan, and Ameet S Talwalkar · 2019
Cited alongside, same era.
A modern introduction to online learning
Francesco Orabona · 2019
Cited alongside, same era.
Meta-learning of sequential strategies
Pedro A Ortega, Jane X Wang, Mark Rowland, Tim Genewein, Zeb Kurth-Nelson, Razvan Pascanu, Nicolas Heess, Joel Veness, Alex Pritzel, Pablo Sprechmann, et al · 2019
Cited alongside, same era.
Meta-learning with stochastic linear bandits
Leonardo Cella, Alessandro Lazaric, and Massimiliano Pontil · 2020
Cited alongside, same era.
Later among the works it cites.
Tsallis-inf: An optimal algorithm for stochastic and adversarial bandits
Julian Zimmert and Yevgeny Seldin · 2021
Later among the works it cites.
Non-stationary bandits and meta-learning with a small set of optimal arms
MohammadJavad Azizi, Thang Duong, Yasin Abbasi-Yadkori, András György, Claire Vernade, and Mohammad Ghavamzadeh · 2022
Closest in time.
Meta-learning adversarial bandits
Maria-Florina Balcan, Keegan Harris, Mikhail Khodak, and Zhiwei Steven Wu · 2022
Closest in time.
Metalearning linear bandits by prior update
Amit Peleg, Naama Pearl, and Ron Meir · 2022
Closest in time.