Fetching the paper…
Reading the bibliography…
We introduce a simple and efficient algorithm for stochastic linear bandits with finitely many actions that is asymptotically optimal and (nearly) worst-case optimal in finite time.
Aggregating strategies
Volodimir G Vovk · 1990
Earlier work this paper cites.
The weighted majority algorithm
Nick Littlestone and Manfred K Warmuth · 1994
Earlier work this paper cites.
Asymptotically efficient adaptive choice of control laws incontrolled markov chains
Todd L Graves and Tze Leung Lai · 1997
Earlier work this paper cites.
Using confidence bounds for exploitation-exploration trade-offs
Peter Auer · 2003
Earlier work this paper cites.
Improved Second-Order Bounds for Prediction with Expert Advice
Nicolò Cesa-Bianchi, Yishay Mansour, and Gilles Stoltz · 2005
Earlier work this paper cites.
Faster and simpler algorithms for multicommodity flow and other fractional packing problems
Naveen Garg and Jochen Koenemann · 2007
Earlier work this paper cites.
Online learning: Theory, algorithms, and applications, 2007
Shai Shalev-Shwartz and Yoram Singer · 2007
Earlier work this paper cites.
Stochastic Linear Optimization under Bandit Feedback
Varsha Dani, Thomas P. Hayes, and Sham M. Kakade · 2008
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Yasin Abbasi-Yadkori, Dávid Pál, and Csaba Szepesvári · 2011
Earlier work this paper cites.
The multiplicative weights update method: a meta-algorithm and applications
Sanjeev Arora, Elad Hazan, and Satyen Kale · 2012
Earlier work this paper cites.
Thompson sampling for contextual bandits with linear payoffs
Shipra Agrawal and Navin Goyal · 2013
Cited alongside, same era.
Learning to optimize via information-directed sampling
Daniel Russo and Benjamin Van Roy · 2014
Cited alongside, same era.
Best-arm identification in linear bandits
Marta Soare, Alessandro Lazaric, and Rémi Munos · 2014
Cited alongside, same era.
18. s997: High dimensional statistics
Philippe Rigollet · 2015
Cited alongside, same era.
Introduction to online convex optimization
Elad Hazan et al · 2016
Cited alongside, same era.
Minimal exploration in structured stochastic bandits
Richard Combes, Stefan Magureanu, and Alexandre Proutiere · 2017
Cited alongside, same era.
Bandit Algorithms
Tor Lattimore and Czsaba Szepesvari · 2019
Later among the works it cites.
A modern introduction to online learning
Francesco Orabona · 2019
Later among the works it cites.
Structure adaptive algorithms for stochastic bandits
Rémy Degenne, Han Shao, and Wouter Koolen · 2020
Closest in time.
Crush optimism with pessimism: Structured bandits beyond asymptotic optimality
Kwang-Sung Jun and Chicheng Zhang · 2020
Closest in time.
Information directed sampling for linear partial monitoring
Johannes Kirschner, Tor Lattimore, and Andreas Krause · 2020
Closest in time.
Simple bayesian algorithms for best-arm identification
Daniel Russo · 2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tor Lattimore and Csaba Szepesvári · 2017
Cited alongside, same era.
Information directed sampling and bandits with heteroscedastic noise
Johannes Kirschner and Andreas Krause · 2018
Cited alongside, same era.
Non-asymptotic pure exploration by solving games
Rémy Degenne, Wouter M Koolen, and Pierre Ménard · 2019
Cited alongside, same era.
Adaptive exploration in linear contextual bandit
Botao Hao, Tor Lattimore, and Csaba Szepesvari · 2019
Cited alongside, same era.
An asymptotically optimal primal-dual incremental algorithm for contextual linear bandits
Andrea Tirinzoni, Matteo Pirotta, Marcello Restelli, and Alessandro Lazaric · 2020
Closest in time.
Optimal learning for structured bandits
Bart PG Van Parys and Negin Golrezaei · 2020
Closest in time.
Experimental design for regret minimization in linear bandits
Andrew Wagenmaker, Julian Katz-Samuels, and Kevin Jamieson · 2020
Closest in time.