Fetching the paper…
Reading the bibliography…
We consider an adversarial variant of the classic $K$-armed linear contextual bandit problem where the sequence of loss functions associated with each arm are allowed to change without restriction over time.
A stochastic approximation method
H. Robbins and S. Monro · 1951
Earlier work this paper cites.
The weighted majority algorithm
N. Littlestone and M. Warmuth · 1994
Earlier work this paper cites.
On the averaged stochastic approximation for linear regression
L. Györfi and H. Walk · 1996
Earlier work this paper cites.
Associative reinforcement learning using linear probabilistic concepts
N. Abe and P. M. Long · 1999
Earlier work this paper cites.
Using confidence bounds for exploitation-exploration trade-offs
P. Auer · 2002
Earlier work this paper cites.
Prediction, Learning, and Games
N. Cesa-Bianchi and G. Lugosi · 2006
Earlier work this paper cites.
The price of bandit information for online optimization
V. Dani, T. Hayes, and S. Kakade · 2008
Earlier work this paper cites.
Efficient bandit algorithms for online multiclass prediction
S. M. Kakade, S. Shalev-Shwartz, and A. Tewari · 2008
Earlier work this paper cites.
Parametric bandits: The generalized linear case
S. Filippi, O. Cappé, A. Garivier, and Cs · 2010
Earlier work this paper cites.
A contextual-bandit approach to personalized news article recommendation
L. Li, W. Chu, J. Langford, and R. E. Schapire · 2010
Earlier work this paper cites.
Linearly parameterized bandits
P. Rusmevichientong and J. Tsitsiklis · 2010
Earlier work this paper cites.
Gaussian process optimization in the bandit setting: No regret and experimental design
N. Srinivas, A. Krause, S. M. Kakade, and M. Seeger · 2010
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Y. Abbasi-Yadkori, D. Pál, and Cs. Szepesvári · 2011
Earlier work this paper cites.
Contextual bandits with linear payoff functions
W. Chu, L. Li, L. Reyzin, and R. Schapire · 2011
Earlier work this paper cites.
Efficient optimal learning for contextual bandits
M. Dudík, D. Hsu, S. Kale, N. Karampatziakis, J. Langford, L. Reyzin, and T. Zhang · 2011
Earlier work this paper cites.
Regret Analysis of Stochastic and Nonstochastic Multi-armed Bandit Problems
S. Bubeck and N. Cesa-Bianchi · 2012
Cited alongside, same era.
Thompson sampling for contextual bandits with linear payoffs
S. Agrawal and N. Goyal · 2013
Cited alongside, same era.
Non-strongly-convex smooth stochastic approximation with convergence rate O ( 1 / n ) O(1/n)
F. Bach and E. Moulines · 2013
Cited alongside, same era.
Concentration inequalities:A Nonasymptotic Theory of Independence
S. Boucheron, G. Lugosi, and P. Massart · 2013
Cited alongside, same era.
An efficient algorithm for learning with semi-bandit feedback
G. Neu and G. Bartók · 2013
Cited alongside, same era.
Taming the monster: A fast and simple algorithm for contextual bandits
A. Agarwal, D. Hsu, S. Kale, J. Langford, L. Li, and R. Schapire · 2014
From ads to interventions: Contextual bandits in mobile health
A. Tewari and S. A. Murphy · 2017
Later among the works it cites.
Make the minority great again: First-order regret bound for contextual bandits
Z. Allen-Zhu, S. Bubeck, and Y. Li · 2018
Later among the works it cites.
Contextual bandits with surrogate losses: Margin bounds and efficient algorithms
D. J. Foster and A. Krishnamurthy · 2018
Later among the works it cites.
Iterate averaging as regularization for stochastic gradient descent
G. Neu and L. Rosasco · 2018
Later among the works it cites.
Gaussian process optimization with adaptive sketching: Scalable and no regret
D. Calandriello, L. Carratino, A. Lazaric, M. Valko, and L. Rosasco · 2019
Later among the works it cites.
Learning to optimize under non-stationarity
W. C. Cheung, D. Simchi-Levi, and R. Zhu · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Regret in online combinatorial optimization
J.-Y. Audibert, S. Bubeck, and G. Lugosi · 2014
Cited alongside, same era.
Online combinatorial optimization with stochastic decision sets and adversarial losses
G. Neu and M. Valko · 2014
Cited alongside, same era.
Importance weighting without importance weights: An efficient algorithm for combinatorial semi-bandits
G. Neu and G. Bartók · 2016
Cited alongside, same era.
BISTRO: An efficient relaxation-based method for contextual bandits
A. Rakhlin and K. Sridharan · 2016
Cited alongside, same era.
Open problem: First-order regret bounds for contextual bandits
A. Agarwal, A. Krishnamurthy, J. Langford, H. Luo, and S. R. E · 2017
Cited alongside, same era.
Efficient online bandit multiclass learning with O ~ ( T ) \widetilde{O}(\sqrt{T}) regret
A. Beygelzimer, F. Orabona, and C. Zhang · 2017
Cited alongside, same era.
Later among the works it cites.
Is a good representation sufficient for sample efficient reinforcement learning?
S. S. Du, S. M. Kakade, R. Wang, and L. F. Yang · 2019
Later among the works it cites.
Model selection for contextual bandits
D. J. Foster, A. Krishnamurthy, and H. Luo · 2019
Later among the works it cites.
Near-optimal oracle-efficient algorithms for stationary and non-stationary stochastic linear bandits
B. Kim and A. Tewari · 2019
Later among the works it cites.
Bandit algorithms
T. Lattimore and Cs. Szepesvári · 2019
Later among the works it cites.
Weighted linear bandits for non-stationary environments
Y. Russac, C. Vernade, and O. Cappé · 2019
Later among the works it cites.
Comments on the Du-Kakade-Wang-Yang lower bounds
B. Van Roy and S. Dong · 2019
Later among the works it cites.
Beyond UCB: Optimal and efficient contextual bandits with regression oracles
D. J. Foster and A. Rakhlin · 2020
Closest in time.
Learning with good feature representations in bandits and in RL with a generative model
T. Lattimore, Cs. Szepesvári, and G. Weisz · 2020
Closest in time.