Fetching the paper…
Reading the bibliography…
Recent works have shown a reduction from contextual bandits to online regression under a realizability assumption [Foster and Rakhlin, 2020, Foster and Krishnamurthy, 2021].
Online learning and online convex optimization
S. Shalev-Shwartz · 1935
Earlier work this paper cites.
A topological property of real analytic subsets (in french)
S. Lojasiewicz · 1963
Earlier work this paper cites.
Gradient methods for the minimisation of functionals
B. Polyak · 1963
Earlier work this paper cites.
Associative reinforcement learning using linear probabilistic concepts
N. Abe and P. M. Long · 1999
Earlier work this paper cites.
Relative loss bounds for on-line density estimation with the exponential family of distributions
K. S. Azoury and M. K. Warmuth · 2001
Earlier work this paper cites.
Reinforcement learning with immediate rewards and linear hypotheses
N. Abe, A. W. Biermann, and P. M. Long · 2003
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Y. Abbasi-Yadkori, D. Pál, and C. Szepesvári · 2011
Earlier work this paper cites.
Contextual bandits with linear payoff functions
W. Chu, L. Li, L. Reyzin, and R. Schapire · 2011
Earlier work this paper cites.
Introduction to the non-asymptotic analysis of random matrices, 2011
R. Vershynin · 2011
Earlier work this paper cites.
Thompson sampling for contextual bandits with linear payoffs
S. Agrawal and N. Goyal · 2013
Earlier work this paper cites.
Finite-time analysis of kernelised contextual bandits
M. Valko, N. Korda, R. Munos, I. Flaounas, and N. Cristianini · 2013
Earlier work this paper cites.
Tight Bounds for the Expected Risk of Linear Classifiers and PAC-Bayes Finite-Sample Guarantees
J. Honorio and T. Jaakkola · 2014
Earlier work this paper cites.
On reverse pinsker inequalities, 2015
I. Sason · 2015
Cited alongside, same era.
Linear convergence of gradient and proximal-gradient methods under the Polyak-lojasiewicz condition
H. Karimi, J. Nutini, and M. Schmidt · 2016
Cited alongside, same era.
Ensemble sampling
X. Lu and B. Van Roy · 2017
Cited alongside, same era.
Stability and generalization of learning algorithms that converge to global optima
Z. Charles and D. Papailiopoulos · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
A. Jacot, F. Gabriel, and C. Hongler · 2018
Cited alongside, same era.
Deep bayesian bandits showdown: An empirical comparison of bayesian deep networks for thompson sampling
C. Riquelme, G. Tucker, and J. Snoek · 2018
Cited alongside, same era.
Local clustering in contextual multi-armed bandits
Y. Ban and J. He · 2021
Later among the works it cites.
Multi-facet contextual bandits: A neural network perspective
Y. Ban, J. He, and C. B. Cook · 2021
Later among the works it cites.
Provable regret bounds for deep online learning and control
X. Chen, E. Minasyan, J. D. Lee, and E. Hazan · 2021
Later among the works it cites.
Efficient first-order contextual bandits: Prediction, allocation, and triangular discrimination
D. J. Foster and A. Krishnamurthy · 2021
Later among the works it cites.
Proxy convexity: A unified framework for the analysis of neural networks trained by gradient descent
S. Frei and Q. Gu · 2021
Later among the works it cites.
Introduction to online convex optimization, 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Gradient descent finds global minima of deep neural networks
S. Du, J. Lee, H. Li, L. Wang, and X. Zhai · 2019
Cited alongside, same era.
Generic outlier detection in multi-armed bandit
Y. Ban and J. He · 2020
Cited alongside, same era.
Beyond ucb: Optimal and efficient contextual bandits with regression oracles
D. Foster and A. Rakhlin · 2020
Cited alongside, same era.
On the linearity of large non-linear models: when and why the tangent kernel is constant
C. Liu, L. Zhu, and M. Belkin · 2020
Cited alongside, same era.
D. Simchi-Levi and Y. Xu · 2020
Cited alongside, same era.
Neural linear bandits: Overcoming catastrophic forgetting through likelihood matching, 2020
T. Zahavy and S. Mannor · 2020
Cited alongside, same era.
E. Hazan · 2021
Later among the works it cites.
Neural thompson sampling
W. Zhang, D. Zhou, L. Li, and Q. Gu · 2021
Later among the works it cites.
Neural contextual bandits with ucb-based exploration
D. Zhou, L. Li, and Q. Gu · 2021
Later among the works it cites.
Loss landscapes and optimization in over-parameterized non-linear systems and neural networks
C. Liu, L. Zhu, and M. Belkin · 2022
Later among the works it cites.
Neural bandit with arm group graph
Y. Qi, Y. Ban, and J. He · 2022
Later among the works it cites.
Restricted strong convexity of deep learning models with smooth activations
A. Banerjee, P. Cisneros-Velarde, L. Zhu, and M. Belkin · 2023
Closest in time.
Graph neural bandits
Y. Qi, Y. Ban, and J. He · 2023
Closest in time.