Fetching the paper…
Reading the bibliography…
We study a general class of contextual bandits, where each context-action pair is associated with a raw feature vector, but the reward generating function is unknown.
Deep neural linear bandits: Overcoming catastrophic forgetting through likelihood matching
Zahavy, T · 1901
Earlier work this paper cites.
Exploration–exploitation tradeoff using variance estimates in multi-armed bandits
Audibert, J.-Y · 1902
Earlier work this paper cites.
A generalization theory of gradient descent for learning over-parameterized deep relu networks
Cao, Y · 1902
Earlier work this paper cites.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
Thompson, W. R · 1933
Earlier work this paper cites.
Approximation by superpositions of a sigmoidal function
Cybenko, G · 1989
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
Auer, P · 2002
Earlier work this paper cites.
Self-supervised contextual bandits in computer vision
Deshmukh, A. A · 2003
Earlier work this paper cites.
Stochastic linear optimization under bandit feedback
Dani, V · 2008
Earlier work this paper cites.
The epoch-greedy algorithm for multi-armed bandits with side information
Langford, J · 2008
Earlier work this paper cites.
Parametric bandits: The generalized linear case
Filippi, S · 2010
Earlier work this paper cites.
A contextual-bandit approach to personalized news article recommendation
Li, L · 2010
Earlier work this paper cites.
Linearly parameterized bandits
Rusmevichientong, P · 2010
Earlier work this paper cites.
Zhang, W · 2010
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Abbasi-Yadkori, Y · 2011
Earlier work this paper cites.
An empirical evaluation of thompson sampling
Chapelle, O · 2011
Cited alongside, same era.
Contextual bandits with linear payoff functions
Chu, W · 2011
Cited alongside, same era.
Finite-time analysis of kernelised contextual bandits
Valko, M · 2013
Cited alongside, same era.
Taming the monster: A fast and simple algorithm for contextual bandits
Agarwal, A · 2014
Cited alongside, same era.
Optimal computational and statistical rates of convergence for sparse nonconvex learning problems
Wang, Z · 2014
Cited alongside, same era.
Deep learning
LeCun, Y · 2015
Cited alongside, same era.
Scalable bayesian optimization using deep neural networks
Neural tangent kernel: Convergence and generalization in neural networks
Jacot, A · 2018
Later among the works it cites.
Deep bayesian bandits showdown: An empirical comparison of bayesian deep networks for thompson sampling
Riquelme, C · 2018
Later among the works it cites.
Stochastic gradient descent optimizes over-parameterized deep relu networks
Zou, D · 2018
Later among the works it cites.
Neural temporal-difference learning converges to global optima
Cai, Q · 2019
Later among the works it cites.
Wide neural networks of any depth evolve as linear models under gradient descent
Lee, J · 2019
Later among the works it cites.
Neural trust region/proximal policy optimization attains globally optimal policy
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Snoek, J · 2015
Cited alongside, same era.
Deep learning
Goodfellow, I · 2016
Cited alongside, same era.
Statistical guarantees for the em algorithm: From population to sample-based analysis
Balakrishnan, S · 2017
Cited alongside, same era.
On kernelized multi-armed bandits
Chowdhury, S. R · 2017
Cited alongside, same era.
UCI machine learning repository. URL http://archive.ics.uci.edu/ml
Dua, D · 2017
Cited alongside, same era.
Provably optimal algorithms for generalized linear contextual bandits
Li, L · 2017
Cited alongside, same era.
Liu, B · 2019
Later among the works it cites.
An improved analysis of training over-parameterized deep neural networks
Zou, D · 2019
Later among the works it cites.
Randomized exploration in generalized linear bandits
Kveton, B · 2020
Closest in time.
Bandit algorithms
Lattimore, T · 2020
Closest in time.
Towards moderate overparameterization: global convergence guarantees for training shallow neural networks
Oymak, S · 2020
Closest in time.
Neural policy gradient methods: Global optimality and rates of convergence
Wang, L · 2020
Closest in time.
A finite-time analysis of q-learning with neural network function approximation
Xu, P · 2020
Closest in time.
Neural contextual bandits with ucb-based exploration
Zhou, D · 2020
Closest in time.