Fetching the paper…
Reading the bibliography…
We study the stochastic contextual bandit problem, where the reward is generated from an unknown function with additive noise.
The jackknife, the bootstrap, and other resampling plans , volume 38
Efron, B · 1982
Earlier work this paper cites.
Gradient-based learning applied to document recognition
LeCun, Y., Bottou, L., Bengio, Y., and Haffner, P · 1998
Earlier work this paper cites.
Using confidence bounds for exploitation-exploration trade-offs
Auer, P · 2002
Earlier work this paper cites.
The nonstochastic multiarmed bandit problem
Auer, P., Cesa-Bianchi, N., Freund, Y., and Schapire, R. E · 2002
Earlier work this paper cites.
Reinforcement learning with immediate rewards and linear hypotheses
Abe, N., Biermann, A. W., and Long, P. M · 2003
Earlier work this paper cites.
Stochastic linear optimization under bandit feedback
Dani, V., Hayes, T. P., and Kakade, S. M · 2008
Earlier work this paper cites.
Efficient bandit algorithms for online multiclass prediction
Kakade, S. M., Shalev-Shwartz, S., and Tewari, A · 2008
Earlier work this paper cites.
Multi-armed bandits in metric spaces
Kleinberg, R., Slivkins, A., and Upfal, E · 2008
Earlier work this paper cites.
The epoch-greedy algorithm for contextual multi-armed bandits
Langford, J. and Zhang, T · 2008
Earlier work this paper cites.
The offset tree for learning with partial labels
Beygelzimer, A. and Langford, J · 2009
Earlier work this paper cites.
Parametric bandits: The generalized linear case
Filippi, S., Cappe, O., Garivier, A., and Szepesvári, C · 2010
Earlier work this paper cites.
A contextual-bandit approach to personalized news article recommendation
Li, L., Chu, W., Langford, J., and Schapire, R. E · 2010
Earlier work this paper cites.
Linearly parameterized bandits
Rusmevichientong, P. and Tsitsiklis, J. N · 2010
Earlier work this paper cites.
Gaussian process optimization in the bandit setting: no regret and experimental design
Srinivas, N., Krause, A., Kakade, S., and Seeger, M · 2010
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Abbasi-Yadkori, Y., Pál, D., and Szepesvári, C · 2011
Earlier work this paper cites.
Contextual bandit algorithms with supervised learning guarantees
Beygelzimer, A., Langford, J., Li, L., Reyzin, L., and Schapire, R. E · 2011
Earlier work this paper cites.
X-armed bandits
Bubeck, S., Munos, R., Stoltz, G., and Szepesvári, C · 2011
Earlier work this paper cites.
An empirical evaluation of thompson sampling
Chapelle, O. and Li, L · 2011
Earlier work this paper cites.
Contextual bandits with linear payoff functions
Chu, W., Li, L., Reyzin, L., and Schapire, R · 2011
Earlier work this paper cites.
Contextual Gaussian process bandit optimization
Krause, A. and Ong, C. S · 2011
Earlier work this paper cites.
Regret analysis of stochastic and nonstochastic multi-armed bandit problems
Bubeck, S. and Cesa-Bianchi, N · 2012
Cited alongside, same era.
Thompson sampling for contextual bandits with linear payoffs
Agrawal, S. and Goyal, N · 2013
Cited alongside, same era.
Finite-time analysis of kernelised contextual bandits
Valko, M., Korda, N., Munos, R., Flaounas, I., and Cristianini, N · 2013
Cited alongside, same era.
Taming the monster: A fast and simple algorithm for contextual bandits
Agarwal, A., Hsu, D., Kale, S., Langford, J., Li, L., and Schapire, R. E · 2014
Cited alongside, same era.
A neural networks committee for the contextual bandit problem
Allesiardo, R., Féraud, R., and Bouneffouf, D · 2014
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Jacot, A., Gabriel, F., and Hongler, C · 2018
Later among the works it cites.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Li, Y. and Liang, Y · 2018
Later among the works it cites.
BBQ-networks: Efficient exploration in deep reinforcement learning for task-oriented dialogue systems
Lipton, Z., Li, X., Gao, J., Li, L., Ahmed, F., and Deng, L · 2018
Later among the works it cites.
Deep Bayesian bandits showdown
Riquelme, C., Tucker, G., and Snoek, J · 2018
Later among the works it cites.
A tutorial on Thompson sampling
Russo, D., Roy, B. V., Kazerouni, A., Osband, I., and Wen, Z · 2018
Later among the works it cites.
Optimal approximation of continuous functions by very deep ReLU networks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Telgarsky, M · 2015
Cited alongside, same era.
Random forest for the contextual bandit problem
Féraud, R., Allesiardo, R., Urvoy, T., and Clérot, F · 2016
Cited alongside, same era.
Deep Learning
Goodfellow, I., Bengio, Y., and Courville, A · 2016
Cited alongside, same era.
Why deep neural networks for function approximation?
Liang, S. and Srikant, R · 2016
Cited alongside, same era.
Benefits of depth in neural networks
Telgarsky, M · 2016
Cited alongside, same era.
Online stochastic linear optimization under one-bit feedback
Zhang, L., Yang, T., Jin, R., Xiao, Y., and Zhou, Z.-H · 2016
Cited alongside, same era.
SGD learns the conjugate kernel class of the network
Daniely, A · 2017
Cited alongside, same era.
Yarotsky, D · 2018
Later among the works it cites.
What can ResNet learn efficiently, going beyond kernels?
Allen-Zhu, Z. and Li, Y · 2019
Closest in time.
A convergence theory for deep learning via over-parameterization
Allen-Zhu, Z., Li, Y., and Song, Z · 2019
Closest in time.
On exact computation with an infinitely wide neural net
Arora, S., Du, S. S., Hu, W., Li, Z., Salakhutdinov, R., and Wang, R · 2019
Closest in time.
Generalization bounds of stochastic gradient descent for wide and deep neural networks
Cao, Y. and Gu, Q · 2019
Closest in time.
How much over-parameterization is sufficient to learn deep relu networks?
Chen, Z., Cao, Y., Zou, D., and Gu, Q · 2019
Closest in time.
Bandit Algorithms
Lattimore, T. and Szepesvári, C · 2019
Closest in time.
Reinforcement leaning in feature space: Matrix bandit, kernels, and regret bound
Yang, L. F. and Wang, M · 2019
Closest in time.
Deep neural linear bandits: Overcoming catastrophic forgetting through likelihood matching
Zahavy, T. and Mannor, S · 2019
Closest in time.
An improved analysis of training over-parameterized deep neural networks
Zou, D. and Gu, Q · 2019
Closest in time.
Stochastic gradient descent optimizes over-parameterized deep ReLU networks
Zou, D., Cao, Y., Zhou, D., and Gu, Q · 2019
Closest in time.
Generalization error bounds of gradient descent for learning over-parameterized deep relu networks
Cao, Y. and Gu, Q · 2020
Closest in time.
Beyond ucb: Optimal and efficient contextual bandits with regression oracles
Foster, D. J. and Rakhlin, A · 2020
Closest in time.
Randomized exploration in generalized linear bandits
Kveton, B., Zaheer, M., Szepesvári, C., Li, L., Ghavamzadeh, M., and Boutilier, C · 2020
Closest in time.