Fetching the paper…
Reading the bibliography…
A major research direction in contextual bandits is to develop algorithms that are computationally efficient, yet support flexible, general-purpose function approximation.
On the complexity of approximating the maximal inscribed ellipsoid for a polytope
L. G. Khachiyan and M. J. Todd · 1990
Earlier work this paper cites.
Associative reinforcement learning using linear probabilistic concepts
N. Abe and P. M. Long · 1999
Earlier work this paper cites.
Using confidence bounds for exploitation-exploration trade-offs
P. Auer · 2002
Earlier work this paper cites.
The nonstochastic multiarmed bandit problem
P. Auer, N. Cesa-Bianchi, Y. Freund, and R. E. Schapire · 2002
Earlier work this paper cites.
Reinforcement learning with immediate rewards and linear hypotheses
N. Abe, A. W. Biermann, and P. M. Long · 2003
Earlier work this paper cites.
Minimum-volume enclosing ellipsoids and core sets
P. Kumar and E. A. Yıldırım · 2005
Earlier work this paper cites.
Logarithmic regret algorithms for online convex optimization
E. Hazan, A. Agarwal, and S. Kale · 2007
Earlier work this paper cites.
On Khachiyan’s algorithm for the computation of minimum-volume enclosing ellipsoids
M.J. Todd and E.A. Yıldırım · 2007
Earlier work this paper cites.
The epoch-greedy algorithm for multi-armed bandits with side information
J. Langford and T. Zhang · 2008
Earlier work this paper cites.
Minimax policies for adversarial and stochastic bandits
J.-Y. Audibert and S. Bubeck · 2009
Earlier work this paper cites.
A contextual-bandit approach to personalized news article recommendation
K. Li, W. Chu, J. Langford, and R. E. Schapire · 2010
Earlier work this paper cites.
Gaussian process optimization in the bandit setting: no regret and experimental design
N. Srinivas, A. Krause, S. Kakade, and M. Seeger · 2010
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Y. Abbasi-Yadkori, D. Pál, and C. Szepesvári · 2011
Earlier work this paper cites.
Contextual bandits with linear payoff functions
W. Chu, L. Li, L. Reyzin, and R. Schapire · 2011
Earlier work this paper cites.
Efficient optimal learning for contextual bandits
M. Dudik, D. Hsu, S. Kale, N. Karampatziakis, J. Langford, L. Reyzin, and T. Zhang · 2011
Earlier work this paper cites.
Efficient learning of generalized linear and single index models with isotonic regression
S. M. Kakade, V. Kanade, O. Shamir, and A. Kalai · 2011
Earlier work this paper cites.
Contextual gaussian process bandit optimization
A. Krause and C.S. Ong · 2011
Earlier work this paper cites.
Online-to-confidence-set conversions and application to sparse stochastic bandits
Y. Abbasi-Yadkori, D. Pal, and C. Szepesvári · 2012
Cited alongside, same era.
Contextual bandit learning with predictable rewards
A. Agarwal, M. Dudik, S. Kale, J. Langford, and R. Schapire · 2012
Cited alongside, same era.
Thompson sampling for contextual bandits with linear payoffs
S. Agrawal and N. Goyal · 2013
Cited alongside, same era.
Multiclass classification with bandit feedback using adaptive regularization
K. Crammer and C. Gentile · 2013
Cited alongside, same era.
High-dimensional gaussian process bandits
J. Djolonga, A. Krause, and V. Cevher · 2013
Cited alongside, same era.
Finite-time analysis of kernelised contextual bandits
M. Valko, N. Korda, R. Munos, I. Flaounas, and N. Cristianini · 2013
Cited alongside, same era.
Practical contextual bandits with regression oracles
D. J. Foster, A. Agarwal, M. Dudik, H. Luo, and R. Schapire · 2018
Later among the works it cites.
Efficient contextual bandits in non-stationary worlds
H. Luo, C-Y. Wei, A. Agarwal, and J. Langford · 2018
Later among the works it cites.
Stochastic bandits robust to adversarial corruptions
T. Lykouris, V. Mirrokni, and R. Paes Leme · 2018
Later among the works it cites.
Semi-parametric efficient policy learning with continuous actions
V. Chernozhukov, M. Demirer, G. Lewis, and V. Syrgkanis · 2019
Later among the works it cites.
Is a good representation sufficient for sample efficient reinforcement learning?
S. S. Du, S. M. Kakade, R. Wang, and L. F. Yang · 2019
Later among the works it cites.
Model selection for contextual bandits
D. J. Foster, A. Krishnamurthy, and H. Luo · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Taming the monster: A fast and simple algorithm for contextual bandits
A. Agarwal, D. Hsu, S. Kale, J. Langford, L. Li, and R. Schapire · 2014
Cited alongside, same era.
One practical algorithm for both stochastic and adversarial bandits
Yevgeny Seldin and Aleksandrs Slivkins · 2014
Cited alongside, same era.
Fighting bandits with a new kind of smoothness
J. D. Abernethy, C. Lee, and A. Tewari · 2015
Cited alongside, same era.
A chaining algorithm for online nonparametric regression
P. Gaillard and S. Gerchinovitz · 2015
Cited alongside, same era.
Safe exploration for optimization with gaussian processes
Y. Sui, A. Gotovos, J. Burdick, and A. Krause · 2015
Cited alongside, same era.
Making contextual decisions with low technical debt
A. Agarwal, S. Bird, M. Cozowicz, L. Hoang, J. Langford, S. Lee, J. Li, D. Melamed, G. Oshri, and O. Ribas · 2016
Cited alongside, same era.
Later among the works it cites.
Better algorithms for stochastic bandits with adversarial corruptions
A. Gupta, T. Koren, and K. Talwar · 2019
Later among the works it cites.
An optimal algorithm for stochastic and adversarial bandits
J. Zimmert and Y. Seldin · 2019
Later among the works it cites.
Corruption-tolerant gaussian process bandit optimization
I. Bogunovic, A. Krause, and J. Scarlett · 2020
Later among the works it cites.
Beyond UCB: Optimal and efficient contextual bandits with regression oracles
D. J. Foster and A. Rakhlin · 2020
Later among the works it cites.
D. J. Foster, A. Rakhlin, D. Simchi-Levi, and Y. Xu · 2020
Later among the works it cites.
Learning with good feature representations in bandits and in RL with a generative model
Tor Lattimore, Csaba Szepesvari, and Gellert Weisz · 2020
Later among the works it cites.
Model selection in contextual stochastic bandit problems
A. Pacchiano, M. Phan, Y. Abbasi-Yadkori, A. Rao, J. Zimmert, T. Lattimore, and C. Szepesvari · 2020
Later among the works it cites.
Bypassing the monster: A faster and simpler optimal algorithm for contextual bandits under realizability
D. Simchi-Levi and Y. Xu · 2020
Later among the works it cites.
Upper counterfactual confidence bounds: a new optimism principle for contextual bandits
Y. Xu and A. Zeevi · 2020
Later among the works it cites.
Learning near optimal policies with low inherent Bellman error
A. Zanette, A. Lazaric, M. Kochenderfer, and E. Brunskill · 2020
Later among the works it cites.