Fetching the paper…
Reading the bibliography…
Tractable contextual bandit algorithms often rely on the realizability assumption - i.e., that the true expected reward model belongs to a known class, such as linear functions.
What is local optimality in nonconvex-nonconcave minimax optimization?
Jin, C., Netrapalli, P., and Jordan, M. I. (2019) · 1902
Earlier work this paper cites.
Smoothness-adaptive stochastic bandits
Gur, Y., Momeni, A., and Wager, S. (2019) · 1910
Earlier work this paper cites.
Associative reinforcement learning using linear probabilistic concepts
Abe, N. and Long, P. M. (1999) · 1999
Earlier work this paper cites.
Beyond ucb: Optimal and efficient contextual bandits with regression oracles
Foster, D. J. and Rakhlin, A. (2020) · 2002
Earlier work this paper cites.
Model selection in contextual stochastic bandit problems
Pacchiano, A., Phan, M., Abbasi-Yadkori, Y., Rao, A., Zimmert, J., Lattimore, T., and Szepesvari, C. (2020) · 2003
Earlier work this paper cites.
Smooth contextual bandits: Bridging the parametric and non-differentiable regret regimes
Hu, Y., Kallus, N., and Mao, X. (2020) · 2010
Earlier work this paper cites.
A contextual-bandit approach to personalized news article recommendation
Li, L., Chu, W., Langford, J., and Schapire, R. E. (2010) · 2010
Earlier work this paper cites.
Nonparametric bandits with covariates
Rigollet, P. and Zeevi, A. (2010) · 2010
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Abbasi-Yadkori, Y., Pál, D., and Szepesvári, C. (2011) · 2011
Earlier work this paper cites.
Contextual bandit algorithms with supervised learning guarantees
Beygelzimer, A., Langford, J., Li, L., Reyzin, L., and Schapire, R. (2011) · 2011
Earlier work this paper cites.
Efficient optimal learning for contextual bandits
Dudik, M., Hsu, D., Kale, S., Karampatziakis, N., Langford, J., Reyzin, L., and Zhang, T. (2011) · 2011
Cited alongside, same era.
Thompson sampling for contextual bandits with linear payoffs
Agrawal, S. and Goyal, N. (2013) · 2013
Cited alongside, same era.
The multi-armed bandit problem with covariates
Perchet, V., Rigollet, P., et al. (2013) · 2013
Cited alongside, same era.
Taming the monster: A fast and simple algorithm for contextual bandits
Agarwal, A., Hsu, D., Kale, S., Langford, J., Li, L., and Schapire, R. (2014) · 2014
Cited alongside, same era.
Interactive machine learning
Langford, J. (2014) · 2014
Cited alongside, same era.
Convex optimization algorithms
Bertsekas, D. P. and Scientific, A. (2015) · 2015
Practical contextual bandits with regression oracles
Foster, D. J., Agarwal, A., Dudík, M., Luo, H., and Schapire, R. E. (2018) · 2018
Later among the works it cites.
A tutorial on thompson sampling
Russo, D. J., Roy, B. V., Kazerouni, A., Osband, I., and Wen, Z. (2018) · 2018
Later among the works it cites.
Model selection for contextual bandits
Foster, D. J., Krishnamurthy, A., and Luo, H. (2019) · 2019
Later among the works it cites.
Adapting to misspecification in contextual bandits
Foster, D. J., Gentile, C., Mohri, M., and Zimmert, J. (2020) · 2020
Closest in time.
Bandit algorithms
Lattimore, T. and Szepesvári, C. (2020) · 2020
Closest in time.
Learning with good feature representations in bandits and in rl with a generative model
Lattimore, T., Szepesvari, C., and Weisz, G. (2020) · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Making contextual decisions with low technical debt
Agarwal, A., Bird, S., Cozowicz, M., Hoang, L., Langford, J., Lee, S., Li, J., Melamed, D., Oshri, G., Ribas, O., et al. (2016) · 2016
Cited alongside, same era.
Corralling a band of bandit algorithms
Agarwal, A., Luo, H., Neyshabur, B., and Schapire, R. E. (2017) · 2017
Cited alongside, same era.
Misspecified linear bandits
Ghosh, A., Chowdhury, S. R., and Gopalan, A. (2017) · 2017
Cited alongside, same era.
From ads to interventions: Contextual bandits in mobile health
Tewari, A. and Murphy, S. A. (2017) · 2017
Cited alongside, same era.
Closest in time.
Efficient and robust algorithms for adversarial linear contextual bandits
Neu, G. and Olkhovskaya, J. (2020) · 2020
Closest in time.
Bypassing the monster: A faster and simpler optimal algorithm for contextual bandits under realizability
Simchi-Levi, D. and Xu, Y. (2020) · 2020
Closest in time.
Learning near optimal policies with low inherent bellman error
Zanette, A., Lazaric, A., Kochenderfer, M., and Brunskill, E. (2020) · 2020
Closest in time.
Oracle Inequalities in Empirical Risk Minimization and Sparse Recovery Problems: Ecole d’Eté de Probabilités de Saint-Flour XXXVIII-2008
Koltchinskii, V. (2011) · 2033
Closest in time.