Fetching the paper…
Reading the bibliography…
We consider the general (stochastic) contextual bandit problem under the realizability assumption, i.e., the expected reward, as a function of contexts and actions, belongs to a general function class $\mathcal{F}$.
1911
Earlier work this paper cites.
Abe N, Long P (1999) Associative reinforcement learning using linear probabilistic concepts. International Conference on Machine Learning
1999
Earlier work this paper cites.
Yang Y, Barron A (1999) Information-theoretic determination of minimax rates of convergence. Annals of Statistics 1564–1599
1999
Earlier work this paper cites.
Auer P, Cesa-Bianchi N, Freund Y, Schapire RE (2002) The nonstochastic multiarmed bandit problem. SIAM Journal on Computing 32(1):48–77
2002
Earlier work this paper cites.
Abe N, Biermann A, Long P (2003) Reinforcement learning with immediate rewards and linear hypotheses. Algorithmica 37(4):263–293
2003
Earlier work this paper cites.
2003
Earlier work this paper cites.
2007
Earlier work this paper cites.
Langford J, Zhang T (2008) The epoch-greedy algorithm for multi-armed bandits with side information. Advances in Neural Information Processing Systems , 817–824
2008
Earlier work this paper cites.
Klivans AR, Sherstov AA (2009) Cryptographic hardness for learning intersections of halfspaces. Journal of Computer and System Sciences 75(1):2–12
2009
Earlier work this paper cites.
McMahan HB, Streeter M (2009) Tighter bounds for multi-armed bandits with expert advice. Conference on Learning Theory
2009
Earlier work this paper cites.
Filippi S, Cappe O, Garivier A, Szepesvári C (2010) Parametric bandits: The generalized linear case. Advances in Neural Information Processing Systems , 586–594
2010
Earlier work this paper cites.
2010
Earlier work this paper cites.
Li L, Chu W, Langford J, Schapire RE (2010) A contextual-bandit approach to personalized news article recommendation. Proceedings of the 19th international conference on World Wide Web , 661–670
2010
Earlier work this paper cites.
Abbasi-Yadkori Y, Pál D, Szepesvári C (2011) Improved algorithms for linear stochastic bandits. Advances in Neural Information Processing Systems , 2312–2320
2011
Earlier work this paper cites.
Beygelzimer A, Langford J, Li L, Reyzin L, Schapire R (2011) Contextual bandit algorithms with supervised learning guarantees. International Conference on Artificial Intelligence and Statistics , 19–26
2011
Earlier work this paper cites.
Chapelle O, Li L (2011) An empirical evaluation of thompson sampling. Advances in Neural Information Processing Systems , 2249–2257
2011
Earlier work this paper cites.
Chu W, Li L, Reyzin L, Schapire R (2011) Contextual bandits with linear payoff functions. International Conference on Artificial Intelligence and Statistics , 208–214
2011
Earlier work this paper cites.
Dudik M, Hsu D, Kale S, Karampatziakis N, Langford J, Reyzin L, Zhang T (2011) Efficient optimal learning for contextual bandits. Conference on Uncertainty in Artificial Intelligence , 169–178
2011
Cited alongside, same era.
Tao T (2011) An introduction to measure theory , volume 126 (American Mathematical Society)
2011
Cited alongside, same era.
2011
Cited alongside, same era.
Agarwal A, Dudík M, Kale S, Langford J, Schapire R (2012) Contextual bandit learning with predictable rewards. International Conference on Artificial Intelligence and Statistics , 19–26
2012
Cited alongside, same era.
Agrawal S, Goyal N (2013) Thompson sampling for contextual bandits with linear payoffs. International Conference on Machine Learning , 127–135
Foster D, Agarwal A, Dudik M, Luo H, Schapire R (2018) Practical contextual bandits with regression oracles. International Conference on Machine Learning , 1539–1548
2018
Later among the works it cites.
Russo D, Van Roy B, Kazerouni A, Osband I, Wen Z (2018) A tutorial on thompson sampling. Foundations and Trends in Machine Learning 11(1):1–96
2018
Later among the works it cites.
Gao Z, Han Y, Ren Z, Zhou Z (2019) Batched multi-armed bandits problem. Advances in Neural Information Processing Systems , 503–513
2019
Later among the works it cites.
Krishnamurthy A, Agarwal A, Huang TK, Daumé III H, Langford J (2019) Active learning for cost-sensitive classification. Journal of Machine Learning Research 20:1–50
2019
Later among the works it cites.
Li Y, Wang Y, Zhou Y (2019) Nearly minimax-optimal regret for linearly parameterized bandits. Conference on Learning Theory , 2173–2174
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2013
Cited alongside, same era.
Agarwal A, Hsu D, Kale S, Langford J, Li L, Schapire R (2014) Taming the monster: A fast and simple algorithm for contextual bandits. International Conference on Machine Learning , 1638–1646
2014
Cited alongside, same era.
Cesa-Bianchi N, Gentile C, Mansour Y (2014) Regret minimization for reserve prices in second-price auctions. IEEE Transactions on Information Theory 61(1):549–564
2014
Cited alongside, same era.
Mendelson S (2014) Learning without concentration. Conference on Learning Theory , 25–39
2014
Cited alongside, same era.
Rakhlin A, Sridharan K (2014) Online non-parametric regression. Conference on Learning Theory , 1232–1264
2014
Cited alongside, same era.
2016
Cited alongside, same era.
Agrawal S, Devanur NR (2016) Linear contextual bandits with knapsacks. Advances in Neural Information Processing Systems , 3458–3467
2016
Cited alongside, same era.
Hazan E, Koren T (2016) The computational power of optimization in online learning. Annual ACM Symposium on Theory of Computing , 128–141
2016
Cited alongside, same era.
2019
Later among the works it cites.
Slivkins A (2019) Introduction to multi-armed bandits. Foundations and Trends in Machine Learning 12(1-2):1–286
2019
Later among the works it cites.
Bastani H, Bayati M (2020) Online decision making with high-dimensional covariates. Operations Research 68(1):276–294
2020
Closest in time.
Foster D, Rakhlin A (2020) Beyond ucb: Optimal and efficient contextual bandits with regression oracles. International Conference on Machine Learning , 3199–3210
2020
Closest in time.
Lattimore T, Szepesvári C (2020) Bandit algorithms (Cambridge University Press)
2020
Closest in time.
Lattimore T, Szepesvari C, Weisz G (2020) Learning with good feature representations in bandits and in rl with a generative model. International Conference on Machine Learning , 5662–5670
2020
Closest in time.
Zhou D, Li L, Gu Q (2020) Neural contextual bandits with ucb-based exploration. International Conference on Machine Learning , 11492–11502
2020
Closest in time.
Farrell MH, Liang T, Misra S (2021) Deep neural networks for estimation and inference. Econometrica 89(1):181–213
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
Li L, Lu Y, Zhou D (2017) Provably optimal algorithms for generalized linear contextual bandits. International Conference on Machine Learning , 2071–2080
2080
Closest in time.