Fetching the paper…
Reading the bibliography…
Motivated by practical needs such as large-scale learning, we study the impact of adaptivity constraints to linear contextual bandits, a central problem in online active learning.
The equivalence of two extremum problems
J. Kiefer and J. Wolfowitz · 1960
Earlier work this paper cites.
Probability inequalities for sums of bounded random variables
Wassily Hoeffding · 1963
Earlier work this paper cites.
Algorithmic complexity: three np-hard problems in computational statistics
William J. Welch · 1982
Earlier work this paper cites.
Associative reinforcement learning using linear probabilistic concepts
Naoki Abe and Philip M. Long · 1999
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
Peter Auer, Nicolo Cesa-Bianchi, and Paul Fischer · 2002
Earlier work this paper cites.
Reinforcement learning with immediate rewards and linear hypotheses
Naoki Abe, Alan W. Biermann, and Philip M. Long · 2003
Earlier work this paper cites.
Using confidence bounds for exploitation-exploration trade-offs
Peter Auer · 2003
Earlier work this paper cites.
Optimal Design of Experiments
Friedrich Pukelsheim · 2006
Earlier work this paper cites.
Optimum experimental designs, with SAS , volume 34
Anthony Atkinson, Alexander Donev, and Randall Tobias · 2007
Earlier work this paper cites.
Stochastic linear optimization under bandit feedback
Varsha Dani, Thomas P Hayes, and Sham M Kakade · 2008
Earlier work this paper cites.
On selecting a maximum volume sub-matrix of a matrix and related problems
Ali Çivril and Malik Magdon-Ismail · 2009
Earlier work this paper cites.
Linearly parameterized bandits
Paat Rusmevichientong and John N. Tsitsiklis · 2010
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Yasin Abbasi-yadkori, Dávid Pál, and Csaba Szepesvári · 2011
Earlier work this paper cites.
Contextual bandits with linear payoff functions
Wei Chu, Lihong Li, Lev Reyzin, and Robert Schapire · 2011
Earlier work this paper cites.
User-friendly tail bounds for sums of random matrices
Joel A Tropp · 2012
Earlier work this paper cites.
Online learning with switching costs and other adaptive adversaries
Nicolò Cesa-Bianchi, Ofer Dekel, and Ohad Shamir · 2013
Cited alongside, same era.
Distributed exploration in multi-armed bandits
Eshcar Hillel, Zohar S Karnin, Tomer Koren, Ronny Lempel, and Oren Somekh · 2013
Cited alongside, same era.
Bandits with switching costs: T2/3 regret
Ofer Dekel, Jian Ding, Tomer Koren, and Yuval Peres · 2014
Cited alongside, same era.
Best-arm identification in linear bandits
Marta Soare, Alessandro Lazaric, and Remi Munos · 2014
Cited alongside, same era.
Concurrent pac rl
Zhaohan Guo and Emma Brunskill · 2015
Cited alongside, same era.
On largest volume simplices and sub-determinants
Marco Di Summa, Friedrich Eisenbrand, Yuri Faenza, and Carsten Moldenhauer · 2015
Cited alongside, same era.
Provably efficient q-learning with low switching cost
Yu Bai, Tengyang Xie, Nan Jiang, and Yu-Xiang Wang · 2019
Later among the works it cites.
Regret bounds for batched bandits
Hossein Esfandiari, Amin Karbasi, Abbas Mehrabian, and Vahab Mirrokni · 2019
Later among the works it cites.
Batched multi-armed bandits problem
Zijun Gao, Yanjun Han, Zhimei Ren, and Zhengqing Zhou · 2019
Later among the works it cites.
Nearly minimax-optimal regret for linearly parameterized bandits
Yingkai Li, Yining Wang, and Yuan Zhou · 2019
Later among the works it cites.
Combinatorial algorithms for optimal design
Vivek Madan, Mohit Singh, Uthaipon Tantipongpipat, and Weijun Xie · 2019
Later among the works it cites.
Phase transitions and cyclic phenomena in bandits with switching constraints
David Simchi-Levi and Yunzong Xu · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Top arm identification in multi-armed bandits with batch arm pulls
Kwang-Sung Jun, Kevin Jamieson, Robert Nowak, and Xiaojin Zhu · 2016
Cited alongside, same era.
Batched bandit problems
Vianney Perchet, Philippe Rigollet, Sylvain Chassang, and Erik Snowberg · 2016
Cited alongside, same era.
Learning with limited rounds of adaptivity: Coin tossing, multi-armed bandits, and ranking from pairwise comparisons
Arpit Agarwal, Shivani Agarwal, Sepehr Assadi, and Sanjeev Khanna · 2017
Cited alongside, same era.
On computationally tractable selection of experiments in measurement-constrained regression models
Yining Wang, Adams Wei Yu, and Aarti Singh · 2017
Cited alongside, same era.
Minimax bounds on stochastic batched convex optimization
John Duchi, Feng Ruan, and Chulhee Yun · 2018
Cited alongside, same era.
Approximate positive correlated distributions and approximation algorithms for d-optimal design
Mohit Singh and Weijun Xie · 2018
Cited alongside, same era.
Later among the works it cites.
Collaborative learning with limited interaction: Tight bounds for distributed exploration in multi-armed bandits
Chao Tao, Qin Zhang, and Yuan Zhou · 2019
Later among the works it cites.
Near-optimal discrete optimization for experimental design: A regret minimization approach
Zeyuan Allen-Zhu, Yuanzhi Li, Aarti Singh, and Yining Wang · 2020
Closest in time.
Multinomial logit bandit with low switching cost
Kefan Dong, Yingkai Li, Qin Zhang, and Yuan Zhou · 2020
Closest in time.
Sequential batch learning in finite-action linear contextual bandits
Yanjun Han, Zhengqing Zhou, Zhengyuan Zhou, Jose Blanchet, Peter W Glynn, and Yinyu Ye · 2020
Closest in time.
Collaborative top distribution identifications with limited interaction (extended abstract)
Nikolai Karpov, Qin Zhang, and Yuan Zhou · 2020
Closest in time.
Bandit Algorithms
Tor Lattimore and Csaba Szepesvári · 2020
Closest in time.
Bypassing the monster: A faster and simpler optimal algorithm for contextual bandits under realizability
David Simchi-Levi and Yunzong Xu · 2020
Closest in time.
Almost optimal model-free reinforcement learning via reference-advantage decomposition
Zihan Zhang, Yuan Zhou, and Xiangyang Ji · 2020
Closest in time.