Fetching the paper…
Reading the bibliography…
We study the contextual linear bandit problem, a version of the standard stochastic multi-armed bandit (MAB) problem where a learner sequentially selects actions to maximize a reward which depends also on a user provided per-round context.
Some aspects of the sequential design of experiments
Herbert Robbins · 1952
Earlier work this paper cites.
Bandit problems: sequential allocation of experiments
Donald A Berry and Bert Fristedt · 1985
Earlier work this paper cites.
Sample mean based index policies with O(log n) regret for the multi-armed bandit problem. , volume 27, pages 1054–1078
Rajeev Agrawal · 1995
Earlier work this paper cites.
Adaptive estimation of a quadratic functional by model selection
B. Laurent and P. Massart · 2000
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
Peter Auer, Nicolò Cesa-Bianchi, and Paul Fischer · 2002
Earlier work this paper cites.
Reinforcement learning with immediate rewards and linear hypotheses
Naoki Abe, Alan W. Biermann, and Philip M. Long · 2003
Earlier work this paper cites.
Using confidence bounds for exploitation-exploration trade-offs
Peter Auer · 2003
Earlier work this paper cites.
Our data, ourselves: Privacy via distributed noise generation
Cynthia Dwork, Krishnaram Kenthapadi, Frank McSherry, Ilya Mironov, and Moni Naor · 2006
Earlier work this paper cites.
Stochastic linear optimization under bandit feedback
Varsha Dani, Thomas Hayes, and Sham Kakade · 2008
Earlier work this paper cites.
Private and continual release of statistics
T.-H. Hubert Chan, Elaine Shi, and Dawn Song · 2010
Earlier work this paper cites.
Boosting and differential privacy
C. Dwork, G. N. Rothblum, and S. Vadhan · 2010
Earlier work this paper cites.
On the geometry of differential privacy
Moritz Hardt and Kunal Talwar · 2010
Earlier work this paper cites.
Linearly parameterized bandits
Paat Rusmevichientong and John N. Tsitsiklis · 2010
Cited alongside, same era.
Introduction to the non-asymptotic analysis of random matrices
Roman Vershynin · 2010
Cited alongside, same era.
Improved algorithms for linear stochastic bandits
Yasin Abbasi-Yadkori, Dávid Pál, and Csaba Szepesvári · 2011
Cited alongside, same era.
Differentially private empirical risk minimization
Kamalika Chaudhuri, Claire Monteleoni, and Anand D. Sarwate · 2011
Cited alongside, same era.
Contextual bandits with linear payoff functions
Wei Chu, Lihong Li, Lev Reyzin, and Robert E. Schapire · 2011
Cited alongside, same era.
Matrix Theory: Basic Results and Techniques
Fuzhen Zhang · 2011
Cited alongside, same era.
The algorithmic foundations of differential privacy
Cynthia Dwork and Aaron Roth · 2014
Later among the works it cites.
Analyze Gauss: Optimal bounds for privacy-preserving principal component analysis
Cynthia Dwork, Kunal Talwar, Abhradeep Thakurta, and Li Zhang · 2014
Later among the works it cites.
Mechanism design in large games: Incentives and privacy
Michael Kearns, Mallesh Pai, Aaron Roth, and Jonathan Ullman · 2014
Later among the works it cites.
(Nearly) optimal differentially private stochastic multi-arm bandits
Nikita Mishra and Abhradeep Thakurta · 2015
Later among the works it cites.
Private approximations of the 2nd-moment matrix using existing techniques in linear regression
Or Sheffet · 2015
Later among the works it cites.
Concentrated differential privacy: Simplifications, extensions, and lower bounds
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Online-to-confidence-set conversions and application to sparse stochastic bandits
Yasin Abbasi-Yadkori, Dávid Pál, and Csaba Szepesvári · 2012
Cited alongside, same era.
Differentially private online learning
Prateek Jain, Pravesh Kothari, and Abhradeep Thakurta · 2012
Cited alongside, same era.
Topics in Random Matrix Theory , volume 132
Terence Tao · 2012
Cited alongside, same era.
(Nearly) optimal algorithms for private online learning in full-information and bandit settings
Adam Smith and Abhradeep Thakurta · 2013
Cited alongside, same era.
Private empirical risk minimization: Efficient algorithms and tight error bounds
Raef Bassily, Adam Smith, and Abhradeep Thakurta · 2014
Cited alongside, same era.
Calibrating noise to sensitivity in private data analysis
Cynthia Dwork, Frank McSherry, Kobbi Nissim, and Adam Smith
Cited in the paper.
Mark Bun and Thomas Steinke · 2016
Later among the works it cites.
Algorithms for differentially private multi-armed bandits
Aristide C. Y. Tossou and Christos Dimitrakakis · 2016
Later among the works it cites.
Finite sample differentially private confidence intervals
Vishesh Karwa and Salil Vadhan · 2017
Later among the works it cites.
The end of optimism? an asymptotic analysis of finite-armed linear bandits
Tor Lattimore and Csaba Szepesvári · 2017
Later among the works it cites.
Achieving privacy in the adversarial multi-armed bandit
Aristide C. Y. Tossou and Christos Dimitrakakis · 2017
Later among the works it cites.
Mitigating bias in adaptive data gathering via differential privacy
Seth Neel and Aaron Roth · 2018
Closest in time.