Fetching the paper…
Reading the bibliography…
In this paper, we study differentially private online learning problems in a stochastic environment under both bandit and full information feedback.
A decision-theoretic generalization of online learning and an application to boosting
Yoav Freund and Robert E Schapire · 1997
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
Peter Auer, Nicolo Cesa-Bianchi, and Paul Fischer · 2002
Earlier work this paper cites.
PAC bounds for multi-armed bandit and Markov decision processes
Eyal Even-Dar, Shie Mannor, and Yishay Mansour · 2002
Earlier work this paper cites.
Minimax policies for adversarial and stochastic bandits
Jean-Yves Audibert and Sébastien Bubeck · 2009
Earlier work this paper cites.
Exploration–exploitation tradeoff using variance estimates in multi-armed bandits
Jean-Yves Audibert, Rémi Munos, and Csaba Szepesvári · 2009
Earlier work this paper cites.
UCB revisited: Improved regret bounds for the stochastic multi-armed bandit problem
Peter Auer and Ronald Ortner · 2010
Earlier work this paper cites.
Differential privacy under continual observation
Cynthia Dwork, Moni Naor, Toniann Pitassi, and Guy N Rothblum · 2010
Earlier work this paper cites.
Private and continual release of statistics
T-H Hubert Chan, Elaine Shi, and Dawn Song · 2011
Earlier work this paper cites.
The KL-UCB algorithm for bounded stochastic bandits and beyond
Aurélien Garivier and Olivier Cappé · 2011
Earlier work this paper cites.
Analysis of Thompson Sampling for the multi-armed bandit problem
Shipra Agrawal and Navin Goyal · 2012
Earlier work this paper cites.
Differentially private online learning
Prateek Jain, Pravesh Kothari, and Abhradeep Thakurta · 2012
Earlier work this paper cites.
(Nearly) optimal algorithms for private online learning in full-information and bandit settings
Abhradeep Guha Thakurta and Adam Smith · 2013
Earlier work this paper cites.
The algorithmic foundations of differential privacy
Cynthia Dwork, Aaron Roth, et al · 2014
Earlier work this paper cites.
Follow the leader with dropout perturbations
Tim Van Erven, Wojciech Kotłowski, and Manfred K Warmuth · 2014
Earlier work this paper cites.
Non-asymptotic analysis of a new bandit algorithm for semi-bounded rewards
Junya Honda and Akimichi Takemura · 2015
Cited alongside, same era.
(Nearly) optimal differentially private stochastic multi-arm bandits
Nikita Mishra and Abhradeep Thakurta · 2015
Cited alongside, same era.
Algorithms for differentially private multi-armed bandits
Aristide CY Tossou and Christos Dimitrakakis · 2016
Cited alongside, same era.
The price of differential privacy for online learning
Naman Agarwal and Karan Singh · 2017
Cited alongside, same era.
Near-optimal regret bounds for Thompson Sampling
Shipra Agrawal and Navin Goyal · 2017
Cited alongside, same era.
On minimaxity of follow the leader strategy in the stochastic setting
Wojciech Kotłowski · 2018
Cited alongside, same era.
Bandit algorithms
Tor Lattimore and Csaba Szepesvári · 2020
Later among the works it cites.
Differentially private Assouad, Fano, and Le Cam
Jayadev Acharya, Ziteng Sun, and Huanyu Zhang · 2021
Closest in time.
Optimal algorithms for private online learning in a stochastic environment
Bingshan Hu, Zhiming Huang, and Nishant A Mehta · 2021
Closest in time.
MOTS: Minimax optimal Thompson Sampling
Tianyuan Jin, Pan Xu, Jieming Shi, Xiaokui Xiao, and Quanquan Gu · 2021
Closest in time.
Private online prediction from experts: Separations and faster rates
Hilal Asi, Vitaly Feldman, Tomer Koren, and Kunal Talwar · 2022
Closest in time.
When privacy meets partial information: A refined analysis of differentially private bandits
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Refining the confidence level for optimistic bandit strategies
Tor Lattimore · 2018
Cited alongside, same era.
Differentially private contextual linear bandits
Roshan Shariff and Or Sheffet · 2018
Cited alongside, same era.
Private communication, February 2019
Wojciech Kotłowski · 2019
Cited alongside, same era.
On the optimality of the Hedge algorithm in the stochastic regime
Jaouad Mourtada and Stéphane Gaïffas · 2019
Cited alongside, same era.
An optimal private stochastic-MAB algorithm based on optimal private stopping rule
Touqir Sajed and Or Sheffet · 2019
Cited alongside, same era.
(Locally) differentially private combinatorial semi-bandits
Xiaoyu Chen, Kai Zheng, Zixin Zhou, Yunchang Yang, Wei Chen, and Liwei Wang · 2020
Cited alongside, same era.
Achraf Azize and Debabrota Basu · 2022
Closest in time.
Maillard sampling: Boltzmann exploration done optimally
Jie Bian and Kwang-Sung Jun · 2022
Closest in time.
Near-optimal Thompson Sampling-based algorithms for differentially private stochastic bandits
Bingshan Hu and Nidhi Hegde · 2022
Closest in time.
Finite-time regret of Thompson Sampling algorithms for exponential family multi-armed bandits
Tianyuan Jin, Pan Xu, Xiaokui Xiao, and Anima Anandkumar · 2022
Closest in time.
Differentially private algorithms for efficient online matroid optimization
Kushagra Chandak, Bingshan Hu, and Nidhi Hegde · 2023
Closest in time.
From 6235149080811616882909238708 to 29: Vanilla Thompson Sampling revisited
Bingshan Hu and Tianyue H. Zhang · 2023
Closest in time.
Optimistic Thompson Sampling-based algorithms for episodic reinforcement learning
Bingshan Hu, Tianyue H Zhang, Nidhi Hegde, and Mark Schmidt · 2023
Closest in time.
Thompson Sampling with less exploration is fast and optimal
Tianyuan Jin, Xianglin Yang, Xiaokui Xiao, and Pan Xu · 2023
Closest in time.