Fetching the paper…
Reading the bibliography…
Reinforcement learning algorithms are widely used in domains where it is desirable to provide a personalized service.
Differentially private summation with multi-message shuffling
Borja Balle, James Bell, Adrià Gascón, and Kobbi Nissim · 1906
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Martin L. Puterman · 1994
Earlier work this paper cites.
The nonstochastic multiarmed bandit problem
Peter Auer, Nicolo Cesa-Bianchi, Yoav Freund, and Robert E Schapire · 2002
Earlier work this paper cites.
Inequalities for the l1 deviation of the empirical distribution
Tsachy Weissman, Erik Ordentlich, Gadiel Seroussi, Sergio Verdu, and Marcelo J Weinberger · 2003
Earlier work this paper cites.
Calibrating noise to sensitivity in private data analysis
Cynthia Dwork, Frank McSherry, Kobbi Nissim, and Adam D. Smith · 2006
Earlier work this paper cites.
Separating local & shuffled differential privacy via histograms
Victor Balcer and Albert Cheu · 2006
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Thomas Jaksch, Ronald Ortner, and Peter Auer · 2010
Earlier work this paper cites.
A bounded-noise mechanism for differential privacy
Yuval Dagan and Gil Kur · 2012
Earlier work this paper cites.
Local privacy, data processing inequalities, and statistical minimax rates
John C Duchi, Michael I Jordan, and Martin J Wainwright · 2013
Earlier work this paper cites.
(more) efficient reinforcement learning via posterior sampling
Ian Osband, Daniel Russo, and Benjamin Van Roy · 2013
Earlier work this paper cites.
Geo-indistinguishability: differential privacy for location-based systems
Miguel E. Andrés, Nicolás Emilio Bordenabe, Konstantinos Chatzikokolakis, and Catuscia Palamidessi · 2013
Earlier work this paper cites.
Rappor: Randomized aggregatable privacy-preserving ordinal response
Úlfar Erlingsson, Vasyl Pihur, and Aleksandra Korolova · 2014
Earlier work this paper cites.
The algorithmic foundations of differential privacy
Cynthia Dwork and Aaron Roth · 2014
Earlier work this paper cites.
(nearly) optimal differentially private stochastic multi-arm bandits
Nikita Mishra and Abhradeep Thakurta · 2015
Earlier work this paper cites.
Local, private, efficient protocols for succinct histograms
Raef Bassily and Adam Smith · 2015
Earlier work this paper cites.
Differentially private multi-agent multi-armed bandits
Aristide Tossou and Christos Dimitrakakis · 2015
Earlier work this paper cites.
Algorithms for differentially private multi-armed bandits
Aristide C. Y. Tossou and Christos Dimitrakakis · 2016
Earlier work this paper cites.
Differentially private policy evaluation
Borja Balle, Maziar Gomrokchi, and Doina Precup · 2016
Earlier work this paper cites.
Discrete distribution estimation under local privacy
Peter Kairouz, Keith Bonawitz, and Daniel Ramage · 2016
Cited alongside, same era.
Minimax optimal procedures for locally private estimation
John C. Duchi, Martin J. Wainwright, and Michael I. Jordan · 2016
Cited alongside, same era.
Unifying pac and regret: Uniform pac bounds for episodic reinforcement learning
Christoph Dann, Tor Lattimore, and Emma Brunskill · 2017
Cited alongside, same era.
Minimax regret bounds for reinforcement learning
Mohammad Gheshlaghi Azar, Ian Osband, and Rémi Munos · 2017
Cited alongside, same era.
Differentially private contextual linear bandits
Roshan Shariff and Or Sheffet · 2018
Cited alongside, same era.
Profile-based privacy for locally private computations
Joseph Geumlek and Kamalika Chaudhuri · 2019
Later among the works it cites.
Privacy-preserving multi-party contextual bandits
Awni Y. Hannun, Brian Knott, Shubho Sengupta, and Laurens van der Maaten · 2019
Later among the works it cites.
Real-world video adaptation with reinforcement learning, 2020
Hongzi Mao, Shannon Chen, Drew Dimmery, Shaun Singh, Drew Blaisdell, Yuandong Tian, Mohammad Alizadeh, and Eytan Bakshy · 2020
Closest in time.
Utility/privacy trade-off through the lens of optimal transport
Etienne Boursier and Vianney Perchet · 2020
Closest in time.
Private reinforcement learning with pac and regret guarantees
Giuseppe Vietri, Borja de Balle Pigem, Akshay Krishnamurthy, and Steven Wu · 2020
Closest in time.
Locally differentially private (contextual) bandits learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
John M Abowd · 2018
Cited alongside, same era.
Corrupt bandits for preserving local privacy
Pratik Gajane, Tanguy Urvoy, and Emilie Kaufmann · 2018
Cited alongside, same era.
Optimal schemes for discrete distribution estimation under locally differential privacy
Min Ye and Alexander Barg · 2018
Cited alongside, same era.
A tutorial on thompson sampling
Daniel Russo, Benjamin Van Roy, Abbas Kazerouni, Ian Osband, and Zheng Wen · 2018
Cited alongside, same era.
How you act tells a lot: Privacy-leaking attack on deep reinforcement learning
Xinlei Pan, Weiyao Wang, Xiaoshuai Zhang, Bo Li, Jinfeng Yi, and Dawn Song · 2019
Cited alongside, same era.
Distributed differential privacy via shuffling
Albert Cheu, Adam Smith, Jonathan Ullman, David Zeber, and Maxim Zhilyaev · 2019
Cited alongside, same era.
An optimal private stochastic-mab algorithm based on optimal private stopping rule
Touqir Sajed and Or Sheffet · 2019
Cited alongside, same era.
Kai Zheng, Tianle Cai, Weiran Huang, Zhenguo Li, and Liwei Wang · 2020
Closest in time.
Hiding among the clones: A simple and nearly optimal analysis of privacy amplification by shuffling, 2020
Vitaly Feldman, Audra McMillan, and Kunal Talwar · 2020
Closest in time.
Encode, shuffle, analyze privacy revisited: Formalizations and empirical evaluation
Úlfar Erlingsson, Vitaly Feldman, Ilya Mironov, Ananth Raghunathan, Shuang Song, Kunal Talwar, and Abhradeep Thakurta · 2020
Closest in time.
(locally) differentially private combinatorial semi-bandits
Xiaoyu Chen, Kai Zheng, Zixin Zhou, Yunchang Yang, Wei Chen, and Liwei Wang · 2020
Closest in time.
Multi-armed bandits with local differential privacy
Wenbo Ren, Xingyu Zhou, Jia Liu, and Ness B Shroff · 2020
Closest in time.
Locally private distributed reinforcement learning
Hajime Ono and Tsubasa Takahashi · 2020
Closest in time.
Bandit algorithms
Tor Lattimore and Csaba Szepesvári · 2020
Closest in time.
A unifying view of optimism in episodic reinforcement learning
Gergely Neu and Ciara Pike-Burke · 2020
Closest in time.
Context aware local differential privacy
Jayadev Acharya, Kallista Bonawitz, Peter Kairouz, Daniel Ramage, and Ziteng Sun · 2020
Closest in time.
Privacy-preserving bandits
Mohammad Malekzadeh, Dimitrios Athanasakis, Hamed Haddadi, and Benjamin Livshits · 2020
Closest in time.
Improved analysis of ucrl2 with empirical bernstein inequality
Ronan Fruit, Matteo Pirotta, and Alessandro Lazaric · 2020
Closest in time.
Robo-advising: Enhancing investment with inverse optimization and deep reinforcement learning, 2021
Haoran Wang and Shi Yu · 2021
Closest in time.
On distributed differential privacy and counting distinct elements
Lijie Chen, Badih Ghazi, Ravi Kumar, and Pasin Manurangsi · 2021
Closest in time.