Fetching the paper…
Reading the bibliography…
Motivated by the wide adoption of reinforcement learning (RL) in real-world personalized services, where users' sensitive and private information needs to be protected, we study regret minimization in finite-horizon Markov decision processes (MDPs) under the constraints of differential privacy (DP).
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
Linear least-squares algorithms for temporal difference learning
Steven J Bradtke and Andrew G Barto · 1996
Earlier work this paper cites.
Least-squares temporal difference learning
Justin A Boyan · 1999
Earlier work this paper cites.
Actor-critic algorithms
Vijay R Konda and John N Tsitsiklis · 2000
Earlier work this paper cites.
Adaptive estimation of a quadratic functional by model selection
Beatrice Laurent and Pascal Massart · 2000
Earlier work this paper cites.
A natural policy gradient
Sham M Kakade · 2001
Earlier work this paper cites.
The concentration of measure phenomenon
Michel Ledoux · 2001
Earlier work this paper cites.
Differential privacy: A survey of results
Cynthia Dwork · 2008
Earlier work this paper cites.
Reinforcement learning design for cancer clinical trials
Yufan Zhao, Michael R Kosorok, and Donglin Zeng · 2009
Earlier work this paper cites.
A contextual-bandit approach to personalized news article recommendation
Lihong Li, Wei Chu, John Langford, and Robert E Schapire · 2010
Earlier work this paper cites.
Introduction to the non-asymptotic analysis of random matrices
Roman Vershynin · 2010
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Yasin Abbasi-Yadkori, Dávid Pál, and Csaba Szepesvári · 2011
Earlier work this paper cites.
Private and continual release of statistics
T-H Hubert Chan, Elaine Shi, and Dawn Song · 2011
Earlier work this paper cites.
(nearly) optimal algorithms for private online learning in full-information and bandit settings
Abhradeep Guha Thakurta and Adam Smith · 2013
Earlier work this paper cites.
Local privacy and statistical minimax rates
John C Duchi, Michael I Jordan, and Martin J Wainwright · 2013
Earlier work this paper cites.
The algorithmic foundations of differential privacy
Cynthia Dwork, Aaron Roth, et al · 2014
Earlier work this paper cites.
Mechanism design in large games: Incentives and privacy
Michael Kearns, Mallesh Pai, Aaron Roth, and Jonathan Ullman · 2014
Earlier work this paper cites.
(nearly) optimal differentially private stochastic multi-arm bandits
Nikita Mishra and Abhradeep Thakurta · 2015
Earlier work this paper cites.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Earlier work this paper cites.
Affective personalization of a social robot tutor for children’s second language skills
Goren Gordon, Samuel Spaulding, Jacqueline Kory Westlund, Jin Joo Lee, Luke Plummer, Marayna Martinez, Madhurima Das, and Cynthia Breazeal · 2016
Earlier work this paper cites.
Algorithms for differentially private multi-armed bandits
Aristide Tossou and Christos Dimitrakakis · 2016
Earlier work this paper cites.
Benchmarking deep reinforcement learning for continuous control
Yan Duan, Xi Chen, Rein Houthooft, John Schulman, and Pieter Abbeel · 2016
Cited alongside, same era.
Differentially private policy evaluation
Borja Balle, Maziar Gomrokchi, and Doina Precup · 2016
Cited alongside, same era.
Private matchings and allocations
Justin Hsu, Zhiyi Huang, Aaron Roth, Tim Roughgarden, and Zhiwei Steven Wu · 2016
Cited alongside, same era.
Concentrated differential privacy: Simplifications, extensions, and lower bounds
Mark Bun and Thomas Steinke · 2016
Cited alongside, same era.
Literature survey of statistical, deep and reinforcement learning in natural language processing
Akanksha Rai Sharma and Pranav Kaushik · 2017
Cited alongside, same era.
Achieving privacy in the adversarial multi-armed bandit
Aristide Tossou and Christos Dimitrakakis · 2017
Local differential privacy for bayesian optimization
Xingyu Zhou and Jian Tan · 2020
Later among the works it cites.
Private reinforcement learning with pac and regret guarantees
Giuseppe Vietri, Borja Balle, Akshay Krishnamurthy, and Steven Wu · 2020
Later among the works it cites.
Local differentially private regret minimization in reinforcement learning
Evrard Garcelon, Vianney Perchet, Ciara Pike-Burke, and Matteo Pirotta · 2020
Later among the works it cites.
Provably efficient exploration in policy optimization
Qi Cai, Zhuoran Yang, Chi Jin, and Zhaoran Wang · 2020
Later among the works it cites.
Optimistic policy optimization with bandit feedback
Yonathan Efroni, Lior Shani, Aviv Rosenberg, and Shie Mannor · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Mastering the game of go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, et al · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
The price of differential privacy for online learning
Naman Agarwal and Karan Singh · 2017
Cited alongside, same era.
Differentially private contextual linear bandits
Roshan Shariff and Or Sheffet · 2018
Cited alongside, same era.
Deep reinforcement learning for nlp
William Yang Wang, Jiwei Li, and Xiaodong He · 2018
Cited alongside, same era.
Differential privacy for multi-armed bandits: What is it and what is its cost?
Debabrota Basu, Christos Dimitrakakis, and Aristide Tossou · 2019
Cited alongside, same era.
Multi-armed bandits with local differential privacy
Wenbo Ren, Xingyu Zhou, Jia Liu, and Ness B Shroff · 2020
Later among the works it cites.
Locally differentially private (contextual) bandits learning
Kai Zheng, Tianle Cai, Weiran Huang, Zhenguo Li, and Liwei Wang · 2020
Later among the works it cites.
Locally private distributed reinforcement learning
Hajime Ono and Tsubasa Takahashi · 2020
Later among the works it cites.
Model-based reinforcement learning with value-targeted regression
Alex Ayoub, Zeyu Jia, Csaba Szepesvari, Mengdi Wang, and Lin Yang · 2020
Later among the works it cites.
Model-based reinforcement learning with value-targeted regression
Zeyu Jia, Lin Yang, Csaba Szepesvari, and Mengdi Wang · 2020
Later among the works it cites.
Provably efficient reinforcement learning with linear function approximation
Chi Jin, Zhuoran Yang, Zhaoran Wang, and Michael I Jordan · 2020
Later among the works it cites.
Learning with good feature representations in bandits and in rl with a generative model
Tor Lattimore, Csaba Szepesvari, and Gellert Weisz · 2020
Later among the works it cites.
On function approximation in reinforcement learning: Optimism in the face of large state spaces
Zhuoran Yang, Chi Jin, Zhaoran Wang, Mengdi Wang, and Michael I Jordan · 2020
Later among the works it cites.
Differentially private regret minimization in episodic markov decision processes
Sayak Ray Chowdhury and Xingyu Zhou · 2021
Later among the works it cites.
Optimal algorithms for private online learning in a stochastic environment
Bingshan Hu, Zhiming Huang, and Nishant A Mehta · 2021
Later among the works it cites.
No-regret algorithms for private gaussian process bandit optimization
Abhimanyu Dubey · 2021
Later among the works it cites.
Locally differentially private reinforcement learning for linear mixture markov decision processes
Chonghua Liao, Jiafan He, and Quanquan Gu · 2021
Later among the works it cites.
Differentially private exploration in reinforcement learning with linear representation
Paul Luyo, Evrard Garcelon, Alessandro Lazaric, and Matteo Pirotta · 2021
Later among the works it cites.
Provably efficient safe exploration via primal-dual policy optimization
Dongsheng Ding, Xiaohan Wei, Zhuoran Yang, Zhaoran Wang, and Mihailo Jovanovic · 2021
Later among the works it cites.
Adaptive control of differentially private linear quadratic systems
Sayak Ray Chowdhury, Xingyu Zhou, and Ness Shroff · 2021
Later among the works it cites.