Fetching the paper…
Reading the bibliography…
We study regret minimization in finite horizon tabular Markov decision processes (MDPs) under the constraints of differential privacy (DP).
Locally differentially private data collection and analysis
Teng Wang, Jun Zhao, Xinyu Yang, and Xuebin Ren · 1906
Earlier work this paper cites.
Neural policy gradient methods: Global optimality and rates of convergence
Lingxiao Wang, Qi Cai, Zhuoran Yang, and Zhaoran Wang · 1909
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Richard S Sutton · 1988
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
Markov Decision Processes — Discrete Stochastic Dynamic Programming
Martin L. Puterman · 1994
Earlier work this paper cites.
Actor-critic algorithms
Vijay R Konda and John N Tsitsiklis · 2000
Earlier work this paper cites.
A natural policy gradient
Sham M Kakade · 2001
Earlier work this paper cites.
Optimistic policy optimization with bandit feedback
Yonathan Efroni, Lior Shani, Aviv Rosenberg, and Shie Mannor · 2002
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
P. Auer, N. Cesa-Bianchi, and P. Fischer · 2002
Earlier work this paper cites.
Mirror descent and nonlinear projected subgradient methods for convex optimization
Amir Beck and Marc Teboulle · 2003
Earlier work this paper cites.
Exploration-exploitation in constrained mdps
Yonathan Efroni, Shie Mannor, and Matteo Pirotta · 2003
Earlier work this paper cites.
Inequalities for the L1
Tsachy Weissman, Erik Ordentlich, Gadiel Seroussi, Sergio Verdu, and Marcelo J Weinberger · 2003
Earlier work this paper cites.
Differential privacy: A survey of results
Cynthia Dwork · 2008
Earlier work this paper cites.
An analysis of model-based interval estimation for M
Alexander L Strehl and Michael L Littman · 2008
Earlier work this paper cites.
Empirical Bernstein bounds and sample variance penalization
A Maurer and M Pontil · 2009
Earlier work this paper cites.
A contextual-bandit approach to personalized news article recommendation
Lihong Li, Wei Chu, John Langford, and Robert E Schapire · 2010
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Thomas Jaksch, Ronald Ortner, and Peter Auer · 2010
Earlier work this paper cites.
Private and continual release of statistics
TH Hubert Chan, Elaine Shi, and Dawn Song · 2010
Earlier work this paper cites.
Differentially private online learning
Prateek Jain, Pravesh Kothari, and Abhradeep Thakurta · 2012
Earlier work this paper cites.
Local privacy and statistical minimax rates
John C Duchi, Michael I Jordan, and Martin J Wainwright · 2013
Earlier work this paper cites.
(nearly) optimal algorithms for private online learning in full-information and bandit settings
Abhradeep Guha Thakurta and Adam Smith · 2013
Earlier work this paper cites.
(more) efficient reinforcement learning via posterior sampling
Ian Osband, Dan Russo, and Benjamin Van Roy · 2013
Cited alongside, same era.
Mechanism design in large games: Incentives and privacy
Michael Kearns, Mallesh Pai, Aaron Roth, and Jonathan Ullman · 2014
Cited alongside, same era.
The algorithmic foundations of differential privacy
Cynthia Dwork, Aaron Roth, et al · 2014
Cited alongside, same era.
(nearly) optimal differentially private stochastic multi-arm bandits
Nikita Mishra and Abhradeep Thakurta · 2015
Cited alongside, same era.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Cited alongside, same era.
Deep reinforcement learning for dialogue generation
Jiwei Li, Will Monroe, Alan Ritter, Michel Galley, Jianfeng Gao, and Dan Jurafsky · 2016
Guidelines for reinforcement learning in healthcare
Omer Gottesman, Fredrik Johansson, Matthieu Komorowski, Aldo Faisal, David Sontag, Finale Doshi-Velez, and Leo Anthony Celi · 2019
Later among the works it cites.
Neural proximal/trust region policy optimization attains globally optimal policy
Boyi Liu, Qi Cai, Zhuoran Yang, and Zhaoran Wang · 2019
Later among the works it cites.
An optimal private stochastic-mab algorithm based on optimal private stopping rule
Touqir Sajed and Or Sheffet · 2019
Later among the works it cites.
Privacy-preserving q-learning with functional noise in continuous state spaces
Baoxiang Wang and Nidhi Hegde · 2019
Later among the works it cites.
Tight regret bounds for model-based reinforcement learning with greedy policies
Yonathan Efroni, Nadav Merlis, Mohammad Ghavamzadeh, and Shie Mannor · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Affective personalization of a social robot tutor for children’s second language skills
Goren Gordon, Samuel Spaulding, Jacqueline Kory Westlund, Jin Joo Lee, Luke Plummer, Marayna Martinez, Madhurima Das, and Cynthia Breazeal · 2016
Cited alongside, same era.
Algorithms for differentially private multi-armed bandits
Aristide Tossou and Christos Dimitrakakis · 2016
Cited alongside, same era.
Benchmarking deep reinforcement learning for continuous control
Yan Duan, Xi Chen, Rein Houthooft, John Schulman, and Pieter Abbeel · 2016
Cited alongside, same era.
Differentially private policy evaluation
Borja Balle, Maziar Gomrokchi, and Doina Precup · 2016
Cited alongside, same era.
Private matchings and allocations
Justin Hsu, Zhiyi Huang, Aaron Roth, Tim Roughgarden, and Zhiwei Steven Wu · 2016
Cited alongside, same era.
Discrete distribution estimation under local privacy
Peter Kairouz, Keith Bonawitz, and Daniel Ramage · 2016
Cited alongside, same era.
Later among the works it cites.
Tighter problem-dependent regret bounds in reinforcement learning without domain knowledge using value function bounds
Andrea Zanette and Emma Brunskill · 2019
Later among the works it cites.
Provably efficient reinforcement learning with linear function approximation
Chi Jin, Zhuoran Yang, Zhaoran Wang, and Michael I Jordan · 2019
Later among the works it cites.
Online learning in kernelized markov decision processes
Sayak Ray Chowdhury and Aditya Gopalan · 2019
Later among the works it cites.
Private reinforcement learning with pac and regret guarantees
Giuseppe Vietri, Borja Balle, Akshay Krishnamurthy, and Steven Wu · 2020
Later among the works it cites.
Local differentially private regret minimization in reinforcement learning
Evrard Garcelon, Vianney Perchet, Ciara Pike-Burke, and Matteo Pirotta · 2020
Later among the works it cites.
Multi-armed bandits with local differential privacy
Wenbo Ren, Xingyu Zhou, Jia Liu, and Ness B Shroff · 2020
Later among the works it cites.
Locally differentially private (contextual) bandits learning
Kai Zheng, Tianle Cai, Weiran Huang, Zhenguo Li, and Liwei Wang · 2020
Later among the works it cites.
Local differential privacy for bayesian optimization
Xingyu Zhou and Jian Tan · 2020
Later among the works it cites.
Provably efficient exploration in policy optimization
Qi Cai, Zhuoran Yang, Chi Jin, and Zhaoran Wang · 2020
Later among the works it cites.
(locally) differentially private combinatorial semi-bandits
Xiaoyu Chen, Kai Zheng, Zixin Zhou, Yunchang Yang, Wei Chen, and Liwei Wang · 2020
Later among the works it cites.
Locally private distributed reinforcement learning
Hajime Ono and Tsubasa Takahashi · 2020
Later among the works it cites.
Model-based reinforcement learning with value-targeted regression
Alex Ayoub, Zeyu Jia, Csaba Szepesvari, Mengdi Wang, and Lin F Yang · 2020
Later among the works it cites.
Learning adversarial markov decision processes with bandit feedback and unknown transition
Chi Jin, Tiancheng Jin, Haipeng Luo, Suvrit Sra, and Tiancheng Yu · 2020
Later among the works it cites.
No-regret algorithms for private gaussian process bandit optimization
Abhimanyu Dubey · 2021
Closest in time.
Adaptive control of differentially private linear quadratic systems
Sayak Ray Chowdhury, Xingyu Zhou, and Ness Shroff · 2021
Closest in time.
Optimal algorithms for private online learning in a stochastic environment
Bingshan Hu, Zhiming Huang, and Nishant A Mehta · 2021
Closest in time.