Fetching the paper…
Reading the bibliography…
Reinforcement learning (RL) algorithms can be used to provide personalized services, which rely on users' private and sensitive data.
A short note on concentration inequalities for random vectors with subgaussian norm
Jin, C · 1902
Earlier work this paper cites.
Differential privacy for multi-armed bandits: What is it and what is its cost?
Basu, D · 1905
Earlier work this paper cites.
Optimism in reinforcement learning with generalized linear function approximation
Wang, Y · 1912
Earlier work this paper cites.
Estimation des densités: risque minimax
Bretagnolle, J · 1979
Earlier work this paper cites.
Learning from delayed rewards
Watkins, C. J. C. H · 1989
Earlier work this paper cites.
Elements of information theory
Cover, T. M · 1999
Earlier work this paper cites.
Calibrating noise to sensitivity in private data analysis
Dwork, C · 2006
Earlier work this paper cites.
Locally differentially private (contextual) bandits learning
Zheng, K · 2006
Earlier work this paper cites.
Provably efficient reinforcement learning for discounted mdps with feature mapping
Zhou, D · 2006
Earlier work this paper cites.
An analysis of model-based interval estimation for markov decision processes
Strehl, A. L · 2008
Earlier work this paper cites.
Local differentially private regret minimization in reinforcement learning
Garcelon, E · 2010
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Jaksch, T · 2010
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Abbasi-Yadkori, Y · 2011
Earlier work this paper cites.
Logarithmic regret for reinforcement learning with linear function approximation
He, J · 2011
Cited alongside, same era.
What can we learn privately?
Kasiviswanathan, S. P · 2011
Cited alongside, same era.
Topics in random matrix theory
Tao, T · 2012
Cited alongside, same era.
Nearly minimax optimal reinforcement learning for linear mixture markov decision processes
Zhou, D · 2012
Cited alongside, same era.
Local privacy, data processing inequalities, and statistical minimax rates
Duchi, J. C · 2013
Cited alongside, same era.
Eluder dimension and the sample complexity of optimistic exploration
Membership inference attacks against machine learning models
Shokri, R · 2017
Later among the works it cites.
Minimax optimal procedures for locally private estimation
Duchi, J. C · 2018
Later among the works it cites.
Corrupt bandits for preserving local privacy
Gajane, P · 2018
Later among the works it cites.
Differentially private contextual linear bandits
Shariff, R · 2018
Later among the works it cites.
Sample-optimal parametric q-learning using linearly additive features
Yang, L · 2019
Later among the works it cites.
Model-based reinforcement learning with value-targeted regression
Ayoub, A · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Russo, D · 2013
Cited alongside, same era.
The algorithmic foundations of differential privacy
Dwork, C · 2014
Cited alongside, same era.
Markov decision processes: discrete stochastic dynamic programming
Puterman, M. L · 2014
Cited alongside, same era.
Deep learning with differential privacy
Abadi, M · 2016
Cited alongside, same era.
Differentially private policy evaluation
Balle, B · 2016
Cited alongside, same era.
Minimax regret bounds for reinforcement learning
Azar, M. G · 2017
Cited alongside, same era.
Contextual decision processes with low bellman rank are pac-learnable
Jiang, N · 2017
Cited alongside, same era.
(locally) differentially private combinatorial semi-bandits
Chen, X · 2020
Later among the works it cites.
Model-based reinforcement learning with value-targeted regression
Jia, Z · 2020
Later among the works it cites.
Provably efficient reinforcement learning with linear function approximation
Jin, C · 2020
Later among the works it cites.
Bandit algorithms
Lattimore, T · 2020
Later among the works it cites.
Private reinforcement learning with pac and regret guarantees
Vietri, G · 2020
Later among the works it cites.
Learning near optimal policies with low inherent bellman error
Zanette, A · 2020
Later among the works it cites.