Fetching the paper…
Reading the bibliography…
Many reinforcement learning applications involve the use of data that is sensitive, such as medical records of patients or financial information.
Linear least-squares algorithms for temporal difference learning
Bradtke, S. J. and Barto, A. G. (1996) · 1996
Earlier work this paper cites.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G. (1998) · 1998
Earlier work this paper cites.
Nonlinear programming
Bertsekas, D. P. (1999) · 1999
Earlier work this paper cites.
Matrix analysis and applied linear algebra
Meyer, C. D. (2000) · 2000
Earlier work this paper cites.
Clinical data based optimal sti strategies for hiv: a reinforcement learning approach
Ernst, D., Stan, G.-B., Goncalves, J., and Wehenkel, L. (2006) · 2006
Earlier work this paper cites.
Fast gradient-descent methods for temporal-difference learning with linear function approximation
Sutton, R. S., Maei, H. R., Precup, D., Bhatnagar, S., Silver, D., Szepesvári, C., and Wiewiora, E. (2009) · 2009
Earlier work this paper cites.
Distributed optimization and statistical learning via the alternating direction method of multipliers
Boyd, S., Parikh, N., Chu, E., Peleato, B., Eckstein, J., et al. (2011) · 2011
Earlier work this paper cites.
Value function approximation in reinforcement learning using the fourier basis
Konidaris, G., Osentoski, S., and Thomas, P. S. (2011) · 2011
Earlier work this paper cites.
Introductory lectures on convex optimization: A basic course
Nesterov, Y. (2013) · 2013
Earlier work this paper cites.
Stochastic gradient descent with differentially private updates
Song, S., Chaudhuri, K., and Sarwate, A. D. (2013) · 2013
Earlier work this paper cites.
Differentially private empirical risk minimization: Efficient algorithms and tight error bounds
Bassily, R., Smith, A., and Thakurta, A. (2014) · 2014
Cited alongside, same era.
The algorithmic foundations of differential privacy
Dwork, C., Roth, A., et al. (2014) · 2014
Cited alongside, same era.
Weighted importance sampling for off-policy learning with linear function approximation
Mahmood, A. R., Hasselt, H., and Sutton, R. S. (2014) · 2014
Cited alongside, same era.
Markov decision processes: discrete stochastic dynamic programming
Puterman, M. L. (2014) · 2014
Cited alongside, same era.
Finite-sample analysis of proximal gradient td algorithms
Liu, B., Liu, J., Ghavamzadeh, M., Mahadevan, S., and Petrik, M. (2015) · 2015
Cited alongside, same era.
Personalized ad recommendation systems for life-time value optimization with guarantees
Concentrated differential privacy
Dwork, C. and Rothblum, G. N. (2016) · 2016
Later among the works it cites.
Stochastic variance reduction methods for saddle-point problems
Palaniappan, B. and Bach, F. (2016) · 2016
Later among the works it cites.
Semi-supervised knowledge transfer for deep learning from private training data
Papernot, N., Abadi, M., Erlingsson, U., Goodfellow, I., and Talwar, K. (2016) · 2016
Later among the works it cites.
Stochastic variance reduction methods for policy evaluation
Du, S. S., Chen, J., Li, L., Xiao, L., and Zhou, D. (2017) · 2017
Later among the works it cites.
Mironov, I. (2017) · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Theocharous, G., Thomas, P. S., and Ghavamzadeh, M. (2015) · 2015
Cited alongside, same era.
Deep learning with differential privacy
Abadi, M., Chu, A., Goodfellow, I., McMahan, H. B., Mironov, I., Talwar, K., and Zhang, L. (2016) · 2016
Cited alongside, same era.
Differentially private policy evaluation
Balle, B., Gomrokchi, M., and Precup, D. (2016) · 2016
Cited alongside, same era.
Concentrated differential privacy: Simplifications, extensions, and lower bounds
Bun, M. and Steinke, T. (2016) · 2016
Cited alongside, same era.
Our data, ourselves: Privacy via distributed noise generation
Dwork, C., Kenthapadi, K., McSherry, F., Mironov, I., and Naor, M. (2006a)
Cited in the paper.
Calibrating noise to sensitivity in private data analysis
Dwork, C., McSherry, F., Nissim, K., and Smith, A. (2006b)
Cited in the paper.
Later among the works it cites.
Deep reinforcement learning for sepsis treatment
Raghu, A., Komorowski, M., Ahmed, I., Celi, L., Szolovits, P., and Ghassemi, M. (2017) · 2017
Later among the works it cites.
Differentially private empirical risk minimization revisited: Faster and more general
Wang, D., Ye, M., and Xu, J. (2017) · 2017
Later among the works it cites.
Privacy amplification by subsampling: Tight analyses via couplings and divergences
Balle, B., Barthe, G., and Gaboardi, M. (2018) · 2018
Later among the works it cites.
Subsampled r \ \backslash ’enyi differential privacy and analytical moments accountant
Wang, Y.-X., Balle, B., and Kasiviswanathan, S. (2018) · 2018
Later among the works it cites.