Fetching the paper…
Reading the bibliography…
The offline reinforcement learning (RL) problem is often motivated by the need to learn data-driven decision policies in financial, legal and healthcare applications.
A measure of asymptotic efficiency for tests of a hypothesis based on the sum of observations
Herman Chernoff et al · 1952
Earlier work this paper cites.
Polynomial-time algorithms for linear programming
George Nemhauser and Laurence Wolsey · 1988
Earlier work this paper cites.
A gentle introduction to concentration inequalities
Karthik Sridharan · 2002
Earlier work this paper cites.
Calibrating noise to sensitivity in private data analysis
Cynthia Dwork, Frank McSherry, Kobbi Nissim, and Adam Smith · 2006
Earlier work this paper cites.
Empirical bernstein bounds and sample variance penalization
Andreas Maurer and Massimiliano Pontil · 2009
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Yasin Abbasi-Yadkori, Dávid Pál, and Csaba Szepesvári · 2011
Earlier work this paper cites.
(nearly) optimal algorithms for private online learning in full-information and bandit settings
Abhradeep Guha Thakurta and Adam Smith · 2013
Earlier work this paper cites.
The algorithmic foundations of differential privacy
Cynthia Dwork, Aaron Roth, et al · 2014
Earlier work this paper cites.
Rényi divergence and kullback-leibler divergence
Tim Van Erven and Peter Harremos · 2014
Earlier work this paper cites.
The simplex method of linear programming
Frederick Arthur Ficken · 2015
Earlier work this paper cites.
Differentially private policy evaluation
Borja Balle, Maziar Gomrokchi, and Doina Precup · 2016
Earlier work this paper cites.
Concentrated differential privacy: Simplifications, extensions, and lower bounds
Mark Bun and Thomas Steinke · 2016
Earlier work this paper cites.
Concentrated differential privacy
Cynthia Dwork and Guy N Rothblum · 2016
Earlier work this paper cites.
The price of differential privacy for online learning
Naman Agarwal and Karan Singh · 2017
Earlier work this paper cites.
Minimax regret bounds for reinforcement learning
Mohammad Gheshlaghi Azar, Ian Osband, and Rémi Munos · 2017
Earlier work this paper cites.
Unifying pac and regret: Uniform pac bounds for episodic reinforcement learning
Christoph Dann, Tor Lattimore, and Emma Brunskill · 2017
Earlier work this paper cites.
iconcur: informed consent for clinical data and bio-sample use for research
Hyeoneui Kim, Elizabeth Bell, Jihoon Kim, Amy Sitapati, Joe Ramsdell, Claudiu Farcas, Dexter Friedman, Stephanie Feudjio Feupe, and Lucila Ohno-Machado · 2017
Earlier work this paper cites.
Continuous state-space models for optimal sepsis treatment: a deep reinforcement learning approach
Aniruddh Raghu, Matthieu Komorowski, Leo Anthony Celi, Peter Szolovits, and Marzyeh Ghassemi · 2017
Earlier work this paper cites.
Deep reinforcement learning framework for autonomous driving
Ahmad EL Sallab, Mohammed Abdou, Etienne Perot, and Senthil Yogamani · 2017
Earlier work this paper cites.
Achieving privacy in the adversarial multi-armed bandit
Aristide Charles Yedia Tossou and Christos Dimitrakakis · 2017
Earlier work this paper cites.
Corrupt bandits for preserving local privacy
Pratik Gajane, Tanguy Urvoy, and Emilie Kaufmann · 2018
Cited alongside, same era.
Differentially private contextual linear bandits
Roshan Shariff and Or Sheffet · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Cited alongside, same era.
Differential privacy for multi-armed bandits: What is it and what is its cost?
Debabrota Basu, Christos Dimitrakakis, and Aristide Tossou · 2019
Cited alongside, same era.
The secret sharer: Evaluating and testing unintended memorization in neural networks
Nicholas Carlini, Chang Liu, Úlfar Erlingsson, Jernej Kos, and Dawn Song · 2019
Cited alongside, same era.
Differentially private regret minimization in episodic markov decision processes
Sayak Ray Chowdhury and Xingyu Zhou · 2021
Later among the works it cites.
Adaptive control of differentially private linear quadratic systems
Sayak Ray Chowdhury, Xingyu Zhou, and Ness Shroff · 2021
Later among the works it cites.
Local differential privacy for regret minimization in reinforcement learning
Evrard Garcelon, Vianney Perchet, Ciara Pike-Burke, and Matteo Pirotta · 2021
Later among the works it cites.
Optimal algorithms for private online learning in a stochastic environment
Bingshan Hu, Zhiming Huang, and Nishant A Mehta · 2021
Later among the works it cites.
Is pessimism provably efficient for offline rl?
Ying Jin, Zhuoran Yang, and Zhaoran Wang · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jonathan Lebensold, William Hamilton, Borja Balle, and Doina Precup · 2019
Cited alongside, same era.
Off-policy policy gradient with state distribution correction
Yao Liu, Adith Swaminathan, Alekh Agarwal, and Emma Brunskill · 2019
Cited alongside, same era.
Privacy-preserving q-learning with functional noise in continuous spaces
Baoxiang Wang and Nidhi Hegde · 2019
Cited alongside, same era.
Privacy preserving off-policy evaluation
Tengyang Xie, Philip S Thomas, and Gerome Miklau · 2019
Cited alongside, same era.
Model-based reinforcement learning with value-targeted regression
Alex Ayoub, Zeyu Jia, Csaba Szepesvari, Mengdi Wang, and Lin Yang · 2020
Cited alongside, same era.
Provably efficient exploration in policy optimization
Qi Cai, Zhuoran Yang, Chi Jin, and Zhaoran Wang · 2020
Cited alongside, same era.
(locally) differentially private combinatorial semi-bandits
Xiaoyu Chen, Kai Zheng, Zixin Zhou, Yunchang Yang, Wei Chen, and Liwei Wang · 2020
Cited alongside, same era.
Chonghua Liao, Jiafan He, and Quanquan Gu · 2021
Later among the works it cites.
Differentially private exploration in reinforcement learning with linear representation
Paul Luyo, Evrard Garcelon, Alessandro Lazaric, and Matteo Pirotta · 2021
Later among the works it cites.
Variance-aware off-policy evaluation with linear function approximation
Yifei Min, Tianhao Wang, Dongruo Zhou, and Quanquan Gu · 2021
Later among the works it cites.
Privately publishable per-instance privacy
Rachel Redberg and Yu-Xiang Wang · 2021
Later among the works it cites.
What are the statistical limits of offline rl with linear function approximation?
Ruosong Wang, Dean P Foster, and Sham M Kakade · 2021
Later among the works it cites.
Near-optimal provable uniform convergence in offline policy evaluation for reinforcement learning
Ming Yin, Yu Bai, and Yu-Xiang Wang · 2021
Later among the works it cites.
Exponential lower bounds for batch reinforcement learning: Batch rl can be exponentially harder than online rl
Andrea Zanette · 2021
Later among the works it cites.
Provable benefits of actor-critic methods for offline reinforcement learning
Andrea Zanette, Martin J Wainwright, and Emma Brunskill · 2021
Later among the works it cites.
Nearly minimax optimal reinforcement learning for linear mixture markov decision processes
Dongruo Zhou, Quanquan Gu, and Csaba Szepesvari · 2021
Later among the works it cites.
Improved regret for differentially private exploration in linear mdp
Dung Daniel Ngo, Giuseppe Vietri, and Zhiwei Steven Wu · 2022
Closest in time.
Near-optimal differentially private reinforcement learning
Dan Qiao and Yu-Xiang Wang · 2022
Closest in time.
Sample-efficient reinforcement learning with loglog(T) switching cost
Dan Qiao, Ming Yin, Ming Min, and Yu-Xiang Wang · 2022
Closest in time.
Ming Yin, Yaqi Duan, Mengdi Wang, and Yu-Xiang Wang · 2022
Closest in time.
Differentially private reinforcement learning with linear function approximation
Xingyu Zhou · 2022
Closest in time.