Fetching the paper…
Reading the bibliography…
When faced with sequential decision-making problems, it is often useful to be able to predict what would happen if decisions were made using a new policy.
Sulla determinazione empirica delle leggi di probabilita
F. P. Cantelli · 1933
Earlier work this paper cites.
Sulla determinazione empirica delle leggi di probabilita
V. Glivenko · 1933
Earlier work this paper cites.
Asymptotic minimax character of the sample distribution function and of the classical multinomial estimator
A. Dvoretzky, J. Kiefer, and J. Wolfowitz · 1956
Earlier work this paper cites.
Confidence limits for the expected value of an arbitrary bounded random variable with a continuous distribution function
T. W. Anderson · 1969
Earlier work this paper cites.
Markov decision processes with a new optimality criterion: Discrete time
S. C. Jaquette · 1973
Earlier work this paper cites.
The variance of discounted Markov decision processes
M. J. Sobel · 1982
Earlier work this paper cites.
Discounted MDP’s: Distribution functions and exponential utility maximization
K.-J. Chung and M. J. Sobel · 1987
Earlier work this paper cites.
Mean, variance, and probabilistic criteria in finite Markov decision processes: A review
D. White · 1988
Earlier work this paper cites.
The tight constant in the Dvoretzky-Kiefer-Wolfowitz inequality
P. Massart · 1990
Earlier work this paper cites.
Markov decision processes
M. L. Puterman · 1990
Earlier work this paper cites.
Bootstrap and wild bootstrap for high dimensional linear models
E. Mammen · 1993
Earlier work this paper cites.
Large Sample Methods in Statistics An Introduction With Applications
P. K. Sen and J. M. Singer · 1993
Earlier work this paper cites.
An Introduction to the Bootstrap
B. Efron and R. J. Tibshirani · 1994
Earlier work this paper cites.
Learning without state-estimation in partially observable Markovian decision processes
S. P. Singh, T. Jaakkola, and M. I. Jordan · 1994
Earlier work this paper cites.
Controlling the false discovery rate: A practical and powerful approach to multiple testing
Y. Benjamini and Y. Hochberg · 1995
Earlier work this paper cites.
Labor market institutions and the distribution of wages, 1973-1992: A semiparametric approach
J. DiNardo, N. M. Fortin, and T. Lemieux · 1995
Earlier work this paper cites.
Python tutorial , volume 620
G. Van Rossum and F. L. Drake Jr · 1995
Earlier work this paper cites.
Optimally combining sampling techniques for Monte Carlo rendering
E. Veach and L. J. Guibas · 1995
Earlier work this paper cites.
Bayesian q-learning
R. Dearden, N. Friedman, and S. Russell · 1998
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
D. Precup · 2000
Earlier work this paper cites.
Asymptotic statistics , volume 3
A. W. Van der Vaart · 2000
Earlier work this paper cites.
On the coherence of expected shortfall
C. Acerbi and D. Tasche · 2002
Earlier work this paper cites.
Explicit nonparametric confidence intervals for the variance with guaranteed coverage
J. P. Romano and M. Wolf · 2002
Earlier work this paper cites.
An empirical likelihood goodness-of-fit test for time series
S. X. Chen, W. Härdle, and M. Li · 2003
Earlier work this paper cites.
Fourier Analysis of Time Series: An Introduction
P. Bloomfield · 2004
Earlier work this paper cites.
The Glivenko-Cantelli Lemma
D. A. Stephens · 2006
Earlier work this paper cites.
Large deviations bounds for estimating conditional value-at-risk
D. B. Brown · 2007
Earlier work this paper cites.
Dual representations for dynamic programming and reinforcement learning
T. Wang, M. Bowling, and D. Schuurmans · 2007
Earlier work this paper cites.
The wild bootstrap, tamed at last
R. Davidson and E. Flachaire · 2008
Earlier work this paper cites.
A probabilistic upper bound on differential entropy
E. Learned-Miller and J. DeStefano · 2008
Earlier work this paper cites.
Causality
J. Pearl · 2009
Earlier work this paper cites.
Nonparametric return distribution approximation for reinforcement learning
T. Morimura, M. Sugiyama, H. Kashima, H. Hachiya, and T. Tanaka · 2010
Earlier work this paper cites.
The Glivenko-Cantelli Theorem
A. M. Shaikh · 2010
Earlier work this paper cites.
PAC bounds for discounted MDPs
T. Lattimore and M. Hutter · 2012
Cited alongside, same era.
Parametric return density estimation for reinforcement learning
T. Morimura, M. Sugiyama, H. Kashima, H. Hachiya, and T. Tanaka · 2012
Cited alongside, same era.
Minimax PAC bounds on the sample complexity of reinforcement learning with a generative model
M. G. Azar, R. Munos, and H. J. Kappen · 2013
Cited alongside, same era.
Inference on counterfactual distributions
V. Chernozhukov, I. Fernández-Val, and B. Melly · 2013
Cited alongside, same era.
Variable risk control via stochastic optimization
S. R. Kuindersma, R. A. Grupen, and A. G. Barto · 2013
Cited alongside, same era.
Model-free intelligent diabetes management using machine learning
M. Bastani · 2014
Concentration inequalities for conditional value at risk
P. Thomas and E. Learned-Miller · 2019
Later among the works it cites.
Preventing undesirable behavior of intelligent machines
P. S. Thomas, B. C. da Silva, A. G. Barto, S. Giguere, Y. Brun, and E. Brunskill · 2019
Later among the works it cites.
Simglucose v0.2.1 (2018) , 2019
J. Xie · 2019
Later among the works it cites.
T. Xie, Y. Ma, and Y.-X. Wang · 2019
Later among the works it cites.
Off-policy evaluation in infinite-horizon reinforcement learning with latent confounders
A. Bennett, N. Kallus, L. Li, and A. Mousavi · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Kolmogorov–Smirnov test: Overview
V. W. Berger and Y. Zhou · 2014
Cited alongside, same era.
Estimation and inference for distribution functions and quantile functions in treatment effect models
S. G. Donald and Y.-C. Hsu · 2014
Cited alongside, same era.
The UVA/PADOVA type 1 diabetes simulator: New features
C. D. Man, F. Micheletto, D. Lv, M. Breton, B. Kovatchev, and C. Cobelli · 2014
Cited alongside, same era.
Universal value function approximators
T. Schaul, D. Horgan, K. Gregor, and D. Silver · 2015
Cited alongside, same era.
Identification and estimation of distributional impacts of interventions using changes in inequality measures
S. Firpo and C. Pinto · 2016
Cited alongside, same era.
Doubly robust off-policy value evaluation for reinforcement learning
N. Jiang and L. Li · 2016
Cited alongside, same era.
D. S. Brown, S. Niekum, and M. Petrik · 2020
Later among the works it cites.
A distributional code for value in dopamine-based reinforcement learning
W. Dabney, Z. Kurth-Nelson, N. Uchida, C. K. Starkweather, D. Hassabis, R. Munos, and M. Botvinick · 2020
Later among the works it cites.
Coindice: Off-policy confidence interval estimation
B. Dai, O. Nachum, Y. Chow, L. Li, C. Szepesvári, and D. Schuurmans · 2020
Later among the works it cites.
Parameter-based value functions
F. Faccio, L. Kirsch, and J. Schmidhuber · 2020
Later among the works it cites.
Deep transfer learning for reducing health care disparities arising from biomedical data inequality
Y. Gao and Y. Cui · 2020
Later among the works it cites.
J. Harb, T. Schaul, D. Precup, and P.-L. Bacon · 2020
Later among the works it cites.
Minimax confidence interval for off-policy evaluation and policy optimization
N. Jiang and J. Huang · 2020
Later among the works it cites.
Double reinforcement learning for efficient off-policy evaluation in Markov decision processes
N. Kallus and M. Uehara · 2020
Later among the works it cites.
Confounding-robust policy evaluation in infinite-horizon reinforcement learning
N. Kallus and A. Zhou · 2020
Later among the works it cites.
Being optimistic to be conservative: Quickly learning a CVaR policy
R. Keramati, C. Dann, A. Tamkin, and E. Brunskill · 2020
Later among the works it cites.
Towards continual reinforcement learning: A review and perspectives
K. Khetarpal, M. Riemer, I. Rish, and D. Precup · 2020
Later among the works it cites.
Statistical bootstrapping for uncertainty estimation in off-policy evaluation
I. Kostrikov and O. Nachum · 2020
Later among the works it cites.
Confident off-policy evaluation and selection through self-normalized importance weighting
I. Kuzborskij, C. Vernade, A. György, and C. Szepesvári · 2020
Later among the works it cites.
Off-policy estimation of long-term average outcomes with applications to mobile health
P. Liao, P. Klasnja, and S. Murphy · 2020
Later among the works it cites.
Importance sampling techniques for policy optimization
A. M. Metelli, M. Papini, N. Montali, and M. Restelli · 2020
Later among the works it cites.
Reinforcement learning via Fenchel-Rockafellar duality
O. Nachum and B. Dai · 2020
Later among the works it cites.
Off-policy policy evaluation for sequential decisions under unobserved confounding
H. Namkoong, R. Keramati, S. Yadlowsky, and E. Brunskill · 2020
Later among the works it cites.
A survey of reinforcement learning algorithms for dynamically varying environments
S. Padakandla · 2020
Later among the works it cites.
Reducing sampling error in batch temporal difference learning
B. Pavse, I. Durugkar, J. P. Hanna, and P. Stone · 2020
Later among the works it cites.
Off-policy evaluation in partially observable environments
G. Tennenholtz, U. Shalit, and S. Mannor · 2020
Later among the works it cites.
Reinforcement learning for strategic recommendations
G. Theocharous, Y. Chandak, P. S. Thomas, and F. de Nijs · 2020
Later among the works it cites.
Minimax weight and q-function learning for off-policy evaluation
M. Uehara, J. Huang, and N. Jiang · 2020
Later among the works it cites.
Deep reinforcement learning amidst lifelong non-stationarity
A. Xie, J. Harrison, and C. Finn · 2020
Later among the works it cites.
Offline policy selection under uncertainty
M. Yang, B. Dai, O. Nachum, G. Tucker, and D. Schuurmans · 2020
Later among the works it cites.
High confidence off-policy (or counterfactual) variance estimation
Y. Chandak, S. Shankar, and P. S. Thomas · 2021
Closest in time.
Non-asymptotic confidence intervals of off-policy evaluation: Primal and dual bounds
Y. Feng, Z. Tang, na zhang, and qiang liu · 2021
Closest in time.
Off-policy risk assessment in contextual bandits
A. Huang, L. Leqi, Z. C. Lipton, and K. Azizzadenesheli · 2021
Closest in time.