Fetching the paper…
Reading the bibliography…
We study the problem of estimating the distribution of the return of a policy using an offline dataset that is not generated from the policy, i.e., distributional offline policy evaluation (OPE).
Analysis of temporal-diffference learning with function approximation
Tsitsiklis, J. and Van Roy, B · 1996
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
Precup, D., Sutton, R. S., and Singh, S. P · 2000
Earlier work this paper cites.
Empirical Processes in M-estimation , volume 6
Van de Geer, S · 2000
Earlier work this paper cites.
Asymptotic statistics , volume 3
Van der Vaart, A. W · 2000
Earlier work this paper cites.
Error bounds for approximate policy iteration
Munos, R · 2003
Earlier work this paper cites.
Tree-based batch mode reinforcement learning
Ernst, D., Geurts, P., and Wehenkel, L · 2005
Earlier work this paper cites.
From ϵ \epsilon -entropy to kl-entropy: Analysis of minimum information complexity density estimation
Zhang, T · 2006
Earlier work this paper cites.
Finite-time bounds for fitted value iteration
Munos, R. and Szepesvári, C · 2008
Earlier work this paper cites.
Optimal transport: old and new , volume 338
Villani, C. et al · 2009
Earlier work this paper cites.
Scherrer, B · 2010
Earlier work this paper cites.
The fixed points of off-policy td
Kolter, J · 2011
Earlier work this paper cites.
Parametric return density estimation for reinforcement learning
Morimura, T., Sugiyama, M., Kashima, H., Hachiya, H., and Tanaka, T · 2012
Earlier work this paper cites.
Nice: Non-linear independent components estimation
Dinh, L., Krueger, D., and Bengio, Y · 2014
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
Sohl-Dickstein, J., Weiss, E., Maheswaranathan, N., and Ganguli, S · 2015
Earlier work this paper cites.
Doubly robust off-policy value evaluation for reinforcement learning
Jiang, N. and Li, L · 2016
Earlier work this paper cites.
A distributional perspective on reinforcement learning
Bellemare, M. G., Dabney, W., and Munos, R · 2017
Cited alongside, same era.
Markov chains and mixing times , volume 107
Levin, D. A. and Peres, Y · 2017
Cited alongside, same era.
Distributional reinforcement learning with quantile regression
Dabney, W., Rowland, M., Bellemare, M., and Munos, R · 2018
Cited alongside, same era.
Doan, T., Mazoure, B., and Lyle, C · 2018
Cited alongside, same era.
An analysis of categorical distributional reinforcement learning
Rowland, M., Bellemare, M., Dabney, W., Munos, R., and Teh, Y. W · 2018
Cited alongside, same era.
Off-policy risk assessment in contextual bandits
Huang, A., Leqi, L., Lipton, Z., and Azizzadenesheli, K · 2021
Later among the works it cites.
Bayesian distributional policy gradients
Li, L. and Faisal, A. A · 2021
Later among the works it cites.
Conservative offline distributional reinforcement learning
Ma, Y., Jayaraman, D., and Bastani, O · 2021
Later among the works it cites.
Pessimistic model-based offline reinforcement learning under partial coverage
Uehara, M. and Sun, W · 2021
Later among the works it cites.
Representation learning for online and offline rl in low-rank mdps
Uehara, M., Zhang, X., and Sun, W · 2021
Later among the works it cites.
Topics in optimal transportation , volume 58
Villani, C · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Singh, S. and Póczos, B · 2018
Cited alongside, same era.
A kernel loss for solving the bellman equation
Feng, Y., Li, L., and Liu, Q · 2019
Cited alongside, same era.
Distributional multivariate policy evaluation and exploration with the bellman gan
Freirich, D., Shimkin, T., Meir, R., and Tamar, A · 2019
Cited alongside, same era.
Denoising diffusion probabilistic models
Ho, J., Jain, A., and Abbeel, P · 2020
Cited alongside, same era.
Being optimistic to be conservative: Quickly learning a cvar policy
Keramati, R., Dann, C., Tamkin, A., and Brunskill, E · 2020
Cited alongside, same era.
Kinematic state abstraction and provably efficient rich-observation reinforcement learning
Misra, D., Henaff, M., Krishnamurthy, A., and Langford, J · 2020
Cited alongside, same era.
Minimax weight and q-function learning for off-policy evaluation
Uehara, M., Huang, J., and Jiang, N · 2020
Cited alongside, same era.
Bellman-consistent pessimism for offline reinforcement learning
Xie, T., Cheng, C.-A., Jiang, N., Mineiro, P., and Agarwal, A · 2021
Later among the works it cites.
Distributional reinforcement learning for multi-dimensional reward functions
Zhang, P., Chen, X., Zhao, L., Xiong, W., Qin, T., and Liu, T.-Y · 2021
Later among the works it cites.
Learning bellman complete representations for offline policy evaluation
Chang, J., Wang, K., Kallus, N., and Sun, W · 2022
Later among the works it cites.
Off-policy risk assessment for markov decision processes
Huang, A., Leqi, L., Lipton, Z., and Azizzadenesheli, K · 2022
Later among the works it cites.
A wasserstein distance approach for concentration of empirical risk estimates
Prashanth, L. and Bhat, S. P · 2022
Later among the works it cites.
Pac reinforcement learning for predictive state representations
Zhan, W., Uehara, M., Sun, W., and Lee, J. D · 2022
Later among the works it cites.
Distributional Reinforcement Learning
Bellemare, M. G., Dabney, W., and Rowland, M · 2023
Closest in time.
An analysis of quantile temporal-difference learning
Rowland, M., Munos, R., Azar, M. G., Tang, Y., Ostrovski, G., Harutyunyan, A., Tuyls, K., Bellemare, M. G., and Dabney, W · 2023
Closest in time.