Fetching the paper…
Reading the bibliography…
In this paper, we study distributional reinforcement learning from the perspective of statistical efficiency.
Dynamic programming under uncertainty with a quadratic criterion function
H. A. Simon · 1956
Earlier work this paper cites.
A note on certainty equivalence in dynamic planning
H. Theil · 1957
Earlier work this paper cites.
Perturbation theory and finite markov chains
P. J. Schweitzer · 1968
Earlier work this paper cites.
Inequalities in theorems of ergodicity and stability for markov chains with common phase space. i
N. Kartashov · 1986
Earlier work this paper cites.
Asymptotic statistics , volume 3
A. van der Vaart · 2000
Earlier work this paper cites.
Dynamic treatment regimes: practical design considerations
P. W. Lavori and R. Dawson · 2004
Earlier work this paper cites.
The reward hypothesis, 2004
R. S. Sutton · 2004
Earlier work this paper cites.
There is a risk-return trade-off after all
E. Ghysels, P. Santa-Clara, and R. Valkanov · 2005
Earlier work this paper cites.
Sensitivity and convergence of uniformly ergodic markov chains
A. Y. Mitrophanov · 2005
Earlier work this paper cites.
Measure theory , volume 1
V. I. Bogachev · 2007
Earlier work this paper cites.
Nonparametric return distribution approximation for reinforcement learning
T. Morimura, M. Sugiyama, H. Kashima, H. Hachiya, and T. Tanaka · 2010
Earlier work this paper cites.
Fundamentals of stein’s method
N. Ross · 2011
Earlier work this paper cites.
Regular perturbation of v-geometrically ergodic markov chains
D. Ferré, L. Hervé, and J. Ledoux · 2013
Earlier work this paper cites.
Reinforcement learning in robotics: A survey
J. Kober, J. A. Bagnell, and J. Peters · 2013
Earlier work this paper cites.
Probability in Banach Spaces: isoperimetry and processes
M. Ledoux and M. Talagrand · 2013
Earlier work this paper cites.
Towards scaling up markov chain monte carlo: an adaptive subsampling approach
R. Bardenet, A. Doucet, and C. Holmes · 2014
Earlier work this paper cites.
High-confidence off-policy evaluation
P. Thomas, G. Theocharous, and M. Ghavamzadeh · 2015
Earlier work this paper cites.
Noisy monte carlo: Convergence of markov chains with approximate transition kernels
P. Alquier, N. Friel, R. Everitt, and A. Boland · 2016
Earlier work this paper cites.
Mathematical foundations of infinite-dimensional statistical models , volume 40
E. Giné and R. Nickl · 2016
Earlier work this paper cites.
Doubly robust off-policy value evaluation for reinforcement learning
N. Jiang and L. Li · 2016
Earlier work this paper cites.
A distributional perspective on reinforcement learning
M. G. Bellemare, W. Dabney, and R. Munos · 2017
Cited alongside, same era.
Error bounds for approximations of markov chains used in bayesian sampling
J. E. Johndrow and J. C. Mattingly · 2017
Cited alongside, same era.
T. Doan, B. Mazoure, and C. Lyle · 2018
Cited alongside, same era.
An analysis of categorical distributional reinforcement learning
M. Rowland, M. Bellemare, W. Dabney, R. Munos, and Y. W. Teh · 2018
Cited alongside, same era.
Perturbation theory for markov chains via wasserstein distance
D. Rudolf and N. Schweizer · 2018
Cited alongside, same era.
Universal off-policy evaluation
Y. Chandak, S. Niekum, B. da Silva, E. Learned-Miller, E. Brunskill, and P. S. Thomas · 2021
Later among the works it cites.
Bootstrapping fitted q-evaluation for off-policy inference
B. Hao, X. Ji, Y. Duan, H. Lu, C. Szepesvari, and M. Wang · 2021
Later among the works it cites.
Discovering faster matrix multiplication algorithms with reinforcement learning
A. Fawzi, M. Balog, A. Huang, T. Hubert, B. Romera-Paredes, M. Barekatain, A. Novikov, F. J. R Ruiz, J. Schrittwieser, G. Swirszcz, et al · 2022
Later among the works it cites.
Off-policy risk assessment for markov decision processes
A. Huang, L. Leqi, Z. Lipton, and K. Azizzadenesheli · 2022
Later among the works it cites.
Distributional reinforcement learning for risk-sensitive policies
S. H. Lim and I. MALIK · 2022
Later among the works it cites.
A review of uncertainty for deep reinforcement learning
O. Lockwood and M. Si · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
D. Silver, T. Hubert, J. Schrittwieser, I. Antonoglou, M. Lai, A. Guez, M. Lanctot, L. Sifre, D. Kumaran, T. Graepel, et al · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
R. S. Sutton and A. G. Barto · 2018
Cited alongside, same era.
High-Dimensional Probability: An Introduction with Applications in Data Science
R. Vershynin · 2018
Cited alongside, same era.
Information-theoretic considerations in batch reinforcement learning
J. Chen and N. Jiang · 2019
Cited alongside, same era.
Estimating risk and uncertainty in deep reinforcement learning
W. R. Clements, B. Van Delft, B.-M. Robaglia, R. B. Slaoui, and S. Toth · 2019
Cited alongside, same era.
Distributional multivariate policy evaluation and exploration with the bellman gan
D. Freirich, T. Shimkin, R. Meir, and A. Tamar · 2019
Cited alongside, same era.
Gan-powered deep distributional reinforcement learning for resource management in network slicing
Y. Hua, R. Li, Z. Zhao, X. Chen, and H. Zhang · 2019
Cited alongside, same era.
Later among the works it cites.
Training language models to follow instructions with human feedback
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. L. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al · 2022
Later among the works it cites.
Statistical inference of the value function for reinforcement learning in infinite-horizon settings
C. Shi, S. Zhang, W. Lu, and R. Song · 2022
Later among the works it cites.
Interpreting distributional reinforcement learning: A regularization perspective, 2022
K. Sun, Y. Zhao, Y. Liu, E. Shi, Y. Wang, X. Yan, B. Jiang, and L. Kong · 2022
Later among the works it cites.
Toward theoretical understandings of robust markov decision processes: Sample complexity and asymptotics
W. Yang, L. Zhang, and Z. Zhang · 2022
Later among the works it cites.
Distributional Reinforcement Learning
M. G. Bellemare, W. Dabney, and M. Rowland · 2023
Closest in time.
Regret bounds for risk-sensitive reinforcement learning with lipschitz dynamic risk measures
H. Liang and Z.-q. Luo · 2023
Closest in time.
Gpt-4 technical report, 2023
OpenAI · 2023
Closest in time.
An analysis of quantile temporal-difference learning
M. Rowland, R. Munos, M. G. Azar, Y. Tang, G. Ostrovski, A. Harutyunyan, K. Tuyls, M. G. Bellemare, and W. Dabney · 2023
Closest in time.
The benefits of being distributional: Small-loss bounds for reinforcement learning
K. Wang, K. Zhou, R. Wu, N. Kallus, and W. Sun · 2023
Closest in time.
Distributional offline policy evaluation with predictive error guarantees
R. Wu, M. Uehara, and W. Sun · 2023
Closest in time.
Uncertainty quantification and exploration for reinforcement learning
Y. Zhu, J. Dong, and H. Lam · 2023
Closest in time.
Perturbations of markov chains
D. Rudolf, A. Smith, and M. Quiroz · 2024
Closest in time.
A distributional analogue to the successor representation
H. Wiltzer, J. Farebrother, A. Gretton, Y. Tang, A. Barreto, W. Dabney, M. G. Bellemare, and M. Rowland · 2024
Closest in time.