Fetching the paper…
Reading the bibliography…
The remarkable empirical performance of distributional reinforcement learning (RL) has garnered increasing attention to understanding its theoretical advantages over classical RL.
Asymptotic evaluation of certain markov process expectations for large time—iii
Donsker, M. D. and Varadhan, S. S · 1976
Earlier work this paper cites.
Function optimization using connectionist reinforcement learning algorithms
Williams, R. J. and Peng, J · 1991
Earlier work this paper cites.
Robust estimation of a location parameter
Huber, P. J · 1992
Earlier work this paper cites.
Q-learning
Watkins, C. J. and Dayan, P · 1992
Earlier work this paper cites.
Robust Statistics , volume 523
Huber, P. J · 2004
Earlier work this paper cites.
Neural fitted q iteration–first experiences with a data efficient neural reinforcement learning method
Riedmiller, M · 2005
Earlier work this paper cites.
All of nonparametric statistics
Wasserman, L · 2006
Earlier work this paper cites.
Parametric return density estimation for reinforcement learning
Morimura, T., Sugiyama, M., Kashima, H., Hachiya, H., and Tanaka, T · 2011
Earlier work this paper cites.
Distilling the knowledge in a neural network
Hinton, G., Vinyals, O., and Dean, J · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Earlier work this paper cites.
Categorical reparameterization with gumbel-softmax
Jang, E., Gu, S., and Poole, B · 2016
Earlier work this paper cites.
Towards principled methods for training generative adversarial networks
Arjovsky, M. and Bottou, L · 2017
Earlier work this paper cites.
Wasserstein generative adversarial networks
Arjovsky, M., Chintala, S., and Bottou, L · 2017
Earlier work this paper cites.
Reinforcement learning with deep energy-based policies
Haarnoja, T., Tang, H., Abbeel, P., and Levine, S · 2017
Earlier work this paper cites.
Efficient exploration through bayesian deep q-networks
Azizzadenesheli, K., Brunskill, E., and Anandkumar, A · 2018
Earlier work this paper cites.
Addressing function approximation error in actor-critic methods
Fujimoto, S., Hoof, H., and Meger, D · 2018
Earlier work this paper cites.
Improving regression performance with distributional losses
Imani, E. and White, M · 2018
Earlier work this paper cites.
Foundations of machine learning, 2018
Mohri, M · 2018
Earlier work this paper cites.
An analysis of categorical distributional reinforcement learning
Rowland, M., Bellemare, M., Dabney, W., Munos, R., and Teh, Y. W · 2018
Earlier work this paper cites.
Reinforcement learning: An Introduction
Sutton, R. S. and Barto, A. G · 2018
Earlier work this paper cites.
Exploration by distributional reinforcement learning
Tang, Y. and Agrawal, S · 2018
Cited alongside, same era.
A comparative analysis of expected and distributional reinforcement learning
Lyle, C., Bellemare, M. G., and Castro, P. S · 2019
Cited alongside, same era.
Distributional reinforcement learning for efficient exploration
Mavrin, B., Zhang, S., Yao, H., Kong, L., Wu, K., and Yu, Y · 2019
Cited alongside, same era.
Propagating uncertainty in reinforcement learning via wasserstein barycenters
Metelli, A. M., Likmeta, A., and Restelli, M · 2019
Cited alongside, same era.
When does label smoothing help?
Müller, R., Kornblith, S., and Hinton, G · 2019
Cited alongside, same era.
Fuzzy tiling activations: A simple approach to learning sparse representations online
Pan, Y., Banman, K., and White, M · 2019
Cited alongside, same era.
Distributional reinforcement learning for risk-sensitive policies
Lim, S. H. and Malik, I · 2022
Closest in time.
A review of uncertainty for deep reinforcement learning
Lockwood, O. and Si, M · 2022
Closest in time.
Dynamic causal effects evaluation in a/b testing with a reinforcement learning framework
Shi, C., Wang, X., Luo, S., Zhu, H., Ye, J., and Song, R · 2022
Closest in time.
Composite quantile regression and the oracle model selection theory
Zou, H. and Yuan, M · 2022
Closest in time.
Distributional reinforcement learning
Bellemare, M. G., Dabney, W., and Rowland, M · 2023
Closest in time.
Pitfall of optimism: Distributional reinforcement learning by randomizing risk criterion
Cho, T., Han, S., Lee, H., Lee, K., and Lee, J · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Statistics and samples in distributional reinforcement learning
Rowland, M., Dadashi, R., Kumar, S., Munos, R., Bellemare, M. G., and Dabney, W · 2019
Cited alongside, same era.
Fully parameterized quantile function for distributional reinforcement learning
Yang, D., Zhao, L., Lin, Z., Qin, T., Bian, J., and Liu, T.-Y · 2019
Cited alongside, same era.
A theoretical analysis of deep q-learning
Fan, J., Wang, Z., Xie, Y., and Yang, Z · 2020
Cited alongside, same era.
Dsac: Distributional soft actor critic for risk-sensitive reinforcement learning
Ma, X., Xia, L., Zhou, Z., Yang, J., and Zhao, Q · 2020
Cited alongside, same era.
On the global convergence rates of softmax policy gradient methods
Mei, J., Xiao, C., Szepesvari, C., and Schuurmans, D · 2020
Cited alongside, same era.
Distributional reinforcement learning with maximum mean discrepancy
Nguyen, T. T., Gupta, S., and Venkatesh, S · 2020
Cited alongside, same era.
Closest in time.
Exploration in deep reinforcement learning: From single-agent to multiagent domain
Hao, J., Yang, T., Tang, H., Bai, C., Liu, J., Meng, Z., Liu, P., and Wang, Z · 2023
Closest in time.
Variance control for distributional reinforcement learning
Kuang, Q., Zhu, Z., Zhang, L., and Zhou, F · 2023
Closest in time.
The statistical benefits of quantile temporal-difference learning for value estimation
Rowland, M., Tang, Y., Lyle, C., Munos, R., Bellemare, M. G., and Dabney, W · 2023
Closest in time.
Adversarial learning of distributional reinforcement learning
Sui, Y., Huang, Y., Zhu, H., and Zhou, F · 2023
Closest in time.
Exploring the training robustness of distributional reinforcement learning against noisy state observations
Sun, K., Liu, Y., Zhao, Y., Yao, H., Jui, S., and Kong, L · 2023
Closest in time.
The benefits of being distributional: Small-loss bounds for reinforcement learning
Wang, K., Zhou, K., Wu, R., Kallus, N., and Sun, W · 2023
Closest in time.
Distributional offline policy evaluation with predictive error guarantees
Wu, R., Uehara, M., and Sun, W · 2023
Closest in time.
Estimation and inference in distributional reinforcement learning
Zhang, L., Peng, Y., Liang, J., Yang, W., and Zhang, Z · 2023
Closest in time.
Provable risk-sensitive distributional reinforcement learning with general function approximation
Chen, Y., Zhang, X., Wang, S., and Huang, L · 2024
Closest in time.
Stop regressing: Training value functions via classification for scalable deep rl
Farebrother, J., Orbay, J., Vuong, Q., Taïga, A. A., Chebotar, Y., Xiao, T., Irpan, A., Levine, S., Castro, P. S., Faust, A., et al · 2024
Closest in time.
How does return distribution in distributional reinforcement learning help optimization?
Sun, K., Jiang, B., and Kong, L · 2024
Closest in time.
More benefits of being distributional: Second-order bounds for reinforcement learning
Wang, K., Oertell, O., Agarwal, A., Kallus, N., and Sun, W · 2024
Closest in time.
Distributional bellman operators over mean embeddings
Wenliang, L. K., Déletang, G., Aitchison, M., Hutter, M., Ruoss, A., Gretton, A., and Rowland, M · 2024
Closest in time.