Fetching the paper…
Reading the bibliography…
Distributional reinforcement learning, which focuses on learning the entire return distribution instead of only its expectation in standard RL, has demonstrated remarkable success in enhancing performance.
Q-learning
Christopher JCH Watkins and Peter Dayan · 1992
Earlier work this paper cites.
Neural fitted q iteration–first experiences with a data efficient neural reinforcement learning method
Martin Riedmiller · 2005
Earlier work this paper cites.
All of nonparametric statistics
Larry Wasserman · 2006
Earlier work this paper cites.
Reinforcement learning in finite mdps: Pac analysis
Alexander L Strehl, Lihong Li, and Michael L Littman · 2009
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Earlier work this paper cites.
Train faster, generalize better: Stability of stochastic gradient descent
Moritz Hardt, Ben Recht, and Yoram Singer · 2016
Earlier work this paper cites.
A distributional perspective on reinforcement learning
Marc G Bellemare, Will Dabney, and Rémi Munos · 2017
Earlier work this paper cites.
The cramer distance as a solution to biased wasserstein gradients
Marc G Bellemare, Ivo Danihelka, Will Dabney, Shakir Mohamed, Balaji Lakshminarayanan, Stephan Hoyer, and Rémi Munos · 2017
Earlier work this paper cites.
Improved training of wasserstein gans
Ishaan Gulrajani, Faruk Ahmed, Martin Arjovsky, Vincent Dumoulin, and Aaron C Courville · 2017
Earlier work this paper cites.
Reinforcement learning with deep energy-based policies
Tuomas Haarnoja, Haoran Tang, Pieter Abbeel, and Sergey Levine · 2017
Earlier work this paper cites.
Implicit quantile networks for distributional reinforcement learning
Will Dabney, Georg Ostrovski, David Silver, and Rémi Munos · 2018
Earlier work this paper cites.
Distributional reinforcement learning with quantile regression
Will Dabney, Mark Rowland, Marc G Bellemare, and Rémi Munos · 2018
Earlier work this paper cites.
Soft actor-critic algorithms and applications
Tuomas Haarnoja, Aurick Zhou, Kristian Hartikainen, George Tucker, Sehoon Ha, Jie Tan, Vikash Kumar, Henry Zhu, Abhishek Gupta, Pieter Abbeel, et al · 2018
Cited alongside, same era.
Improving regression performance with distributional losses
Ehsan Imani and Martha White · 2018
Cited alongside, same era.
Spectral normalization for generative adversarial networks
Takeru Miyato, Toshiki Kataoka, Masanori Koyama, and Yuichi Yoshida · 2018
Cited alongside, same era.
How does batch normalization help optimization?
Shibani Santurkar, Dimitris Tsipras, Andrew Ilyas, and Aleksander Madry · 2018
Cited alongside, same era.
Reinforcement learning: An Introduction
Richard S Sutton and Andrew G Barto · 2018
Cited alongside, same era.
Understanding the impact of entropy on policy optimization
On the global convergence rates of softmax policy gradient methods
Jincheng Mei, Chenjun Xiao, Csaba Szepesvari, and Dale Schuurmans · 2020
Later among the works it cites.
Distributional reinforcement learning with maximum mean discrepancy
Thanh Tang Nguyen, Sunil Gupta, and Svetha Venkatesh · 2020
Later among the works it cites.
Towards understanding label smoothing
Yi Xu, Yuanhong Xu, Qi Qian, Hao Li, and Rong Jin · 2020
Later among the works it cites.
Towards deeper deep reinforcement learning
Johan Bjorck, Carla P Gomes, and Kilian Q Weinberger · 2021
Later among the works it cites.
Revisiting rainbow: Promoting more insightful and inclusive deep reinforcement learning research
Johan Samir Obando Ceron and Pablo Samuel Castro · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zafarali Ahmed, Nicolas Le Roux, Mohammad Norouzi, and Dale Schuurmans · 2019
Cited alongside, same era.
A comparative analysis of expected and distributional reinforcement learning
Clare Lyle, Marc G Bellemare, and Pablo Samuel Castro · 2019
Cited alongside, same era.
Distributional reinforcement learning for efficient exploration
Borislav Mavrin, Shangtong Zhang, Hengshuai Yao, Linglong Kong, Kaiwen Wu, and Yaoliang Yu · 2019
Cited alongside, same era.
Fully parameterized quantile function for distributional reinforcement learning
Derek Yang, Li Zhao, Zichuan Lin, Tao Qin, Jiang Bian, and Tie-Yan Liu · 2019
Cited alongside, same era.
Optimality and approximation with policy gradient methods in markov decision processes
Alekh Agarwal, Sham M Kakade, Jason D Lee, and Gaurav Mahajan · 2020
Cited alongside, same era.
A theoretical analysis of deep q-learning
Jianqing Fan, Zhaoran Wang, Yuchen Xie, and Zhuoran Yang · 2020
Cited alongside, same era.
Dsac: Distributional soft actor critic for risk-sensitive reinforcement learning
Xiaoteng Ma, Li Xia, Zhengyuan Zhou, Jun Yang, and Qianchuan Zhao · 2020
Cited alongside, same era.
Florin Gogianu, Tudor Berariu, Mihaela C Rosca, Claudia Clopath, Lucian Busoniu, and Razvan Pascanu · 2021
Later among the works it cites.
Functional regularization for reinforcement learning via learned fourier features
Alexander Li and Deepak Pathak · 2021
Later among the works it cites.
Distributional reinforcement learning with monotonic splines
Yudong Luo, Guiliang Liu, Haonan Duan, Oliver Schulte, and Pascal Poupart · 2021
Later among the works it cites.
Conservative offline distributional reinforcement learning
Yecheng Jason Ma, Dinesh Jayaraman, and Osbert Bastani · 2021
Later among the works it cites.
Interpreting distributional reinforcement learning: A regularization perspective
Ke Sun, Yingnan Zhao, Yi Liu, Shi Enze, Wang Yafei, Yan Xiaodong, Bei Jiang, and Linglong Kong · 2021
Later among the works it cites.
Distributional reinforcement learning via sinkhorn iterations
Ke Sun, Yingnan Zhao, Yi Liu, Bei Jiang, and Linglong Kong · 2022
Closest in time.