Fetching the paper…
Reading the bibliography…
We present Distributional Soft Actor-Critic (DSAC), a distributional reinforcement learning (RL) algorithm that combines the strengths of distributional information of accumulated rewards and entropy-driven exploration from Soft Actor-Critic (SAC) algorithm.
Distributional reinforcement learning with linear function approximation
Marc G. Bellemare, Nicolas Le Roux, Pablo Samuel Castro, and Subhodeep Moitra. 2019b · 1902
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning. In International conference on machine learning . 1928–1937
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu. 2016 · 1937
Earlier work this paper cites.
Portfolio Selection: Efficient Diversification of Investments
Harry M Markowitz. 1959 · 1959
Earlier work this paper cites.
Robust Estimation of a Location Parameter
Peter J. Huber. 1964 · 1964
Earlier work this paper cites.
The variance of discounted Markov decision processes
Matthew J. Sobel. 1982 · 1982
Earlier work this paper cites.
Advances in prospect theory: Cumulative representation of uncertainty
Amos Tversky and Daniel Kahneman. 1992 · 1992
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Martin L. Puterman. 1994 · 1994
Earlier work this paper cites.
Dynamic Programming and Optimal Control (1st ed.)
Dimitri P. Bertsekas. 1995 · 1995
Earlier work this paper cites.
Neuro-dynamic programming: an overview. In Proceedings of 1995 34th IEEE Conference on Decision and Control , Vol. 1. IEEE, 560–564
Dimitri P. Bertsekas and John N. Tsitsiklis. 1995 · 1995
Earlier work this paper cites.
From Stochastic Dominance to Mean-Risk Models: Semideviations as Risk Measures
Włodzimierz Ogryczak and Andrzej Ruszczyński. 1999 · 1999
Earlier work this paper cites.
Optimization of Conditional Value-at-Risk
R. Tyrrell Rockafellar and Stanislav Uryasev. 2000 · 2000
Earlier work this paper cites.
A class of distortion operators for pricing financial and insurance risks
Shaun S. Wang. 2000 · 2000
Earlier work this paper cites.
General duality between optimal control and estimation. In 2008 47th IEEE Conference on Decision and Control . IEEE, 4286–4292
Emanuel Todorov. 2008 · 2008
Earlier work this paper cites.
Maximum Entropy Inverse Reinforcement Learning. In Proceedings of the Twenty-Third AAAI Conference on Artificial Intelligence, AAAI 2008, Chicago, Illinois, USA, July 13-17, 2008
Brian D. Ziebart, Andrew L. Maas, J. Andrew Bagnell, and Anind K. Dey. 2008 · 2008
Earlier work this paper cites.
Properties of distortion risk measures
Alejandro Balbás, José Garrido, and Silvia Mayoral. 2009 · 2009
Earlier work this paper cites.
Nonparametric return distribution approximation for reinforcement learning. In International Conference on Machine Learning . 799–806
Tetsuro Morimura, Masashi Sugiyama, Hisashi Kashima, Hirotaka Hachiya, and Toshiyuki Tanaka. 2010 · 2010
Earlier work this paper cites.
Policy gradients with variance related risk criteria. In Proceedings of the 29th International Coference on International Conference on Machine Learning . 1651–1658
Aviv Tamar, Dotan Di Castro, and Shie Mannor. 2012 · 2012
Earlier work this paper cites.
MuJoCo: A physics engine for model-based control. In Intelligent Robots and Systems (IROS), 2012 IEEE/RSJ International Conference on
Emanuel Todorov, Tom Erez, and Yuval Tassa. 2012 · 2012
Earlier work this paper cites.
Playing Atari with Deep Reinforcement Learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller. 2013 · 2013
Earlier work this paper cites.
On stochastic optimal control and reinforcement learning by approximate inference. In Twenty-Third International Joint Conference on Artificial Intelligence
Konrad Rawlik, Marc Toussaint, and Sethu Vijayakumar. 2013 · 2013
Earlier work this paper cites.
Risk-sensitive reinforcement learning
Yun Shen, Michael J. Tobia, Tobias Sommer, and Klaus Obermayer. 2014 · 2014
Earlier work this paper cites.
Risk-sensitive and robust decision-making: a CVaR optimization approach. In Advances in Neural Information Processing Systems . 1522–1530
Yinlam Chow, Aviv Tamar, Shie Mannor, and Marco Pavone. 2015 · 2015
Earlier work this paper cites.
Learning continuous control policies by stochastic value gradients. In Advances in Neural Information Processing Systems . 2944–2952
Nicolas Heess, Gregory Wayne, David Silver, Timothy Lillicrap, Tom Erez, and Yuval Tassa. 2015 · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A. Rusu, Joel Veness, Marc G. Bellemare, Alex Graves, Martin Riedmiller, Andreas K. Fidjeland, Georg Ostrovski, et al · 2015
Cited alongside, same era.
Trust region policy optimization. In International conference on machine learning . 1889–1897
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz. 2015 · 2015
Cited alongside, same era.
Geoffrey E. Hinton Jimmy Lei Ba, Jamie Ryan Kiros. 2016 · 2016
Cited alongside, same era.
Variance-constrained actor-critic algorithms for discounted and average reward MDPs
L. A. Prashanth and Mohammad Ghavamzadeh. 2016 · 2016
Distributional Deep Reinforcement Learning with a Mixture of Gaussians. In Proceedings - IEEE International Conference on Robotics and Automation . IEEE, 9791–9797
Yunho Choi, Kyungjae Lee, and Songhwai Oh. 2019 · 2019
Later among the works it cites.
A Theory of Regularized Markov Decision Processes. In International Conference on Machine Learning . 2160–2169
Matthieu Geist, Bruno Scherrer, and Olivier Pietquin. 2019 · 2019
Later among the works it cites.
PyTorch: An imperative style, high-performance deep learning library. In Advances in Neural Information Processing Systems . 8024–8035
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al · 2019
Later among the works it cites.
Nonlinear distributional gradient temporal-difference learning. In International Conference on Machine Learning . 5251–5260
Chao Qu, Shie Mannor, and Huan Xu. 2019 · 2019
Later among the works it cites.
Statistics and Samples in Distributional Reinforcement Learning. In International Conference on Machine Learning . 5528–5536
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Mastering the game of Go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al · 2016
Cited alongside, same era.
Deep reinforcement learning with double q-learning. In Thirtieth AAAI conference on artificial intelligence
Hado Van Hasselt, Arthur Guez, and David Silver. 2016 · 2016
Cited alongside, same era.
Optimization of Markov decision processes under the variance criterion
Li Xia. 2016 · 2016
Cited alongside, same era.
A distributional perspective on reinforcement learning. In International Conference on Machine Learning . JMLR. org, 449–458
Marc G Bellemare, Will Dabney, and Rémi Munos. 2017 · 2017
Cited alongside, same era.
Deep reinforcement learning for robotic manipulation with asynchronous off-policy updates. In 2017 IEEE international conference on robotics and automation (ICRA) . IEEE, 3389–3396
Shixiang Gu, Ethan Holly, Timothy Lillicrap, and Sergey Levine. 2017 · 2017
Cited alongside, same era.
Bridging the gap between value and policy based reinforcement learning. In Proceedings of the 31st International Conference on Neural Information Processing Systems (Long Beach, California, USA) (NIPS’17) . Curran Associates Inc., Red Hook, NY, USA, 2772–2782
Ofir Nachum, Mohammad Norouzi, Kelvin Xu, and Dale Schuurmans. 2017 · 2017
Cited alongside, same era.
Equivalence between policy gradients and soft q-learning
John Schulman, Xi Chen, and Pieter Abbeel. 2017a · 2017
Cited alongside, same era.
Mark Rowland, Robert Dadashi, Saurabh Kumar, Remi Munos, Marc G Bellemare, and Will Dabney. 2019 · 2019
Later among the works it cites.
rlpyt: A Research Code Base for Deep Reinforcement Learning in PyTorch
Adam Stooke and Pieter Abbeel. 2019 · 2019
Later among the works it cites.
Fully Parameterized Quantile Function for Distributional Reinforcement Learning. In Advances in Neural Information Processing Systems . 6190–6199
Derek Yang, Li Zhao, Zichuan Lin, Tao Qin, Jiang Bian, and Tie-Yan Liu. 2019 · 2019
Later among the works it cites.
QUOTA: The quantile option architecture for reinforcement learning. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 33. 5797–5804
Shangtong Zhang and Hengshuai Yao. 2019 · 2019
Later among the works it cites.
A distributional code for value in dopamine-based reinforcement learning
Will Dabney, Zeb Kurth-Nelson, Naoshige Uchida, Clara Kwon Starkweather, Demis Hassabis, Rémi Munos, and Matthew Botvinick. 2020 · 2020
Closest in time.
Improving robustness via risk averse distributional reinforcement learning. In Learning for Dynamics and Control . PMLR, 958–968
Rahul Singh, Qinsheng Zhang, and Yongxin Chen. 2020 · 2020
Closest in time.
Non-crossing quantile regression for deep reinforcement learning. In Proceedings of the 34th International Conference on Neural Information Processing Systems (Vancouver, BC, Canada) (NIPS ’20) . Curran Associates Inc., Red Hook, NY, USA, Article 1334, 11 pages
Fan Zhou, Jianing Wang, and Xingdong Feng. 2020 · 2020
Closest in time.
Distributional soft actor-critic: Off-policy reinforcement learning for addressing value estimation errors
Jingliang Duan, Yang Guan, Shengbo Eben Li, Yangang Ren, Qi Sun, and Bo Cheng. 2021 · 2021
Closest in time.
Conservative offline distributional reinforcement learning. In Proceedings of the 35th International Conference on Neural Information Processing Systems (NIPS ’21) . Curran Associates Inc., Red Hook, NY, USA, Article 1471, 13 pages
Yecheng Jason Ma, Dinesh Jayaraman, and Osbert Bastani. 2021 · 2021
Closest in time.
Non-decreasing Quantile Function Network with Efficient Exploration for Distributional Reinforcement Learning. In Proceedings of the Thirtieth International Joint Conference on Artificial Intelligence, IJCAI-21 , Zhi-Hua Zhou (Ed.). International Joint Conferences on Artificial Intelligence Organization, 3455–3461
Fan Zhou, Zhoufan Zhu, Qi Kuang, and Liwen Zhang. 2021 · 2021
Closest in time.
Fast global convergence of natural policy gradient methods with entropy regularization
Shicong Cen, Chen Cheng, Yuxin Chen, Yuting Wei, and Yuejie Chi. 2022 · 2022
Closest in time.
Distributional Reinforcement Learning with Monotonic Splines. In International Conference on Learning Representations
Yudong Luo, Guiliang Liu, Haonan Duan, Oliver Schulte, and Pascal Poupart. 2022 · 2022
Closest in time.
Mean-Semivariance Policy Optimization via Risk-Averse Reinforcement Learning
Xiaoteng Ma, Shuai Ma, Li Xia, and Qianchuan Zhao. 2022 · 2022
Closest in time.
Risk-Sensitive Reinforcement Learning via Policy Gradient Search
LA Prashanth, Michael C Fu, et al · 2022
Closest in time.
Sample-based distributional policy gradient. In Learning for Dynamics and Control Conference . PMLR, 676–688
Rahul Singh, Keuntaek Lee, and Yongxin Chen. 2022 · 2022
Closest in time.
Interpreting Distributional Reinforcement Learning: Regularization and Optimization Perspectives
Ke Sun, Yingnan Zhao, Yi Liu, Enze Shi, Yafei Wang, Aref Sadeghi, Xiaodong Yan, Bei Jiang, and Linglong Kong. 2022 · 2022
Closest in time.
Stop Regressing: Training Value Functions via Classification for Scalable Deep RL. In Proceedings of the 41st International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 235) , Ruslan Salakhutdinov, Zico Kolter, Katherine Heller, Adrian Weller, Nuria Oliver, Jonathan Scarlett, and Felix Berkenkamp (Eds.). PMLR, 13049–13071
Jesse Farebrother, Jordi Orbay, Quan Vuong, Adrien Ali Taiga, Yevgen Chebotar, Ted Xiao, Alex Irpan, Sergey Levine, Pablo Samuel Castro, Aleksandra Faust, Aviral Kumar, and Rishabh Agarwal. 2024 · 2024
Closest in time.
Inference-Time Scaling for Generalist Reward Modeling
Zijun Liu, Peiyi Wang, Runxin Xu, Shirong Ma, Chong Ruan, Peng Li, Yang Liu, and Yu Wu. 2025 · 2025
Closest in time.