Fetching the paper…
Reading the bibliography…
Due to the nature of risk management in learning applicable policies, risk-sensitive reinforcement learning (RSRL) has been realized as an important direction.
Robust estimation of a location parameter
Peter J Huber. 1964 · 1964
Earlier work this paper cites.
Efficient Memory-based Learning for Robot Control
Andrew William Moore. 1990 · 1990
Earlier work this paper cites.
Advances in prospect theory: Cumulative representation of uncertainty
Amos Tversky and Daniel Kahneman. 1992 · 1992
Earlier work this paper cites.
Premium calculation by transforming the layer premium density
Shaun Wang. 1996 · 1996
Earlier work this paper cites.
Curvature of the probability weighting function
George Wu and Richard Gonzalez. 1996 · 1996
Earlier work this paper cites.
Integral probability metrics and their generating classes of functions
Alfred Müller. 1997 · 1997
Earlier work this paper cites.
Quantile regression
Roger Koenker and Kevin F Hallock. 2001 · 2001
Earlier work this paper cites.
Markov decision processes with average-value-at-risk criteria
Nicole Bäuerle and Jonathan Ott. 2011 · 2011
Earlier work this paper cites.
Algorithms for CVaR optimization in MDPs. In Proceedings of the 28th Advances in Neural Information Processing Systems (NeurIPS) . 3509–3517
Yinlam Chow and Mohammad Ghavamzadeh. 2014 · 2014
Earlier work this paper cites.
Risk-sensitive and robust decision-making: a cvar optimization approach. In Proceedings of the 29th Advances in Neural Information Processing Systems (NIPS) , Vol. 28
Yinlam Chow, Aviv Tamar, Shie Mannor, and Marco Pavone. 2015 · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Earlier work this paper cites.
Policy gradient for coherent risk measures. In Proceedings of the 29th Advances in Neural Information Processing Systems (NIPS) , Vol. 28
Aviv Tamar, Yinlam Chow, Mohammad Ghavamzadeh, and Shie Mannor. 2015 · 2015
Earlier work this paper cites.
Constrained policy optimization. In Proceedings of the 34th International Conference on Machine Learning (ICML) . PMLR, 22–31
Joshua Achiam, David Held, Aviv Tamar, and Pieter Abbeel. 2017 · 2017
Cited alongside, same era.
A distributional perspective on reinforcement learning. In Proceedings of the 34th International Conference on Machine Learning (ICML) . PMLR, 449–458
Marc G Bellemare, Will Dabney, and Rémi Munos. 2017 · 2017
Cited alongside, same era.
Mastering the game of go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, et al · 2017
Cited alongside, same era.
Minimalistic Gridworld Environment for OpenAI Gym
Maxime Chevalier-Boisvert, Lucas Willems, and Suman Pal. 2018 · 2018
Cited alongside, same era.
A lyapunov-based approach to safe reinforcement learning
Yinlam Chow, Ofir Nachum, Edgar Duenez-Guzman, and Mohammad Ghavamzadeh. 2018 · 2018
Cited alongside, same era.
Risk-Sensitive Reinforcement Learning: Near-Optimal Risk-Sample Tradeoff in Regret. In Proceedings of the 34th Advances in Neural Information Processing Systems (NeurIPS)
Yingjie Fei, Zhuoran Yang, Yudong Chen, Zhaoran Wang, and Qiaomin Xie. 2020 · 2020
Later among the works it cites.
Being optimistic to be conservative: Quickly learning a cvar policy. In Proceedings of the 34th AAAI Conference on Artificial Intelligence (AAAI) . 4436–4443
Ramtin Keramati, Christoph Dann, Alex Tamkin, and Emma Brunskill. 2020 · 2020
Later among the works it cites.
DSAC: Distributional Soft Actor Critic for Risk-Sensitive Reinforcement Learning
Xiaoteng Ma, Li Xia, Zhengyuan Zhou, Jun Yang, and Qianchuan Zhao. 2020 · 2020
Later among the works it cites.
Conservative Offline Distributional Reinforcement Learning. In Proceedings of the 35th Advances in Neural Information Processing Systems (NeurIPS) . 19235–19247
Yecheng Jason Ma, Dinesh Jayaraman, and Osbert Bastani. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Gal Dalal, Krishnamurthy Dvijotham, Matej Vecerik, Todd Hester, Cosmin Paduraru, and Yuval Tassa. 2018 · 2018
Cited alongside, same era.
Addressing function approximation error in actor-critic methods. In Proceedings of the 35th International Conference on Machine Learning (ICML) . PMLR, 1587–1596
Scott Fujimoto, Herke Hoof, and David Meger. 2018 · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor. In Proceedings of the 35th International Conference on Machine Learning (ICML) . PMLR, 1861–1870
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine. 2018 · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto. 2018 · 2018
Cited alongside, same era.
Statistics and Samples in Distributional Reinforcement Learning. In Proceedings of the 36th International Conference on Machine Learning (ICML) , Vol. 97. 5528–5536
Mark Rowland, Robert Dadashi, Saurabh Kumar, Rémi Munos, Marc G. Bellemare, and Will Dabney. 2019 · 2019
Cited alongside, same era.
Worst Cases Policy Gradients. In Proceedings of the 3rd Annual Conference on Robot Learning (CoRL) , Vol. 100. 1078–1093
Yichuan Charlie Tang, Jian Zhang, and Ruslan Salakhutdinov. 2019 · 2019
Cited alongside, same era.
Implicit quantile networks for distributional reinforcement learning. In Proceedings of the 35th International Conference on Machine Learning (ICML) . PMLR, 1096–1105
Will Dabney, Georg Ostrovski, David Silver, and Rémi Munos. 2018a
Cited in the paper.
Núria Armengol Urpí, Sebastian Curi, and Andreas Krause. 2021 · 2021
Later among the works it cites.
Markov decision processes with recursive risk measures
Nicole Bäuerle and Alexander Glauner. 2022 · 2022
Later among the works it cites.
Distributional Reinforcement Learning for Risk-Sensitive Policies. In Proceedings of the 36th Advances in Neural Information Processing Systems (NeurIPS)
Shiau Hong Lim and Ilyas Malik. 2022 · 2022
Later among the works it cites.
Marc Rigter, Bruno Lacerda, and Nick Hawes. 2022 · 2022
Later among the works it cites.
PerfectDou: Dominating DouDizhu with Perfect Information Distillation
Guan Yang, Minghuan Liu, Weijun Hong, Weinan Zhang, Fei Fang, Guangjun Zeng, and Yue Lin. 2022 · 2022
Later among the works it cites.
Distributional reinforcement learning
Marc G Bellemare, Will Dabney, and Mark Rowland. 2023 · 2023
Closest in time.
Off-policy deep reinforcement learning without exploration. In Proceedings of the 36th International Conference on Machine Learning (ICML) . PMLR, 2052–2062
Scott Fujimoto, David Meger, and Doina Precup. 2019 · 2062
Closest in time.