Fetching the paper…
Reading the bibliography…
In risk-averse reinforcement learning (RL), the goal is to optimize some risk measure of the returns.
The Lord of the Rings: The Fellowship of the Ring
J. R. R. Tolkien · 1954
Earlier work this paper cites.
Survey Sampling
Leslie Kish · 1965
Earlier work this paper cites.
Probability inequalities for sums of bounded random variables
Wassily Hoeffding · 1994
Earlier work this paper cites.
TD algorithm for the variance of return and mean-variance reinforcement learning
Makoto Sato, Hajime Kimura, and Syumpei Kobayashi · 2001
Earlier work this paper cites.
Risk-sensitive optimal control for Markov decision processes with monotone cost
V. S. Borkar and S. P. Meyn · 2002
Earlier work this paper cites.
Rare event estimation for static models via cross-entropy and importance sampling, 2003
Tito Homem de Mello and Reuven Y. Rubinstein · 2003
Earlier work this paper cites.
The cross entropy method for fast policy search
Shie Mannor, Reuven Rubinstein, and Yohai Gat · 2003
Earlier work this paper cites.
A tutorial on the cross-entropy method
P. T. de Boer, Dirk P. Kroese, Shie Mannor, and Reuven Y. Rubinstein · 2005
Earlier work this paper cites.
Cross-entropy method: convergence issues for extended implementation, 2006
Frederic Dambreville · 2006
Earlier work this paper cites.
Simulating sensitivities of conditional value at risk
L. Jeff Hong and Guangwu Liu · 2009
Earlier work this paper cites.
Concentration inequalities: A nonasymptotic theory of independence
Stéphane Boucheron, Gábor Lugosi, and Pascal Massart · 2013
Earlier work this paper cites.
Optimal cloud resource auto-scaling for web applications
Jing Jiang, Jie Lu, Guangquan Zhang, and Guodong Long · 2013
Earlier work this paper cites.
Actor-critic algorithms for risk-sensitive MDPs
L.A. Prashanth and M. Ghavamzadeh · 2013
Earlier work this paper cites.
Risk-constrained Markov decision processes
Vivek Borkar and Rahul Jain · 2014
Earlier work this paper cites.
Algorithms for CVaR optimization in MDPs
Y. Chow and M. Ghavamzadeh · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization, 2014
Jimmy Ba Diederik P. Kingma · 2014
Cited alongside, same era.
Effective sample size, 2014
Tom Leinster · 2014
Cited alongside, same era.
Risk-sensitive and robust decision-making: a CVaR optimization approach
Y. Chow, A. Tamar, S. Mannor, and M. Pavone · 2015
Cited alongside, same era.
A comprehensive survey on safe reinforcement learning
Javier García and Fernando Fernández · 2015
Cited alongside, same era.
Contextual markov decision processes
Assaf Hallak, Dotan Di Castro, and Shie Mannor · 2015
Cited alongside, same era.
Variance-constrained actor-critic algorithms for discounted and average reward MDPs
Worst cases policy gradients
Yichuan Tang, Jian Zhang, and Ruslan Salakhutdinov · 2019
Later among the works it cites.
Quantile QT-Opt for risk-aware vision-based robotic grasping
Cristian Bodnar, Adrian Li, Karol Hausman, Peter Pastor, and Mrinal Kalakrishnan · 2020
Later among the works it cites.
Adaptive sampling for stochastic risk-averse learning
Sebastian Curi, Kfir Y. Levy, Stefanie Jegelka, and Andreas Krause · 2020
Later among the works it cites.
Being optimistic to be conservative: Quickly learning a CVaR policy
Ramtin Keramati, Christoph Dann, Alex Tamkin, and Emma Brunskill · 2020
Later among the works it cites.
Option hedging with risk averse reinforcement learning
Edoardo Vittori, Michele Trapletti, and Marcello Restelli · 2020
Later among the works it cites.
An improved convergence analysis of stochastic variance-reduced policy gradient
Pan Xu, Felicia Gao, and Quanquan Gu · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
L.A. Prashanth and M. Ghavamzadeh · 2016
Cited alongside, same era.
A distributional perspective on reinforcement learning
Marc G Bellemare, Will Dabney, and Rémi Munos · 2017
Cited alongside, same era.
Risk-sensitive inverse reinforcement learning via coherent risk models
Anirudha Majumdar, Sumeet Singh, Ajay Mandlekar, and Marco Pavone · 2017
Cited alongside, same era.
Robust adversarial reinforcement learning
Lerrel Pinto, James Davidson, Rahul Sukthankar, and Abhinav Gupta · 2017
Cited alongside, same era.
Epopt: Learning robust neural network policies using model ensembles
Aravind Rajeswaran, Sarvjeet Ghotra, Balaraman Ravindran, and Sergey Levine · 2017
Cited alongside, same era.
Predictive scaling for EC2, powered by machine learning, 2018
Jeff Barr · 2018
Cited alongside, same era.
Risk-sensitive inverse reinforcement learning via semi- and non-parametric methods
Sumeet Singh, Jonathan Lacotte, Anirudha Majumdar, and Marco Pavone · 2018
Cited alongside, same era.
Later among the works it cites.
Exponential bellman equation and improved regret bounds for risk-sensitive reinforcement learning
Yingjie Fei, Zhuoran Yang, Yudong Chen, and Zhaoran Wang · 2021
Later among the works it cites.
CARL: Conditional-value-at-risk adversarial reinforcement learning
Mathieu Godbout, Maxime Heuillet, Sharath Chandra, Rupali Bhati, and Audrey Durand · 2021
Later among the works it cites.
Challenges of real-world reinforcement learning: Definitions, benchmarks and analysis
Cosmin Paduraru, Daniel J. Mankowitz, Gabriel Dulac-Arnold, Jerry Li, Nir Levine, Sven Gowal, and Todd Hester · 2021
Later among the works it cites.
RMIX: Learning risk-sensitive policies for cooperative reinforcement learning agents
Wei Qiu, Xinrun Wang, Runsheng Yu, Xu He, R. Wang, Bo An, Svetlana Obraztsova, and Zinovi Rabinovich · 2021
Later among the works it cites.
Stable-baselines3: Reliable reinforcement learning implementations
Antonin Raffin, Ashley Hill, Adam Gleave, Anssi Kanervisto, Maximilian Ernestus, and Noah Dormann · 2021
Later among the works it cites.
Nithia Vijayan and L. A. Prashanth · 2021
Later among the works it cites.
Cross entropy method with non-stationary score function
Ido Greenberg · 2022
Closest in time.
Intel confirms acquisition of AI-based workload optimization startup granulate, reportedly for up to $650M, 2022
Ingrid Lunden · 2022
Closest in time.
Reinforcement learning for datacenter congestion control
Chen Tessler, Yuval Shpigelman, Gal Dalal, Amit Mandelbaum, Doron Haritan Kazakov, Benjamin Fuhrer, Gal Chechik, and Shie Mannor · 2022
Closest in time.