Fetching the paper…
Reading the bibliography…
Offline reinforcement learning (RL) is suitable for safety-critical domains where online exploration is too costly or dangerous.
Dyna, an integrated architecture for learning, planning, and reacting
Richard S Sutton · 1991
Earlier work this paper cites.
Coherent measures of risk
Philippe Artzner, Freddy Delbaen, Jean-Marc Eber, and David Heath · 1999
Earlier work this paper cites.
A class of distortion operators for pricing financial and insurance risks
Shaun S Wang · 2000
Earlier work this paper cites.
Optimal execution of portfolio transactions
Robert Almgren and Neil Chriss · 2001
Earlier work this paper cites.
Conditional value-at-risk for general loss distributions
R Tyrrell Rockafellar and Stanislav Uryasev · 2002
Earlier work this paper cites.
Time consistent dynamic risk measures
Kang Boda and Jerzy A Filar · 2006
Earlier work this paper cites.
Clinical data based optimal STI strategies for HIV: a reinforcement learning approach
Damien Ernst, Guy-Bart Stan, Jorge Goncalves, and Louis Wehenkel · 2006
Earlier work this paper cites.
Ornstein–uhlenbeck processes and extensions
Ross A Maller, Gernot Müller, and Alex Szimayer · 2009
Earlier work this paper cites.
Parametric return density estimation for reinforcement learning
Tetsuro Morimura, Masashi Sugiyama, Hisashi Kashima, Hirotaka Hachiya, and Toshiyuki Tanaka · 2010
Earlier work this paper cites.
Risk-averse dynamic programming for Markov decision processes
Andrzej Ruszczyński · 2010
Earlier work this paper cites.
Policy gradients with variance related risk criteria
Aviv Tamar, Dotan Di Castro, and Shie Mannor · 2012
Earlier work this paper cites.
Markov decision processes: Discrete stochastic dynamic programming
Martin L Puterman · 2014
Earlier work this paper cites.
A comprehensive survey on safe reinforcement learning
Javier Garcıa and Fernando Fernández · 2015
Earlier work this paper cites.
RLPy: a value-function-based reinforcement learning framework for education and research
Alborz Geramifard, Christoph Dann, Robert H Klein, William Dabney, and Jonathan P How · 2015
Earlier work this paper cites.
Bayesian reinforcement learning: A survey
Mohammad Ghavamzadeh, Shie Mannor, Joelle Pineau, Aviv Tamar, et al · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Policy gradient for coherent risk measures
Aviv Tamar, Yinlam Chow, Mohammad Ghavamzadeh, and Shie Mannor · 2015
Earlier work this paper cites.
Optimizing the CVaR via sampling
Aviv Tamar, Yonatan Glassner, and Shie Mannor · 2015
Earlier work this paper cites.
Conditional value-at-risk for elliptical distributions
Valentyn Khokhlov · 2016
Earlier work this paper cites.
Verification of continuous time random walk wind model
Wojciech Popko, Matthias Wächter, and Philipp Thomas · 2016
Earlier work this paper cites.
A distributional perspective on reinforcement learning
Marc G Bellemare, Will Dabney, and Rémi Munos · 2017
Earlier work this paper cites.
Risk-constrained reinforcement learning with percentile risk criteria
Yinlam Chow, Mohammad Ghavamzadeh, Lucas Janson, and Marco Pavone · 2017
Earlier work this paper cites.
EPOpt: Learning robust neural network policies using model ensembles
Aravind Rajeswaran, Sarvjeet Ghotra, Balaraman Ravindran, and Sergey Levine · 2017
Earlier work this paper cites.
Deep reinforcement learning in a handful of trials using probabilistic dynamics models
Kurtland Chua, Roberto Calandra, Rowan McAllister, and Sergey Levine · 2018
Earlier work this paper cites.
Implicit quantile networks for distributional reinforcement learning
Will Dabney, Georg Ostrovski, David Silver, and Rémi Munos · 2018
Earlier work this paper cites.
Decomposition of uncertainty in Bayesian deep learning for efficient and risk-sensitive learning
Stefan Depeweg, Jose-Miguel Hernandez-Lobato, Finale Doshi-Velez, and Steffen Udluft · 2018
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Cited alongside, same era.
Model-ensemble trust-region policy optimization
Thanard Kurutach, Ignasi Clavera, Yan Duan, Aviv Tamar, and Pieter Abbeel · 2018
Cited alongside, same era.
Multi-agent deep reinforcement learning for liquidation strategy analysis
Wenhang Bao and Xiao-yang Liu · 2019
Cited alongside, same era.
Off-policy deep reinforcement learning without exploration
Scott Fujimoto, David Meger, and Doina Precup · 2019
Cited alongside, same era.
When to trust your model: Model-based policy optimization
Michael Janner, Justin Fu, Marvin Zhang, and Sergey Levine · 2019
Cited alongside, same era.
Stabilizing off-policy Q-learning via bootstrapping error reduction
Offline meta reinforcement learning – identifiability challenges and effective data collection strategies
Ron Dorfman, Idan Shenfeld, and Aviv Tamar · 2021
Later among the works it cites.
A minimalist approach to offline reinforcement learning
Scott Fujimoto and Shixiang Shane Gu · 2021
Later among the works it cites.
Two steps to risk sensitivity
Christopher Gagne and Peter Dayan · 2021
Later among the works it cites.
Mastering atari with discrete world models
Danijar Hafner, Timothy Lillicrap, Mohammad Norouzi, and Jimmy Ba · 2021
Later among the works it cites.
Offline reinforcement learning with Fisher divergence critic regularization
Ilya Kostrikov, Rob Fergus, Jonathan Tompson, and Ofir Nachum · 2021
Later among the works it cites.
Conservative offline distributional reinforcement learning
Yecheng Ma, Dinesh Jayaraman, and Osbert Bastani · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Aviral Kumar, Justin Fu, Matthew Soh, George Tucker, and Sergey Levine · 2019
Cited alongside, same era.
Off-policy policy gradient with state distribution correction
Yao Liu, Adith Swaminathan, Alekh Agarwal, and Emma Brunskill · 2019
Cited alongside, same era.
AlgaeDICE: Policy gradient from arbitrary experience
Ofir Nachum, Bo Dai, Ilya Kostrikov, Yinlam Chow, Lihong Li, and Dale Schuurmans · 2019
Cited alongside, same era.
Behavior regularized offline reinforcement learning
Yifan Wu, George Tucker, and Ofir Nachum · 2019
Cited alongside, same era.
An optimistic perspective on offline reinforcement learning
Rishabh Agarwal, Dale Schuurmans, and Mohammad Norouzi · 2020
Cited alongside, same era.
Estimating risk and uncertainty in deep reinforcement learning
William R Clements, Bastien Van Delft, Benoît-Marie Robaglia, Reda Bahi Slaoui, and Sébastien Toth · 2020
Cited alongside, same era.
Epistemic risk-sensitive reinforcement learning
Hannes Eriksson and Christos Dimitrakakis · 2020
Cited alongside, same era.
Later among the works it cites.
Risk-averse Bayes-adaptive reinforcement learning
Marc Rigter, Bruno Lacerda, and Nick Hawes · 2021
Later among the works it cites.
Lectures on stochastic programming: modeling and theory
Alexander Shapiro, Darinka Dentcheva, and Andrzej Ruszczynski · 2021
Later among the works it cites.
Lectures on stochastic programming: modeling and theory
Alexander Shapiro, Darinka Dentcheva, and Andrzej Ruszczynski · 2021
Later among the works it cites.
Overcoming model bias for robust offline deep reinforcement learning
Phillip Swazinna, Steffen Udluft, and Thomas Runkler · 2021
Later among the works it cites.
Risk-averse offline reinforcement learning
Núria Armengol Urpí, Sebastian Curi, and Andreas Krause · 2021
Later among the works it cites.
Offline reinforcement learning with reverse model-based imagination
Jianhao Wang, Wenzhe Li, Haozhe Jiang, Guangxiang Zhu, Siyuan Li, and Chongjie Zhang · 2021
Later among the works it cites.
Bellman-consistent pessimism for offline reinforcement learning
Tengyang Xie, Ching-An Cheng, Nan Jiang, Paul Mineiro, and Alekh Agarwal · 2021
Later among the works it cites.
Pareto policy pool for model-based offline reinforcement learning
Yijun Yang, Jing Jiang, Tianyi Zhou, Jie Ma, and Yuhui Shi · 2021
Later among the works it cites.
COMBO: Conservative offline model-based policy optimization
Tianhe Yu, Aviral Kumar, Rafael Rafailov, Aravind Rajeswaran, Sergey Levine, and Chelsea Finn · 2021
Later among the works it cites.
Adversarially trained actor critic for offline reinforcement learning
Ching-An Cheng, Tengyang Xie, Nan Jiang, and Alekh Agarwal · 2022
Closest in time.
Sentinel: Taming uncertainty with ensemble based distributional reinforcement learning
Hannes Eriksson, Debabrota Basu, Mina Alibeigi, and Christos Dimitrakakis · 2022
Closest in time.
Offline RL policies should be trained to be adaptive
Dibya Ghosh, Anurag Ajay, Pulkit Agrawal, and Sergey Levine · 2022
Closest in time.
Model-based offline reinforcement learning with pessimism-modulated dynamics belief
Kaiyang Guo, Yunfeng Shao, and Yanhui Geng · 2022
Closest in time.
Offline reinforcement learning with implicit Q-learning
Ilya Kostrikov, Ashvin Nair, and Sergey Levine · 2022
Closest in time.
Revisiting design choices in offline model based reinforcement learning
Cong Lu, Philip Ball, Jack Parker-Holder, Michael Osborne, and Stephen J Roberts · 2022
Closest in time.
Challenges and opportunities in offline reinforcement learning from visual observations
Cong Lu, Philip J Ball, Tim GJ Rudner, Jack Parker-Holder, Michael A Osborne, and Yee Whye Teh · 2022
Closest in time.
Planning for risk-aversion and expected value in MDPs
Marc Rigter, Paul Duckworth, Bruno Lacerda, and Nick Hawes · 2022
Closest in time.
RAMBO-RL: Robust adversarial model-based offline reinforcement learning
Marc Rigter, Bruno Lacerda, and Nick Hawes · 2022
Closest in time.
RORL: Robust offline reinforcement learning via conservative smoothing
Rui Yang, Chenjia Bai, Xiaoteng Ma, Zhaoran Wang, Chongjie Zhang, and Lei Han · 2022
Closest in time.