Fetching the paper…
Reading the bibliography…
We consider the offline constrained reinforcement learning (RL) problem, in which the agent aims to compute a policy that maximizes expected return while satisfying given cost constraints, learning only from a pre-collected dataset.
AlgaeDICE: Policy gradient from arbitrary experience
Ofir Nachum, Bo Dai, Ilya Kostrikov, Yinlam Chow, Lihong Li, and Dale Schuurmans · 1912
Earlier work this paper cites.
An Introduction to the Bootstrap
Bradley Efron and Robert J. Tibshirani · 1993
Earlier work this paper cites.
Mixture density networks
Christopher M. Bishop · 1994
Earlier work this paper cites.
Constrained Markov Decision Processes
Eitan Altman · 1999
Earlier work this paper cites.
An actor-critic algorithm for constrained markov decision processes
V.S. Borkar · 2005
Earlier work this paper cites.
Robust dynamic programming
Garud N. Iyengar · 2005
Earlier work this paper cites.
Acme: A research framework for distributed reinforcement learning
Matt Hoffman, Bobak Shahriari, John Aslanides, Gabriel Barth-Maron, Feryal Behbahani, Tamara Norman, Abbas Abdolmaleki, Albin Cassirer, Fan Yang, Kate Baumli, Sarah Henderson, Alex Novikov, Sergio Gómez Colmenarejo, Serkan Cabi, Caglar Gulcehre, Tom Le Paine, Andrew Cowie, Ziyu Wang, Bilal Piot, and Nando de Freitas · 2006
Earlier work this paper cites.
Reinforcement learning: State-of-the-art
Sascha Lange, Thomas Gabel, and Martin Riedmiller · 2012
Earlier work this paper cites.
Scaling up robust mdps using function approximation
Aviv Tamar, Shie Mannor, and Huan Xu · 2014
Earlier work this paper cites.
Risk-sensitive and robust decision-making: a cvar optimization approach
Yinlam Chow, Aviv Tamar, Shie Mannor, and Marco Pavone · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A. Rusu, Joel Veness, Marc G. Bellemare, Alex Graves, Martin Riedmiller, Andreas K. Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis · 2015
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Timothy P. Lillicrap, Jonathan J. Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2016
Earlier work this paper cites.
Constrained policy optimization
Joshua Achiam, David Held, Aviv Tamar, and Pieter Abbeel · 2017
Earlier work this paper cites.
Improved training of Wasserstein GANs
Ishaan Gulrajani, Faruk Ahmed, Martin Arjovsky, Vincent Dumoulin, and Aaron C Courville · 2017
Earlier work this paper cites.
Bootstrapping with models: Confidence intervals for off-policy evaluation
Josiah P. Hanna, Peter Stone, and Scott Niekum · 2017
Earlier work this paper cites.
Mastering the game of Go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, Yutian Chen, Timothy Lillicrap, Fan Hui, Laurent Sifre, George van den Driessche, Thore Graepel, and Demis Hassabis · 2017
Cited alongside, same era.
Maximum a posteriori policy optimisation
Abbas Abdolmaleki, Jost Tobias Springenberg, Yuval Tassa, Remi Munos, Nicolas Heess, and Martin Riedmiller · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Cited alongside, same era.
Off-policy deep reinforcement learning without exploration
Scott Fujimoto, David Meger, and Doina Precup · 2019
Cited alongside, same era.
Way off-policy batch deep reinforcement learning of implicit human preferences in dialog, 2019
Natasha Jaques, Asma Ghandeharioun, Judy Hanwen Shen, Craig Ferguson, Agata Lapedriza, Noah Jones, Shixiang Gu, and Rosalind Picard · 2019
Cited alongside, same era.
MOReL : Model-based offline reinforcement learning
Rahul Kidambi, Aravind Rajeswaran, Praneeth Netrapalli, and Thorsten Joachims · 2020
Later among the works it cites.
Conservative Q-learning for offline reinforcement learning
Aviral Kumar, Aurick Zhou, George Tucker, and Sergey Levine · 2020
Later among the works it cites.
Batch reinforcement learning with hyperparameter gradients
Byungjun Lee, Jongmin Lee, Peter Vrancx, Dongho Kim, and Kee-Eung Kim · 2020
Later among the works it cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems, 2020
Sergey Levine, Aviral Kumar, George Tucker, and Justin Fu · 2020
Later among the works it cites.
Constrained Markov decision processes via backward value functions
Harsh Satija, Philip Amortila, and Joelle Pineau · 2020
Later among the works it cites.
Keep doing what worked: Behavioral modelling priors for offline reinforcement learning, 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Imitation learning via off-policy distribution matching
Ilya Kostrikov, Ofir Nachum, and Jonathan Tompson · 2019
Cited alongside, same era.
Stabilizing off-policy Q-learning via bootstrapping error reduction
Aviral Kumar, Justin Fu, Matthew Soh, George Tucker, and Sergey Levine · 2019
Cited alongside, same era.
Safe policy improvement with baseline bootstrapping
Romain Laroche, Paul Trichelair, and Remi Tachet Des Combes · 2019
Cited alongside, same era.
Batch policy learning under constraints
Hoang Le, Cameron Voloshin, and Yisong Yue · 2019
Cited alongside, same era.
Benchmarking Safe Exploration in Deep Reinforcement Learning
Alex Ray, Joshua Achiam, and Dario Amodei · 2019
Cited alongside, same era.
Reward constrained policy optimization
Chen Tessler, Daniel J. Mankowitz, and Shie Mannor · 2019
Cited alongside, same era.
Behavior regularized offline reinforcement learning, 2019
Yifan Wu, George Tucker, and Ofir Nachum · 2019
Cited alongside, same era.
Noah Y. Siegel, Jost Tobias Springenberg, Felix Berkenkamp, Abbas Abdolmaleki, Michael Neunert, Thomas Lampe, Roland Hafner, Nicolas Heess, and Martin Riedmiller · 2020
Later among the works it cites.
Critic regularized regression
Ziyu Wang, Alexander Novikov, Konrad Zolna, Josh S Merel, Jost Tobias Springenberg, Scott E Reed, Bobak Shahriari, Noah Siegel, Caglar Gulcehre, Nicolas Heess, and Nando de Freitas · 2020
Later among the works it cites.
MOPO: Model-based offline policy optimization
Tianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon, James Zou, Sergey Levine, Chelsea Finn, and Tengyu Ma · 2020
Later among the works it cites.
Offline reinforcement learning with fisher divergence critic regularization
Ilya Kostrikov, Rob Fergus, Jonathan Tompson, and Ofir Nachum · 2021
Later among the works it cites.
OptiDICE: Offline policy optimization via stationary distribution correction estimation
Jongmin Lee, Wonseok Jeon, Byungjun Lee, Joelle Pineau, and Kee-Eung Kim · 2021
Later among the works it cites.
Robust constrained reinforcement learning for continuous control with model misspecification, 2021
Daniel J. Mankowitz, Dan A. Calian, Rae Jeong, Cosmin Paduraru, Nicolas Heess, Sumanth Dathathri, Martin Riedmiller, and Timothy Mann · 2021
Later among the works it cites.
Risk-averse offline reinforcement learning
Núria Armengol Urpí, Sebastian Curi, and Andreas Krause · 2021
Later among the works it cites.
Constraints penalized q-learning for safe offline reinforcement learning, 2021
Haoran Xu, Xianyuan Zhan, and Xiangyu Zhu · 2021
Later among the works it cites.
Wcsac: Worst-case soft actor critic for safety-constrained reinforcement learning
Qisong Yang, Thiago D. Simão, Simon H Tindemans, and Matthijs T. J. Spaan · 2021
Later among the works it cites.