Fetching the paper…
Reading the bibliography…
Offline constrained reinforcement learning (RL) aims to learn a policy that maximizes the expected cumulative reward subject to constraints on expected cumulative cost using an existing dataset.
“Constrained Markov Decision Processes: Stochastic Modeling”, 1999
Eitan Altman · 1999
Earlier work this paper cites.
“A Natural Policy Gradient”
Sham Kakade · 2001
Earlier work this paper cites.
“Approximately Optimal Approximate Reinforcement Learning”
Sham Kakade and John Langford · 2002
Earlier work this paper cites.
“Error bounds for approximate policy iteration”
Rémi Munos · 2003
Earlier work this paper cites.
“Convex optimization”, 2004
Stephen. Boyd and Lieven Vandenberghe · 2004
Earlier work this paper cites.
“Error bounds for approximate value iteration”
Rémi Munos · 2005
Earlier work this paper cites.
“Learning near-optimal policies with Bellman-residual minimization based fitted policy iteration and a single sample path”
András Antos, Csaba Szepesvári and Rémi Munos · 2008
Earlier work this paper cites.
“Finite-Time Bounds for Fitted Value Iteration”
Rémi Munos and Csaba Szepesvári · 2008
Earlier work this paper cites.
“Deep reinforcement learning for robotic manipulation with asynchronous off-policy updates”
Shixiang Gu, Ethan Holly, Timothy Lillicrap and Sergey Levine · 2017
Earlier work this paper cites.
“Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor”
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel and Sergey Levine · 2018
Earlier work this paper cites.
“Learning hand-eye coordination for robotic grasping with deep learning and large-scale data collection”
Sergey Levine, Peter Pastor, Alex Krizhevsky, Julian Ibarz and Deirdre Quillen · 2018
Earlier work this paper cites.
“Information-Theoretic Considerations in Batch Reinforcement Learning”
Jinglin Chen and Nan Jiang · 2019
Earlier work this paper cites.
“Batch Policy Learning under Constraints”
Hoang Le, Cameron Voloshin and Yisong Yue · 2019
Earlier work this paper cites.
“Natural Policy Gradient Primal-Dual Method for Constrained Markov Decision Processes”
Dongsheng Ding, Kaiqing Zhang, Tamer Basar and Mihailo Jovanovic · 2020
Cited alongside, same era.
“An empirical investigation of the challenges of real-world reinforcement learning”, 2020
Gabriel Dulac-Arnold, Nir Levine, Daniel. Mankowitz, Jerry Li, Cosmin Paduraru, Sven Gowal and Todd Hester · 2020
Cited alongside, same era.
“Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems”, 2020
Sergey Levine, Aviral Kumar, George Tucker and Justin Fu · 2020
Cited alongside, same era.
“Safe Off-Policy Deep Reinforcement Learning Algorithm for Volt-VAR Control in Power Distribution Systems”
Wei Wang, Nanpeng Yu, Yuanqi Gao and Jie Shi · 2020
Cited alongside, same era.
“Q* Approximation Schemes for Batch Reinforcement Learning: A Theoretical Comparison”
Tengyang Xie and Nan Jiang · 2020
Cited alongside, same era.
“Offline reinforcement learning under value and density-ratio realizability: The power of gaps”
Jinglin Chen and Nan Jiang · 2022
Later among the works it cites.
“Adversarially Trained Actor Critic for Offline Reinforcement Learning”
Ching-An Cheng, Tengyang Xie, Nan Jiang and Alekh Agarwal · 2022
Later among the works it cites.
“Offline Reinforcement Learning: Fundamental Barriers for Value Function Approximation”, 2022
Dylan. Foster, Akshay Krishnamurthy, David Simchi-Levi and Yunzong Xu · 2022
Later among the works it cites.
“Bullet-Safety-Gym: A Framework for Constrained Reinforcement Learning”, 2022
Sven Gronauer · 2022
Later among the works it cites.
“COptiDICE: Offline Constrained Reinforcement Learning via Stationary Distribution Correction Estimation”, 2022
Jongmin Lee, Cosmin Paduraru, Daniel. Mankowitz, Nicolas Heess, Doina Precup, Kee-Eung Kim and Arthur Guez · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Decision transformer: Reinforcement learning via sequence modeling”
Lili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee, Aditya Grover, Misha Laskin, Pieter Abbeel, Aravind Srinivas and Igor Mordatch · 2021
Cited alongside, same era.
“A Primal-Dual Approach to Constrained Markov Decision Processes”, 2021
Yi Chen, Jing Dong and Zhaoran Wang · 2021
Cited alongside, same era.
“A Workflow for Offline Model-Free Robotic Reinforcement Learning”, 2021
Aviral Kumar, Anikait Singh, Stephen Tian, Chelsea Finn and Sergey Levine · 2021
Cited alongside, same era.
“OptiDICE: Offline Policy Optimization via Stationary Distribution Correction Estimation”
Jongmin Lee, Wonseok Jeon, Byungjun Lee, Joelle Pineau and Kee-Eung Kim · 2021
Cited alongside, same era.
“Model Selection for Offline Reinforcement Learning: Practical Considerations for Healthcare Settings”
Shengpu Tang and Jenna Wiens · 2021
Cited alongside, same era.
“Bellman-consistent pessimism for offline reinforcement learning”
Tengyang Xie, Ching-An Cheng, Nan Jiang, Paul Mineiro and Alekh Agarwal · 2021
Cited alongside, same era.
“Batch Value-function Approximation with Only Realizability”
Tengyang Xie and Nan Jiang · 2021
Cited alongside, same era.
“Constraints penalized q-learning for safe offline reinforcement learning”
Haoran Xu, Xianyuan Zhan and Xiangyu Zhu · 2022
Later among the works it cites.
“When is Realizability Sufficient for Off-Policy Reinforcement Learning?”, 2022
Andrea Zanette · 2022
Later among the works it cites.
“Bellman Residual Orthogonalization for Offline Reinforcement Learning”, 2022
Andrea Zanette and Martin. Wainwright · 2022
Later among the works it cites.
“Offline Reinforcement Learning with Realizability and Single-policy Concentrability”
Wenhao Zhan, Baihe Huang, Audrey Huang, Nan Jiang and Jason Lee · 2022
Later among the works it cites.
“Constrained decision transformer for offline safe reinforcement learning”
Zuxin Liu, Zijian Guo, Yihang Yao, Zhepeng Cen, Wenhao Yu, Tingnan Zhang and Ding Zhao · 2023
Closest in time.
“Datasets and Benchmarks for Offline Safe Reinforcement Learning”
Zuxin Liu et al · 2023
Closest in time.
“Revisiting the linear-programming framework for offline rl with general function approximation”
Asuman Ozdaglar, Sarath Pattathil, Jiawei Zhang and Kaiqing Zhang · 2023
Closest in time.
“Importance Weighted Actor-Critic for Optimal Conservative Offline Reinforcement Learning”, 2023
Hanlin Zhu, Paria Rashidinejad and Jiantao Jiao · 2023
Closest in time.