Fetching the paper…
Reading the bibliography…
Constrained reinforcement learning (RL) seeks high-performance policies under safety constraints.
Learning to predict by the methods of temporal differences
Richard S Sutton · 1988
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Sham Kakade and John Langford · 2002
Earlier work this paper cites.
Denoising diffusion implicit models
Jiaming Song, Chenlin Meng, and Stefano Ermon · 2010
Earlier work this paper cites.
Information theory: coding theorems for discrete memoryless systems
Imre Csiszár and János Körner · 2011
Earlier work this paper cites.
Score-based generative modeling through stochastic differential equations
Yang Song, Jascha Sohl-Dickstein, Diederik P Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole · 2011
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P Kingma · 2013
Earlier work this paper cites.
Issues in using function approximation for reinforcement learning
Sebastian Thrun and Anton Schwartz · 2014
Earlier work this paper cites.
Trust region policy optimization
John Schulman · 2015
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli · 2015
Earlier work this paper cites.
Learning structured output representation using deep conditional generative models
Kihyuk Sohn, Honglak Lee, and Xinchen Yan · 2015
Earlier work this paper cites.
Constrained policy optimization
Joshua Achiam, David Held, Aviv Tamar, and Pieter Abbeel · 2017
Earlier work this paper cites.
Risk-constrained reinforcement learning with percentile risk criteria
Yinlam Chow, Mohammad Ghavamzadeh, Lucas Janson, and Marco Pavone · 2018
Earlier work this paper cites.
Information-theoretic considerations in batch reinforcement learning
Jinglin Chen and Nan Jiang · 2019
Earlier work this paper cites.
Off-policy deep reinforcement learning without exploration
Scott Fujimoto, David Meger, and Doina Precup · 2019
Earlier work this paper cites.
Bootstrapping upper confidence bound
Botao Hao, Yasin Abbasi Yadkori, Zheng Wen, and Guang Cheng · 2019
Earlier work this paper cites.
Stabilizing off-policy q-learning via bootstrapping error reduction
Aviral Kumar, Justin Fu, Matthew Soh, George Tucker, and Sergey Levine · 2019
Earlier work this paper cites.
Benchmarking safe exploration in deep reinforcement learning
Alex Ray, Joshua Achiam, and Dario Amodei · 2019
Earlier work this paper cites.
Behavior regularized offline reinforcement learning
Yifan Wu, George Tucker, and Ofir Nachum · 2019
Earlier work this paper cites.
Balancing constraints and rewards with meta-gradient d4pg
Dan A Calian, Daniel J Mankowitz, Tom Zahavy, Zhongwen Xu, Junhyuk Oh, Nir Levine, and Timothy Mann · 2020
Earlier work this paper cites.
Natural policy gradient primal-dual method for constrained markov decision processes
Dongsheng Ding, Kaiqing Zhang, Tamer Basar, and Mihailo Jovanovic · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel · 2020
Earlier work this paper cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Sergey Levine, Aviral Kumar, George Tucker, and Justin Fu · 2020
Earlier work this paper cites.
Responsive safety in reinforcement learning by pid lagrangian methods
Adam Stooke, Joshua Achiam, and Pieter Abbeel · 2020
Cited alongside, same era.
Projection-based constrained policy optimization
Tsung-Yen Yang, Justinian Rosca, Karthik Narasimhan, and Peter J Ramadge · 2020
Cited alongside, same era.
On the theory of policy gradient methods: Optimality, approximation, and distribution shift
Alekh Agarwal, Sham M Kakade, Jason D Lee, and Gaurav Mahajan · 2021
Cited alongside, same era.
Constrained Markov decision processes
Eitan Altman · 2021
Cited alongside, same era.
Decision transformer: Reinforcement learning via sequence modeling
Lili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee, Aditya Grover, Misha Laskin, Pieter Abbeel, Aravind Srinivas, and Igor Mordatch · 2021
Cited alongside, same era.
Score regularized policy optimization through diffusion behavior
Huayu Chen, Cheng Lu, Zhengyi Wang, Hang Su, and Jun Zhu · 2023
Later among the works it cites.
Diffusion policy: Visuomotor policy learning via action diffusion
Cheng Chi, Zhenjia Xu, Siyuan Feng, Eric Cousineau, Yilun Du, Benjamin Burchfiel, Russ Tedrake, and Shuran Song · 2023
Later among the works it cites.
Voce: Variational optimization with conservative estimation for offline safe reinforcement learning
Jiayi Guan, Guang Chen, Jiaming Ji, Long Yang, Zhijun Li, et al · 2023
Later among the works it cites.
Idql: Implicit q-learning as an actor-critic method with diffusion policies
Philippe Hansen-Estruch, Ilya Kostrikov, Michael Janner, Jakub Grudzien Kuba, and Sergey Levine · 2023
Later among the works it cites.
Trust region-based safe distributional reinforcement learning for multiple constraints
Dohyeong Kim, Kyungjae Lee, and Songhwai Oh · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dongsheng Ding, Xiaohan Wei, Zhuoran Yang, Zhaoran Wang, and Mihailo Jovanovic · 2021
Cited alongside, same era.
A minimalist approach to offline reinforcement learning
Scott Fujimoto and Shixiang Shane Gu · 2021
Cited alongside, same era.
Offline reinforcement learning with implicit q-learning
Ilya Kostrikov, Ashvin Nair, and Sergey Levine · 2021
Cited alongside, same era.
Offline reinforcement learning for autonomous driving with safety and exploration enhancement
Tianyu Shi, Dong Chen, Kaian Chen, and Zhaojian Li · 2021
Cited alongside, same era.
Crpo: A new approach for safe reinforcement learning with convergence guarantee
Tengyu Xu, Yingbin Liang, and Guanghui Lan · 2021
Cited alongside, same era.
Is conditional generative modeling all you need for decision-making?
Anurag Ajay, Yilun Du, Abhi Gupta, Joshua Tenenbaum, Tommi Jaakkola, and Pulkit Agrawal · 2022
Cited alongside, same era.
Convergence of denoising diffusion models under the manifold hypothesis
Valentin De Bortoli · 2022
Cited alongside, same era.
Later among the works it cites.
Safe offline reinforcement learning with real-time budget constraints
Qian Lin, Bo Tang, Zifan Wu, Chao Yu, Shangqin Mao, Qianlong Xie, Xingxing Wang, and Dong Wang · 2023
Later among the works it cites.
Contrastive energy prediction for exact energy-guided diffusion sampling in offline reinforcement learning
Cheng Lu, Huayu Chen, Jianfei Chen, Hang Su, Chongxuan Li, and Jun Zhu · 2023
Later among the works it cites.
Safediffuser: Safe planning with diffusion probabilistic models
Wei Xiao, Tsun-Hsuan Wang, Chuang Gan, and Daniela Rus · 2023
Later among the works it cites.
Saformer: A conditional sequence modeling approach to offline safe reinforcement learning
Qin Zhang, Linrui Zhang, Haoran Xu, Li Shen, Bowen Wang, Yongzhe Chang, Xueqian Wang, Bo Yuan, and Dacheng Tao · 2023
Later among the works it cites.
Gradient-adaptive pareto optimization for constrained reinforcement learning
Zixian Zhou, Mengda Huang, Feiyang Pan, Jia He, Xiang Ao, Dandan Tu, and Qing He · 2023
Later among the works it cites.
Diffusion policies creating a trust region for offline reinforcement learning
Tianyu Chen, Zhendong Wang, and Mingyuan Zhou · 2024
Later among the works it cites.
Linjiajie Fang, Ruoxue Liu, Jing Zhang, Wenjia Wang, and Bing-Yi Jing · 2024
Later among the works it cites.
Iterative reachability estimation for safe reinforcement learning
Milan Ganai, Zheng Gong, Chenning Yu, Sylvia Herbert, and Sicun Gao · 2024
Later among the works it cites.
Diffusion model is an effective planner and data synthesizer for multi-task reinforcement learning
Haoran He, Chenjia Bai, Kang Xu, Zhuoran Yang, Weinan Zhang, Dong Wang, Bin Zhao, and Xuelong Li · 2024
Later among the works it cites.
Scale-invariant gradient aggregation for constrained multi-objective reinforcement learning
Dohyeong Kim, Mineui Hong, Jeongho Park, and Songhwai Oh · 2024
Later among the works it cites.
Prajwal Koirala, Zhanhong Jiang, Soumik Sarkar, and Cody Fleming · 2024
Later among the works it cites.
Off-policy primal-dual safe reinforcement learning
Zifan Wu, Bo Tang, Qian Lin, Chao Yu, Shangqin Mao, Qianlong Xie, Xingxing Wang, and Dong Wang · 2024
Later among the works it cites.
Oasis: Conditional distribution shaping for offline safe reinforcement learning
Yihang Yao, Zhepeng Cen, Wenhao Ding, Haohong Lin, Shiqi Liu, Tingnan Zhang, Wenhao Yu, and Ding Zhao · 2024
Later among the works it cites.
Safe offline reinforcement learning with feasibility-guided diffusion model
Yinan Zheng, Jianxiong Li, Dongjie Yu, Yujie Yang, Shengbo Eben Li, Xianyuan Zhan, and Jingjing Liu · 2024
Later among the works it cites.
Constraint-adaptive policy switching for offline safe reinforcement learning
Yassine Chemingui, Aryan Deshwal, Honghao Wei, Alan Fern, and Jana Doppa · 2025
Closest in time.
Offline safe reinforcement learning using trajectory classification
Ze Gong, Akshat Kumar, and Pradeep Varakantham · 2025
Closest in time.
Constraint-conditioned actor-critic for offline safe reinforcement learning
Zijian Guo, Weichao Zhou, Shengao Wang, and Wenchao Li · 2025
Closest in time.