Fetching the paper…
Reading the bibliography…
In many real-world applications, a reinforcement learning (RL) agent should consider multiple objectives and adhere to safety guidelines.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
Nonlinear multiobjective optimization , volume 12
Kaisa Miettinen · 1999
Earlier work this paper cites.
Multiobjective evolutionary algorithms: a comparative case study and the strength pareto approach
E. Zitzler and L. Thiele · 1999
Earlier work this paper cites.
A fast elitist non-dominated sorting genetic algorithm for multi-objective optimization: NSGA-II
Kalyanmoy Deb, Samir Agrawal, Amrit Pratap, and T. Meyarivan · 2000
Earlier work this paper cites.
Multiple-gradient descent algorithm (MGDA) for multiobjective optimization
Jean-Antoine Désidéri · 2012
Earlier work this paper cites.
Scalarized multi-objective reinforcement learning: Novel design techniques
Kristof Van Moffaert, Madalina M. Drugan, and Ann Nowé · 2013
Earlier work this paper cites.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Earlier work this paper cites.
Safe and efficient off-policy reinforcement learning
Rémi Munos, Tom Stepleton, Anna Harutyunyan, and Marc Bellemare · 2016
Earlier work this paper cites.
Constrained policy optimization
Joshua Achiam, David Held, Aviv Tamar, and Pieter Abbeel · 2017
Earlier work this paper cites.
Distributional reinforcement learning with quantile regression
Will Dabney, Mark Rowland, Marc Bellemare, and Rémi Munos · 2018
Earlier work this paper cites.
Addressing function approximation error in actor-critic methods
Scott Fujimoto, Herke van Hoof, and David Meger · 2018
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Earlier work this paper cites.
A generalized algorithm for multi-objective reinforcement learning and policy adaptation
Runzhe Yang, Xingyuan Sun, and Karthik Narasimhan · 2019
Cited alongside, same era.
A distributional view on multi-objective policy optimization
Abbas Abdolmaleki, Sandy Huang, Leonard Hasenclever, Michael Neunert, Francis Song, Martina Zambelli, Murilo Martins, Nicolas Heess, Raia Hadsell, and Martin Riedmiller · 2020
Cited alongside, same era.
Combining a gradient-based method and an evolution strategy for multi-objective reinforcement learning
Diqi Chen, Yizhou Wang, and Wen Gao · 2020
Cited alongside, same era.
Controlling overestimation bias with truncated mixture of continuous distributional quantile critics
Arsenii Kuznetsov, Pavel Shvechikov, Alexander Grishin, and Dmitry Vetrov · 2020
Cited alongside, same era.
Responsive safety in reinforcement learning by PID Lagrangian methods
Adam Stooke, Joshua Achiam, and Pieter Abbeel · 2020
Cited alongside, same era.
Dongsheng Ding, Kaiqing Zhang, Jiali Duan, Tamer Başar, and Mihailo R. Jovanović · 2022
Later among the works it cites.
A practical guide to multi-objective reinforcement learning and planning
Conor F Hayes, Roxana Rădulescu, Eugenio Bargiacchi, Johan Källström, Matthew Macfarlane, Mathieu Reymond, Timothy Verstraeten, Luisa M Zintgraf, Richard Dazeley, Fredrik Heintz, et al · 2022
Later among the works it cites.
A constrained multi-objective reinforcement learning framework
Sandy Huang, Abbas Abdolmaleki, Giulia Vezzani, Philemon Brakel, Daniel J. Mankowitz, Michael Neunert, Steven Bohez, Yuval Tassa, Nicolas Heess, Martin Riedmiller, and Raia Hadsell · 2022
Later among the works it cites.
Efficient off-policy safe reinforcement learning using trust region conditional value at risk
Dohyeong Kim and Songhwai Oh · 2022
Later among the works it cites.
Pareto policy adaptation
Panagiotis Kyriakis and Jyotirmoy Deshmukh · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Prediction-guided multi-objective reinforcement learning for continuous robot control
Jie Xu, Yunsheng Tian, Pingchuan Ma, Daniela Rus, Shinjiro Sueda, and Wojciech Matusik · 2020
Cited alongside, same era.
Gradient surgery for multi-task learning
Tianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine, Karol Hausman, and Chelsea Finn · 2020
Cited alongside, same era.
On the theory of policy gradient methods: Optimality, approximation, and distribution shift
Alekh Agarwal, Sham M Kakade, Jason D Lee, and Gaurav Mahajan · 2021
Cited alongside, same era.
Conflict-averse gradient descent for multi-task learning
Bo Liu, Xingchao Liu, Xiaojie Jin, Peter Stone, and Qiang Liu · 2021
Cited alongside, same era.
CRPO: A new approach for safe reinforcement learning with convergence guarantee
Tengyu Xu, Yingbin Liang, and Guanghui Lan · 2021
Cited alongside, same era.
MO-Gym: A library of multi-objective reinforcement learning environments
Lucas N. Alegre, Florian Felten, El-Ghazali Talbi, Grégoire Danoy, Ann Nowé, Ana L. C. Bazzan, and Bruno C. da Silva · 2022
Cited alongside, same era.
Achieving zero constraint violation for constrained reinforcement learning via primal-dual approach
Qinbo Bai, Amrit Singh Bedi, Mridul Agarwal, Alec Koppel, and Vaneet Aggarwal · 2022
Cited alongside, same era.
Multi-task learning as a bargaining game
Aviv Navon, Aviv Shamsian, Idan Achituve, Haggai Maron, Kenji Kawaguchi, Gal Chechik, and Ethan Fetaya · 2022
Later among the works it cites.
Sample-efficient multi-objective learning via generalized policy improvement prioritization
Lucas N Alegre, Ana LC Bazzan, Diederik M Roijers, Ann Nowé, and Bruno C da Silva · 2023
Later among the works it cites.
PD-MORL: Preference-driven multi-objective reinforcement learning algorithm
Toygun Basaklar, Suat Gumussoy, and Umit Ogras · 2023
Later among the works it cites.
A toolkit for reliable benchmarking and research in multi-objective reinforcement learning
Florian Felten, Lucas Nunes Alegre, Ann Nowe, Ana L. C. Bazzan, El Ghazali Talbi, Grégoire Danoy, and Bruno Castro da Silva · 2023
Later among the works it cites.
Safety-gymnasium
Jiaming Ji, Borong Zhang, Xuehai Pan, Jiayi Zhou, Juntao Dai, and Yaodong Yang · 2023
Later among the works it cites.
Trust region-based safe distributional reinforcement learning for multiple constraints
Dohyeong Kim, Kyungjae Lee, and Songhwai Oh · 2023
Later among the works it cites.
Multi-objective reinforcement learning: Convexity, stationarity and Pareto optimality
Haoye Lu, Daniel Herman, and Yaoliang Yu · 2023
Later among the works it cites.