Fetching the paper…
Reading the bibliography…
We study policy optimization for Markov decision processes (MDPs) with multiple reward value functions, which are to be jointly optimized according to given criteria such as proportional fairness (smooth concave scalarization), hard constraints (constrained MDP), and max-min trade-off.
Rate control for communication networks: shadow prices, proportional fairness and stability
Frank P Kelly, Aman K Maulloo, and David Kim Hong Tan · 1998
Earlier work this paper cites.
Constrained Markov decision processes
Eitan Altman · 1999
Earlier work this paper cites.
Mirror descent and nonlinear projected subgradient methods for convex optimization
Amir Beck and Marc Teboulle · 2003
Earlier work this paper cites.
Prox-method with rate of convergence o (1/t) for variational inequalities with lipschitz continuous monotone operators and smooth convex-concave saddle point problems
Arkadi Nemirovski · 2004
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Earlier work this paper cites.
Optimization, learning, and games with predictable sequences
Sasha Rakhlin and Karthik Sridharan · 2013
Earlier work this paper cites.
A survey of multi-objective sequential decision-making
Diederik M Roijers, Peter Vamplew, Shimon Whiteson, and Richard Dazeley · 2013
Earlier work this paper cites.
Scalarized multi-objective reinforcement learning: Novel design techniques
Kristof Van Moffaert, Madalina M Drugan, and Ann Nowé · 2013
Earlier work this paper cites.
Multi-objective mdps with conditional lexicographic reward preferences
Kyle Hollins Wray, Shlomo Zilberstein, and Abdel-Illah Mouaddib · 2015
Earlier work this paper cites.
Openai gym, 2016
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Earlier work this paper cites.
Hao Yu and Michael J Neely · 2017
Earlier work this paper cites.
A simple parallel algorithm with an O(1/t) convergence rate for general convex programs
Hao Yu and Michael J Neely · 2017
Earlier work this paper cites.
Reinforcement learning for joint optimization of multiple rewards
Mridul Agarwal and Vaneet Aggarwal · 2019
Earlier work this paper cites.
A generalized algorithm for multi-objective reinforcement learning and policy adaptation
Runzhe Yang, Xingyuan Sun, and Karthik Narasimhan · 2019
Cited alongside, same era.
Optimality and approximation with policy gradient methods in markov decision processes
Alekh Agarwal, Sham M Kakade, Jason D Lee, and Gaurav Mahajan · 2020
Cited alongside, same era.
Natural policy gradient primal-dual method for constrained markov decision processes
Dongsheng Ding, Kaiqing Zhang, Tamer Basar, and Mihailo R Jovanovic · 2020
Cited alongside, same era.
First-order and Stochastic Optimization Methods for Machine Learning
Guanghui Lan · 2020
Cited alongside, same era.
On the global convergence rates of softmax policy gradient methods
Jincheng Mei, Chenjun Xiao, Csaba Szepesvari, and Dale Schuurmans · 2020
Cited alongside, same era.
Linear last-iterate convergence in constrained saddle-point optimization
Achieving zero constraint violation for constrained reinforcement learning via primal-dual approach
Qinbo Bai, Amrit Singh Bedi, Mridul Agarwal, Alec Koppel, and Vaneet Aggarwal · 2021
Later among the works it cites.
Data networks
Dimitri Bertsekas and Robert Gallager · 2021
Later among the works it cites.
Fast global convergence of natural policy gradient methods with entropy regularization
Shicong Cen, Chen Cheng, Yuxin Chen, Yuting Wei, and Yuejie Chi · 2021
Later among the works it cites.
On the linear convergence of natural policy gradient algorithm
Sajad Khodadadian, Prakirt Raj Jhunjhunwala, Sushil Mahavir Varma, and Siva Theja Maguluri · 2021
Later among the works it cites.
Guanghui Lan · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chen-Yu Wei, Chung-Wei Lee, Mengxiao Zhang, and Haipeng Luo · 2020
Cited alongside, same era.
Online primal-dual mirror descent under stochastic constraints
Xiaohan Wei, Hao Yu, and Michael J Neely · 2020
Cited alongside, same era.
Jingfeng Wu, Vladimir Braverman, and Lin F Yang · 2020
Cited alongside, same era.
Variational policy gradient method for reinforcement learning with general utilities
Junyu Zhang, Alec Koppel, Amrit Singh Bedi, Csaba Szepesvari, and Mengdi Wang · 2020
Cited alongside, same era.
First order constrained optimization in policy space
Yiming Zhang, Quan Vuong, and Keith Ross · 2020
Cited alongside, same era.
Provable multi-objective reinforcement learning with generative models
Dongruo Zhou, Jiahao Chen, and Quanquan Gu · 2020
Cited alongside, same era.
On the theory of policy gradient methods: Optimality, approximation, and distribution shift
Alekh Agarwal, Sham M Kakade, Jason D Lee, and Gaurav Mahajan · 2021
Cited alongside, same era.
Later among the works it cites.
Faster algorithm and sharper analysis for constrained markov decision process
Tianjiao Li, Ziwei Guan, Shaofeng Zou, Tengyu Xu, Yingbin Liang, and Guanghui Lan · 2021
Later among the works it cites.
Learning policies with zero or bounded constraint violation for constrained mdps
Tao Liu, Ruida Zhou, Dileep Kalathil, Panganamala Kumar, and Chao Tian · 2021
Later among the works it cites.
A multi-objective approach to mitigate negative side effects
Sandhya Saisubramanian, Ece Kamar, and Shlomo Zilberstein · 2021
Later among the works it cites.
CRPO: A new approach for safe reinforcement learning with convergence guarantee
Tengyu Xu, Yingbin Liang, and Guanghui Lan · 2021
Later among the works it cites.
A dual approach to constrained markov decision processes with entropy regularization
Donghao Ying, Yuhao Ding, and Javad Lavaei · 2021
Later among the works it cites.
Beyond cumulative returns via reinforcement learning over state-action occupancy measures
Junyu Zhang, Amrit Singh Bedi, Mengdi Wang, and Alec Koppel · 2021
Later among the works it cites.
Towards painless policy optimization for constrained mdps
Arushi Jain, Sharan Vaswani, Reza Babanezhad, Csaba Szepesvari, and Doina Precup · 2022
Closest in time.
Pareto policy adaptation
Panagiotis Kyriakis, Jyotirmoy Deshmukh, and Paul Bogdan · 2022
Closest in time.