Fetching the paper…
Reading the bibliography…
We study multi-objective reinforcement learning (RL) where an agent's reward is represented as a vector.
Zur theorie der gesellschaftsspiele
J v Neumann · 1928
Earlier work this paper cites.
Stochastic games
Lloyd S Shapley · 1953
Earlier work this paper cites.
An analog of the minimax theorem for vector payoffs
David Blackwell · 1956
Earlier work this paper cites.
Convex analysis
R Tyrrell Rockafellar · 1970
Earlier work this paper cites.
Markov games as a framework for multi-agent reinforcement learning
Michael L Littman · 1994
Earlier work this paper cites.
R-max-a general polynomial time algorithm for near-optimal reinforcement learning
Ronen I Brafman and Moshe Tennenholtz · 2002
Earlier work this paper cites.
Empirical bernstein bounds and sample variance penalization
Andreas Maurer and Massimiliano Pontil · 2009
Earlier work this paper cites.
Provably efficient online agnostic learning in markov games
Yi Tian, Yuanhao Wang, Tiancheng Yu, and Suvrit Sra · 2010
Earlier work this paper cites.
Blackwell approachability and no-regret learning are equivalent
Jacob Abernethy, Peter L Bartlett, and Elad Hazan · 2011
Earlier work this paper cites.
Bandits with knapsacks
Ashwinkumar Badanidiyuru, Robert Kleinberg, and Aleksandrs Slivkins · 2013
Earlier work this paper cites.
Bandits with concave rewards and convex knapsacks
Shipra Agrawal and Nikhil R Devanur · 2014
Earlier work this paper cites.
Approachability in unknown games: Online learning meets multi-objective optimization
Shie Mannor, Vianney Perchet, and Gilles Stoltz · 2014
Cited alongside, same era.
Near-optimal reinforcement learning in factored mdps
Ian Osband and Benjamin Van Roy · 2014
Cited alongside, same era.
Adaptive algorithms for online convex optimization with long-term constraints
Rodolphe Jenatton, Jim Huang, and Cédric Archambeau · 2016
Cited alongside, same era.
An online convex optimization approach to blackwell’s approachability
Nahum Shimkin · 2016
Cited alongside, same era.
Minimax regret bounds for reinforcement learning
Mohammad Gheshlaghi Azar, Ian Osband, and Rémi Munos · 2017
Cited alongside, same era.
Is Q-learning provably efficient?
Chi Jin, Zeyuan Allen-Zhu, Sebastien Bubeck, and Michael I Jordan · 2018
Cited alongside, same era.
Efficient reinforcement learning in factored mdps with application to constrained rl
Xiaoyu Chen, Jiachen Hu, Lihong Li, and Liwei Wang · 2020
Later among the works it cites.
Provably efficient safe exploration via primal-dual policy optimization
Dongsheng Ding, Xiaohan Wei, Zhuoran Yang, Zhaoran Wang, and Mihailo R Jovanović · 2020
Later among the works it cites.
Episodic reinforcement learning in finite mdps: Minimax lower bounds revisited
Omar Darwiche Domingues, Pierre Ménard, Emilie Kaufmann, and Michal Valko · 2020
Later among the works it cites.
Exploration-exploitation in constrained mdps
Yonathan Efroni, Shie Mannor, and Matteo Pirotta · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Online convex optimization for cumulative constraints
Jianjun Yuan and Andrew Lamperski · 2018
Cited alongside, same era.
Non-stationary reinforcement learning: The blessing of (more) optimism
Wang Chi Cheung, David Simchi-Levi, and Ruihao Zhu · 2019
Cited alongside, same era.
Provable self-play algorithms for competitive reinforcement learning
Yu Bai and Chi Jin · 2020
Cited alongside, same era.
Near-optimal reinforcement learning with self-play
Yu Bai, Chi Jin, and Tiancheng Yu · 2020
Cited alongside, same era.
Constrained episodic reinforcement learning in concave-convex and knapsack settings
Kianté Brantley, Miroslav Dudik, Thodoris Lykouris, Sobhan Miryoosefi, Max Simchowitz, Aleksandrs Slivkins, and Wen Sun · 2020
Cited alongside, same era.
Towards minimax optimal reinforcement learning in factored markov decision processes
Yi Tian, Jian Qian, and Suvrit Sra
Cited in the paper.
Chi Jin, Akshay Krishnamurthy, Max Simchowitz, and Tiancheng Yu · 2020
Later among the works it cites.
A sharp analysis of model-based reinforcement learning with self-play
Qinghua Liu, Tiancheng Yu, Yu Bai, and Chi Jin · 2020
Later among the works it cites.
Upper confidence primal-dual reinforcement learning for cmdp with adversarial loss
Shuang Qiu, Xiaohan Wei, Zhuoran Yang, Jieping Ye, and Zhaoran Wang · 2020
Later among the works it cites.
Learning in markov decision processes under constraints
Rahul Singh, Abhishek Gupta, and Ness B Shroff · 2020
Later among the works it cites.
Jingfeng Wu, Vladimir Braverman, and Lin F Yang · 2020
Later among the works it cites.
Qiaomin Xie, Yudong Chen, Zhaoran Wang, and Zhuoran Yang · 2020
Later among the works it cites.