Fetching the paper…
Reading the bibliography…
In this paper we consider multi-objective reinforcement learning where the objectives are balanced using preferences.
Extensions of lipschitz mappings into a hilbert space
William B Johnson and Joram Lindenstrauss · 1984
Earlier work this paper cites.
Dynamic preferences in multi-criteria reinforcement learning
Sriraam Natarajan and Prasad Tadepalli · 2005
Earlier work this paper cites.
Empirical bernstein bounds and sample variance penalization
Andreas Maurer and Massimiliano Pontil · 2009
Earlier work this paper cites.
Nearly minimax optimal reward-free reinforcement learning
Zihan Zhang, Simon S Du, and Xiangyang Ji · 2010
Earlier work this paper cites.
The adversarial stochastic shortest path problem with unknown transition probabilities
Gergely Neu, Andras Gyorgy, and Csaba Szepesvári · 2012
Earlier work this paper cites.
A survey of multi-objective sequential decision-making
Diederik M Roijers, Peter Vamplew, Shimon Whiteson, and Richard Dazeley · 2013
Earlier work this paper cites.
A comprehensive survey on safe reinforcement learning
Javier Garcıa and Fernando Fernández · 2015
Earlier work this paper cites.
Multi-objective deep reinforcement learning
Hossam Mossalam, Yannis M Assael, Diederik M Roijers, and Shimon Whiteson · 2016
Earlier work this paper cites.
Constrained policy optimization
Joshua Achiam, David Held, Aviv Tamar, and Pieter Abbeel · 2017
Earlier work this paper cites.
Minimax regret bounds for reinforcement learning
Mohammad Gheshlaghi Azar, Ian Osband, and Rémi Munos · 2017
Earlier work this paper cites.
Unifying pac and regret: Uniform pac bounds for episodic reinforcement learning
Christoph Dann, Tor Lattimore, and Emma Brunskill · 2017
Cited alongside, same era.
Dynamic weights in multi-objective deep reinforcement learning
Axel Abels, Diederik M Roijers, Tom Lenaerts, Ann Nowé, and Denis Steckelmacher · 2018
Cited alongside, same era.
Controlled markov processes with safety state constraints
Mahmoud El Chamie, Yue Yu, Behçet Açıkmeşe, and Masahiro Ono · 2018
Cited alongside, same era.
A general safety framework for learning-based control in uncertain robotic systems
Jaime F Fisac, Anayo K Akametalu, Melanie N Zeilinger, Shahab Kaynama, Jeremy Gillula, and Claire J Tomlin · 2018
Cited alongside, same era.
Is q-learning provably efficient?
Chi Jin, Zeyuan Allen-Zhu, Sebastien Bubeck, and Michael I Jordan · 2018
Cited alongside, same era.
A generalized algorithm for multi-objective reinforcement learning and policy adaptation
Runzhe Yang, Xingyuan Sun, and Karthik Narasimhan · 2019
Later among the works it cites.
Tighter problem-dependent regret bounds in reinforcement learning without domain knowledge using value function bounds
Andrea Zanette and Emma Brunskill · 2019
Later among the works it cites.
Constrained episodic reinforcement learning in concave-convex and knapsack settings
Kianté Brantley, Miroslav Dudik, Thodoris Lykouris, Sobhan Miryoosefi, Max Simchowitz, Aleksandrs Slivkins, and Wen Sun · 2020
Closest in time.
Exploration-exploitation in constrained mdps
Yonathan Efroni, Shie Mannor, and Matteo Pirotta · 2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Cited alongside, same era.
Wang Chi Cheung · 2019
Cited alongside, same era.
Learning adversarial mdps with bandit feedback and unknown transition, 2019
Chi Jin, Tiancheng Jin, Haipeng Luo, Suvrit Sra, and Tiancheng Yu · 2019
Cited alongside, same era.
Online convex optimization in adversarial markov decision processes
Aviv Rosenberg and Yishay Mansour · 2019
Cited alongside, same era.
Task-agnostic exploration in reinforcement learning, 2020a
Xuezhou Zhang, Yuzhe Ma, and Adish Singla
Cited in the paper.
Chi Jin, Akshay Krishnamurthy, Max Simchowitz, and Tiancheng Yu · 2020
Closest in time.
Adaptive reward-free exploration
Emilie Kaufmann, Pierre Ménard, Omar Darwiche Domingues, Anders Jonsson, Edouard Leurent, and Michal Valko · 2020
Closest in time.
Fast active learning for pure exploration in reinforcement learning
Pierre Ménard, Omar Darwiche Domingues, Anders Jonsson, Emilie Kaufmann, Edouard Leurent, and Michal Valko · 2020
Closest in time.
Safe reinforcement learning in constrained markov decision processes
Akifumi Wachi and Yanan Sui · 2020
Closest in time.
On reward-free reinforcement learning with linear function approximation
Ruosong Wang, Simon S Du, Lin F Yang, and Ruslan Salakhutdinov · 2020
Closest in time.