Fetching the paper…
Reading the bibliography…
Existing multi-objective reinforcement learning (MORL) algorithms do not account for objectives that arise from players with differing beliefs.
The bargaining problem
John F Nash · 1950
Earlier work this paper cites.
On informationally decentralized systems
Leonid Hurwicz · 1972
Earlier work this paper cites.
Incentive compatibility and the bargaining problem
Roger B Myerson · 1979
Earlier work this paper cites.
Cardinal welfare, individualistic ethics, and interpersonal comparisons of utility
John C Harsanyi · 1980
Earlier work this paper cites.
Efficient mechanisms for bilateral trading
Roger B Myerson and Mark A Satterthwaite · 1983
Earlier work this paper cites.
Multi-criteria reinforcement learning
Zoltán Gábor, Zsolt Kalmár, and Csaba Szepesvári · 1998
Earlier work this paper cites.
Learning agents for uncertain environments
Stuart Russell · 1998
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
Andrew Y Ng, Stuart J Russell, et al · 2000
Earlier work this paper cites.
A gentle introduction to the universal algorithmic agent { \{ AIXI } \} , 2003
Marcus Hutter · 2003
Earlier work this paper cites.
Artificial intelligence: a modern approach (Chapter 17.1) , volume 2
Stuart Russell, Peter Norvig, John F Canny, Jitendra M Malik, and Douglas D Edwards · 2003
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
Pieter Abbeel and Andrew Y Ng · 2004
Cited alongside, same era.
Multiagent systems: Algorithmic, game-theoretic, and logical foundations
Yoav Shoham and Kevin Leyton-Brown · 2008
Cited alongside, same era.
Modeling and reasoning with Bayesian networks (Chapter 4)
Adnan Darwiche · 2009
Cited alongside, same era.
Causality
Judea Pearl · 2009
Cited alongside, same era.
Evolving policies for multi-reward partially observable markov decision processes (mr-pomdps)
Harold Soh and Yiannis Demiris · 2011
Cited alongside, same era.
Multiple attribute decision making: methods and applications
Gwo-Hshiung Tzeng and Jih-Jeng Huang · 2011
Cited alongside, same era.
Fairness in multi-agent sequential decision-making
Chongjie Zhang and Julie A Shah · 2014
Later among the works it cites.
Reflective variants of solomonoff induction and aixi
Benja Fallenstein, Nate Soares, and Jessica Taylor · 2015
Later among the works it cites.
Point-based planning for multi-objective pomdps
Diederik M Roijers, Shimon Whiteson, and Frans A Oliehoek · 2015
Later among the works it cites.
Toward idealized decision theory
Nate Soares and Benja Fallenstein · 2015
Later among the works it cites.
Multi-objective pomdps with lexicographic reward preferences
Kyle Hollins Wray and Shlomo Zilberstein · 2015
Later among the works it cites.
Racing to the precipice: a model of artificial intelligence development
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Space-time embedded intelligence
Laurent Orseau and Mark Ring · 2012
Cited alongside, same era.
Game theory
Roger B Myerson · 2013
Cited alongside, same era.
Superintelligence: Paths, dangers, strategies
Nick Bostrom · 2014
Cited alongside, same era.
Multi-objective sequential decision making
Weijia Wang · 2014
Cited alongside, same era.
Stuart Armstrong, Nick Bostrom, and Carl Shulman · 2016
Later among the works it cites.
On the promotion of safe and socially beneficial artificial intelligence
Seth D Baum · 2016
Later among the works it cites.
Parametric bounded lob’s theorem and robust cooperation of bounded agents
Andrew Critch · 2016
Later among the works it cites.
Scott Garrabrant, Tsvi Benson-Tilsen, Andrew Critch, Nate Soares, and Jessica Taylor · 2016
Later among the works it cites.
Cooperative inverse reinforcement learning, 2016
Dylan Hadfield-Menell, Anca Dragan, Pieter Abbeel, and Stuart Russell · 2016
Later among the works it cites.