Fetching the paper…
Reading the bibliography…
Most of the prior work on multi-agent reinforcement learning (MARL) achieves optimal collaboration by directly controlling the agents to maximize a common reward.
Moral hazard and observability
Bengt Hölmstrom · 1979
Earlier work this paper cites.
Optimal auction design
Roger B. Myerson · 1981
Earlier work this paper cites.
Optimal coordination mechanisms in generalized principal–agent problems
Roger B. Myerson · 1982
Earlier work this paper cites.
Multitask principal-agent analyses: Incentive contracts, asset ownership, and job design
Bengt Hölmstrom and Paul Milgrom · 1991
Earlier work this paper cites.
Markov games as a framework for multi-agent reinforcement learning
Michael L Littman · 1994
Earlier work this paper cites.
Computing factored value functions for policies in structured mdps
Daphne Koller and Ronald Parr · 1999
Earlier work this paper cites.
Multiagent planning with factored mdps
Carlos Guestrin, Daphne Koller, and Ronald Parr · 2001
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
Peter Auer, Nicolo Cesa-Bianchi, and Paul Fischer · 2002
Earlier work this paper cites.
Complexity of mechanism design
Vincent Conitzer and Tuomas Sandholm · 2002
Earlier work this paper cites.
The theory of incentives: the principal-agent model
Jean-Jacques Laffont and D. Martimort · 2002
Earlier work this paper cites.
Coordination and adaptation in impromptu teams
Michael Bowling and Peter McCracken · 2005
Earlier work this paper cites.
A comprehensive survey of multiagent reinforcement learning
Lucian Busoniu, Robert Babuska, and Bart De Schutter · 2008
Earlier work this paper cites.
Optimal and approximate q-value functions for decentralized pomdps
Frans A. Oliehoek, Matthijs TJ Spaan, and Nikos Vlassis · 2008
Earlier work this paper cites.
A continuous-time version of the principal-agent problem
Yuliy Sannikov · 2008
Earlier work this paper cites.
Value-based policy teaching with active indirect elicitation
Haoqi Zhang and David Parkes · 2008
Earlier work this paper cites.
Action understanding as inverse planning
Chris L Baker, Rebecca Saxe, and Joshua B Tenenbaum · 2009
Cited alongside, same era.
Policy teaching through reward function learning
Haoqi Zhang, David C. Parkes, and Yiling Chen · 2009
Cited alongside, same era.
Reward design via online gradient ascent
Jonathan Sorg, Richard L. Lewis, and Satinder P. Singh · 2010
Cited alongside, same era.
Ad hoc autonomous agent teams: Collaboration without pre-coordination
Peter Stone, Gal A. Kaminka, Sarit Kraus, and Jeffrey S. Rosenschein · 2010
Cited alongside, same era.
Lecture 6.5—rmsprop: Divide the gradient by a running average of its recent magnitude
Tijmen Tieleman and Geoffrey Hinto · 2012
Cited alongside, same era.
Gradient-based hyperparameter optimization through reversible learning
Dougal Maclaurin, David Duvenaud, and Ryan P. Adams · 2015
Cited alongside, same era.
Model-agnostic meta-learning for fast adaptation of deep networks
Chelsea Finn, Pieter Abbeel, and Sergey Levine · 2017
Later among the works it cites.
Low-shot visual recognition by shrinking and hallucinating features
Bharath Hariharan and Ross Girshick · 2017
Later among the works it cites.
Multi-agent actor-critic for mixed cooperative-competitive environments
Ryan Lowe, Yi Wu, Aviv Tamar, Jean Harb, Pieter Abbeel, and Igor Mordatch · 2017
Later among the works it cites.
Deep decentralized multi-task multi-agent reinforcement learning under partial observability
Shayegan Omidshafiei, Jason Pazis, Christopher Amato, Jonathan P. How, and John Vian · 2017
Later among the works it cites.
Peng Peng, Ying Wen, Yaodong Yang, Yuan Quan, Zhenkun Tang, Haitao Long, and Jun Wang · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning to communicate with deep multi-agent reinforcement learning
Jakob Foerster, Ioannis Alexandros Assael, Nando de Freitas, and Shimon Whiteson · 2016
Cited alongside, same era.
Deep successor reinforcement learning
Tejas D. Kulkarni, Ardavan Saeedi, Simanta Gautam, and Samuel J Gershman · 2016
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Cited alongside, same era.
Learning to reinforcement learn
Jane X Wang, Zeb Kurth-Nelson, Dhruva Tirumala, Hubert Soyer, Joel Z Leibo, Remi Munos, Charles Blundell, Dharshan Kumaran, and Matt Botvinick · 2016
Cited alongside, same era.
Modular multitask reinforcement learning with policy sketches
Jacob Andreas, Dan Klein, and Sergey Levine · 2017
Cited alongside, same era.
Designing neural network architectures using reinforcement learning
Bowen Baker, Otkrist Gupta, Nikhil Naik, and Ramesh Raskar · 2017
Cited alongside, same era.
Learned optimizers that scale and generalize
Olga Wichrowska, Niru Maheswaranathan, Matthew W Hoffman, Sergio Gomez Colmenarejo, Misha Denil, Nando de Freitas, and Jascha Sohl-Dickstein · 2017
Later among the works it cites.
Visual semantic planning using deep successor representations
Yuke Zhu, Daniel Gordon, Eric Kolve, Dieter Fox, Li Fei-Fei, Abhinav Gupta, Roozbeh Mottaghi, and Ali Farhadi · 2017
Later among the works it cites.
Continuous adaptation via meta-learning in nonstationary and competitive environments
Maruan Al-Shedivat, Trapit Bansal, Yuri Burda, Ilya Sutskever, Igor Mordatch, and Pieter Abbeel · 2018
Closest in time.
Universal successor representations for transfer reinforcement learning
Chen Ma, Junfeng Wen, and Yoshua Bengio · 2018
Closest in time.
Neil C. Rabinowitz, Frank Perbet, H. Francis Song, Chiyuan Zhang, S.M. Ali Eslami, and Matthew Botvinick · 2018
Closest in time.
Qmix: Monotonic value function factorisation for deep multi-agent reinforcement learning
Tabish Rashid, Mikayel Samvelyan, Christian Schroeder de Witt, Gregory Farquhar, Jakob Foerster, and Shimon Whiteson · 2018
Closest in time.
Simplifying reward design through divide-and-conquer
Ellis Ratner, Dylan Hadfield-Menell, and Anca D. Dragan · 2018
Closest in time.
Value-decomposition networks for cooperative multi-agent learning based on team reward
Peter Sunehag, Guy Lever, Audrunas Gruslys, Wojciech Marian Czarnecki, Vinicius Zambaldi, Max Jaderberg, Marc Lanctot, Nicolas SOnnerat, Joel Z. Leibo, Karl Tuyls, and Thore Graepel · 2018
Closest in time.
One-shot imitation from observing humans via domain-adaptive meta-learning
Tianhe Yu, Chelsea Finn, Annie Xie, Sudeep Dasari, Pieter Abbeel, and Sergey Levine · 2018
Closest in time.