Fetching the paper…
Reading the bibliography…
We consider the problem of making AI agents that collaborate well with humans in partially observable fully cooperative environments given datasets of human behavior.
Aviral Kumar, Xue Bin Peng, and Sergey Levine · 1912
Earlier work this paper cites.
A framework for behavioural cloning
Michael Bain and Claude Sammut · 1996
Earlier work this paper cites.
Cognition and behavior in two-person guessing games: An experimental study
Miguel Costa-Gomes and Vincent P. Crawford · 2006
Earlier work this paper cites.
Decentralized stochastic control with partial history sharing: A common information approach
Ashutosh Nayyar, Aditya Mahajan, and Demosthenis Teneketzis · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Tom Schaul, John Quan, Ioannis Antonoglou, and David Silver · 2015
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Mastering the game of go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, et al · 2017
Earlier work this paper cites.
Value-decomposition networks for cooperative multi-agent learning
Peter Sunehag, Guy Lever, Audrunas Gruslys, Wojciech Marian Czarnecki, Vinícius Flores Zambaldi, Max Jaderberg, Marc Lanctot, Nicolas Sonnerat, Joel Z. Leibo, Karl Tuyls, and Thore Graepel · 2017
Earlier work this paper cites.
QMIX: monotonic value function factorisation for deep multi-agent reinforcement learning
Tabish Rashid, Mikayel Samvelyan, Christian Schröder de Witt, Gregory Farquhar, Jakob N. Foerster, and Shimon Whiteson · 2018
Cited alongside, same era.
A general reinforcement learning algorithm that masters chess, shogi, and go through self-play
David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, et al · 2018
Cited alongside, same era.
Superhuman AI for multiplayer poker
Noam Brown and Tuomas Sandholm · 2019
Cited alongside, same era.
On the utility of learning about humans for human-ai coordination
Micah Carroll, Rohin Shah, Mark K Ho, Tom Griffiths, Sanjit Seshia, Pieter Abbeel, and Anca Dragan · 2019
Cited alongside, same era.
Recurrent experience replay in distributed reinforcement learning
Steven Kapturowski, Georg Ostrovski, John Quan, Rémi Munos, and Will Dabney · 2019
Cited alongside, same era.
Learning personalized models of human behavior in chess
Reid McIlroy-Young, Russell Wang, Siddhartha Sen, Jon Kleinberg, and Ashton Anderson · 2020
Later among the works it cites.
No-press diplomacy from scratch
Anton Bakhtin, David Wu, Adam Lerer, and Noam Brown · 2021
Later among the works it cites.
Decision transformer: Reinforcement learning via sequence modeling
Lili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee, Aditya Grover, Michael Laskin, Pieter Abbeel, Aravind Srinivas, and Igor Mordatch · 2021
Later among the works it cites.
K-level reasoning for zero-shot coordination in hanabi
Brandon Cui, Hengyuan Hu, Luis Pineda, and Jakob Foerster · 2021
Later among the works it cites.
Trajectory diversity for zero-shot coordination
Andrei Lupu, Brandon Cui, Hengyuan Hu, and Jakob Foerster · 2021
Later among the works it cites.
Collaborating with humans without human data
DJ Strouse, Kevin R. McKee, Matthew Botvinick, Edward Hughes, and Richard Everett · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning existing social conventions via observationally augmented self-play
Adam Lerer and Alexander Peysakhovich · 2019
Cited alongside, same era.
The hanabi challenge: A new frontier for ai research
Nolan Bard, Jakob N Foerster, Sarath Chandar, Neil Burch, Marc Lanctot, H Francis Song, Emilio Parisotto, Vincent Dumoulin, Subhodeep Moitra, Edward Hughes, et al · 2020
Cited alongside, same era.
“other-play” for zero-shot coordination
Hengyuan Hu, Adam Lerer, Alex Peysakhovich, and Jakob Foerster · 2020
Cited alongside, same era.
Improving policies via search in cooperative partially observable games
Adam Lerer, Hengyuan Hu, Jakob Foerster, and Noam Brown · 2020
Cited alongside, same era.
K-level resoning for zero-shot coordination in hanabi
Brandon Cui, Hengyuan Hu, Luis Pineda, and Jakob Foerster
Cited in the paper.
Learned belief search: Efficiently improving policies in partially observable settings
Hengyuan Hu, Adam Lerer, Noam Brown, and Jakob Foerster
Cited in the paper.
Learned belief search: Efficiently improving policies in partially observable settings
Hengyuan Hu, Adam Lerer, Noam Brown, and Jakob N. Foerster
Cited in the paper.
Later among the works it cites.
Discovering diverse multi-agent strategic behavior via reward randomization
Zhenggang Tang, Chao Yu, Boyuan Chen, Huazhe Xu, Xiaolong Wang, Fei Fang, Simon Shaolei Du, Yu Wang, and Yi Wu · 2021
Later among the works it cites.
The surprising effectiveness of MAPPO in cooperative, multi-agent games
Chao Yu, Akash Velu, Eugene Vinitsky, Yu Wang, Alexandre M. Bayen, and Yi Wu · 2021
Later among the works it cites.
Modeling strong and human-like gameplay with KL-regularized search
Athul Paul Jacob, David J Wu, Gabriele Farina, Adam Lerer, Hengyuan Hu, Anton Bakhtin, Jacob Andreas, and Noam Brown · 2022
Closest in time.