Fetching the paper…
Reading the bibliography…
There is a recent trend of applying multi-agent reinforcement learning (MARL) to train an agent that can cooperate with humans in a zero-shot fashion without using any human data.
Multi-attribute utility theory: models and assessment procedures
Detlof Von Winterfeldt and Gregory W Fischer · 1975
Earlier work this paper cites.
Risk aversion in the small and in the large
John W Pratt · 1978
Earlier work this paper cites.
Bounded rationality
Reinhard Selten · 1990
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
Andrew Y Ng and Stuart Russell · 2000
Earlier work this paper cites.
A cognitive hierarchy model of games
Colin F Camerer, Teck-Hua Ho, and Juin-Kuan Chong · 2004
Earlier work this paper cites.
Ten challenges for making automation a "team player" in joint human-agent activity
Glen Klien, David D Woods, Jeffrey M Bradshaw, Robert R Hoffman, and Paul J Feltovich · 2004
Earlier work this paper cites.
Security in multiagent systems by policy randomization
Praveen Paruchuri, Milind Tambe, Fernando Ordónez, and Sarit Kraus · 2006
Earlier work this paper cites.
Robots at home: Understanding long-term human-robot interaction
Cory D Kidd and Cynthia Breazeal · 2008
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
Brian D Ziebart, Andrew L Maas, J Andrew Bagnell, Anind K Dey, et al · 2008
Earlier work this paper cites.
Ad hoc autonomous agent teams: Collaboration without pre-coordination
Peter Stone, Gal A Kaminka, Sarit Kraus, and Jeffrey S Rosenschein · 2010
Earlier work this paper cites.
Behavioral game theory: Experiments in strategic interaction
Colin F Camerer · 2011
Earlier work this paper cites.
Thirty years of prospect theory in economics: A review and assessment
Nicholas C Barberis · 2013
Earlier work this paper cites.
Security games with interval uncertainty
Christopher Kiekintveld, Towhidul Islam, and Vladik Kreinovich · 2013
Earlier work this paper cites.
Analyzing the effectiveness of adversary modeling in security games
Thanh Nguyen, Rong Yang, Amos Azaria, Sarit Kraus, and Milind Tambe · 2013
Earlier work this paper cites.
Robots that can adapt like animals
Antoine Cully, Jeff Clune, Danesh Tarapore, and Jean-Baptiste Mouret · 2015
Earlier work this paper cites.
Manifesto for a new (computational) cognitive revolution
Thomas L Griffiths · 2015
Earlier work this paper cites.
A survey on handling computationally expensive multiobjective optimization problems using surrogates: non-nature inspired methods
Mohammad Tabatabaei, Jussi Hakanen, Markus Hartikainen, Kaisa Miettinen, and Karthik Sindhya · 2015
Earlier work this paper cites.
Maximum entropy deep inverse reinforcement learning
Markus Wulfmeier, Peter Ondruska, and Ingmar Posner · 2015
Earlier work this paper cites.
Learning the preferences of ignorant, inconsistent agents
Owain Evans, Andreas Stuhlmüller, and Noah Goodman · 2016
Earlier work this paper cites.
Cooperative inverse reinforcement learning
Dylan Hadfield-Menell, Stuart J Russell, Pieter Abbeel, and Anca Dragan · 2016
Earlier work this paper cites.
Quality diversity: A new frontier for evolutionary computation
Justin K Pugh, Lisa B Soros, and Kenneth O Stanley · 2016
Earlier work this paper cites.
Human-robot interaction: status and challenges
Thomas B Sheridan · 2016
Earlier work this paper cites.
Successor features for transfer in reinforcement learning
Andre Barreto, Will Dabney, Remi Munos, Jonathan J Hunt, Tom Schaul, Hado P van Hasselt, and David Silver · 2017
Earlier work this paper cites.
A unified game-theoretic approach to multiagent reinforcement learning
Marc Lanctot, Vinicius Zambaldi, Audrunas Gruslys, Angeliki Lazaridou, Karl Tuyls, Julien Pérolat, David Silver, and Thore Graepel · 2017
Earlier work this paper cites.
Multi-agent actor-critic for mixed cooperative-competitive environments
Ryan Lowe, Yi Wu, Aviv Tamar, Jean Harb, Pieter Abbeel, and Igor Mordatch · 2017
Earlier work this paper cites.
An introduction to behavioral economics
Nick Wilkinson and Matthias Klaes · 2017
Earlier work this paper cites.
Progress and prospects of the human-robot collaboration
Arash Ajoudani, Andrea Maria Zanchettin, Serena Ivaldi, Alin Albu-Schäffer, Kazuhiro Kosuge, and Oussama Khatib · 2018
Cited alongside, same era.
Variational inverse control with events: A general framework for data-driven reward definition
Justin Fu, Avi Singh, Dibya Ghosh, Larry Yang, and Sergey Levine · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Cited alongside, same era.
Qmix: Monotonic value function factorisation for deep multi-agent reinforcement learning
Tabish Rashid, Mikayel Samvelyan, Christian Schroeder, Gregory Farquhar, Jakob Foerster, and Shimon Whiteson · 2018
Cited alongside, same era.
A general reinforcement learning algorithm that masters chess, shogi, and go through self-play
David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, et al · 2018
Cited alongside, same era.
Discovering diverse multi-agent strategic behavior via reward randomization
Zhenggang Tang, Chao Yu, Boyuan Chen, Huazhe Xu, Xiaolong Wang, Fei Fang, Simon Shaolei Du, Yu Wang, and Yi Wu · 2020
Later among the works it cites.
Multi-agent collaboration via reward attribution decomposition
Tianjun Zhang, Huazhe Xu, Xiaolong Wang, Yi Wu, Kurt Keutzer, Joseph E Gonzalez, and Yuandong Tian · 2020
Later among the works it cites.
Learning to cooperate with unseen agents through meta-reinforcement learning
Rujikorn Charakorn, Poramate Manoonpong, and Nat Dilokthanakul · 2021
Later among the works it cites.
K-level reasoning for zero-shot coordination in hanabi
Brandon Cui, Hengyuan Hu, Luis Pineda, and Jakob Foerster · 2021
Later among the works it cites.
Cooperative AI: machines must learn to find common ground, 2021
Allan Dafoe, Yoram Bachrach, Gillian Hadfield, Eric Horvitz, Kate Larson, and Thore Graepel · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Cited alongside, same era.
Dota 2 with large scale deep reinforcement learning
Christopher Berner, Greg Brockman, Brooke Chan, Vicki Cheung, Przemysław Dębiak, Christy Dennison, David Farhi, Quirin Fischer, Shariq Hashme, Chris Hesse, et al · 2019
Cited alongside, same era.
Human–vehicle cooperation in automated driving: A multidisciplinary review and appraisal
Francesco Biondi, Ignacio Alvarez, and Kyeong-Ah Jeong · 2019
Cited alongside, same era.
On the utility of learning about humans for human-AI coordination
Micah Carroll, Rohin Shah, Mark K Ho, Tom Griffiths, Sanjit Seshia, Pieter Abbeel, and Anca Dragan · 2019
Cited alongside, same era.
A survey on handling computationally expensive multiobjective optimization problems with evolutionary algorithms
Tinkle Chugh, Karthik Sindhya, Jussi Hakanen, and Kaisa Miettinen · 2019
Cited alongside, same era.
Impossibility and uncertainty theorems in ai value alignment (or why your agi should not have a utility function)
Peter Eckersley · 2019
Cited alongside, same era.
Robust multi-agent reinforcement learning via minimax deep deterministic policy gradient
Shihui Li, Yi Wu, Xinyue Cui, Honghua Dong, Fei Fang, and Stuart Russell · 2019
Cited alongside, same era.
Introduction to game theory
Fei Fang, Shutian Liu, Anjon Basak, Quanyan Zhu, Christopher D Kiekintveld, and Charles A Kamhoua · 2021
Later among the works it cites.
Pick your battles: Interaction graphs as population-level objectives for strategic diversity
Marta Garnelo, Wojciech Marian Czarnecki, Siqi Liu, Dhruva Tirumala, Junhyuk Oh, Gauthier Gidel, Hado van Hasselt, and David Balduzzi · 2021
Later among the works it cites.
Dynamic population-based meta-learning for multi-agent communication with natural language
Abhinav Gupta, Marc Lanctot, and Angeliki Lazaridou · 2021
Later among the works it cites.
Off-belief learning
Hengyuan Hu, Adam Lerer, Brandon Cui, Luis Pineda, Noam Brown, and Jakob Foerster · 2021
Later among the works it cites.
A survey of generalisation in deep reinforcement learning
Robert Kirk, Amy Zhang, Edward Grefenstette, and Tim Rocktäschel · 2021
Later among the works it cites.
Evaluating the robustness of collaborative agents
Paul Knott, Micah Carroll, Sam Devlin, Kamil Ciosek, Katja Hofmann, Anca Dragan, and Rohin Shah · 2021
Later among the works it cites.
Formalizing and guaranteeing human-robot interaction
Hadas Kress-Gazit, Kerstin Eder, Guy Hoffman, Henny Admoni, Brenna Argall, Ruediger Ehlers, Christoffer Heckman, Nils Jansen, Ross Knepper, Jan Křetínskỳ, et al · 2021
Later among the works it cites.
Towards unifying behavioral and response diversity for open-ended learning in zero-sum games
Xiangyu Liu, Hangtian Jia, Ying Wen, Yujing Hu, Yingfeng Chen, Changjie Fan, Zhipeng Hu, and Yaodong Yang · 2021
Later among the works it cites.
Trajectory diversity for zero-shot coordination
Andrei Lupu, Brandon Cui, Hengyuan Hu, and Jakob Foerster · 2021
Later among the works it cites.
Continuous coordination as a realistic scenario for lifelong learning
Hadi Nekoei, Akilesh Badrinaaraayanan, Aaron Courville, and Sarath Chandar · 2021
Later among the works it cites.
Collaborating with humans without human data
DJ Strouse, Kevin McKee, Matt Botvinick, Edward Hughes, and Richard Everett · 2021
Later among the works it cites.
A new formalism, method and open issues for zero-shot coordination
Johannes Treutlein, Michael Dennis, Caspar Oesterheld, and Jakob Foerster · 2021
Later among the works it cites.
Too many cooks: Bayesian inference for coordinating multi-agent collaboration
Sarah A Wu, Rose E Wang, James A Evans, Joshua B Tenenbaum, David C Parkes, and Max Kleiman-Weiner · 2021
Later among the works it cites.
Learning latent representations to influence multi-agent interaction
Annie Xie, Dylan Losey, Ryan Tolsma, Chelsea Finn, and Dorsa Sadigh · 2021
Later among the works it cites.
The surprising effectiveness of ppo in cooperative, multi-agent games
Chao Yu, Akash Velu, Eugene Vinitsky, Yu Wang, Alexandre Bayen, and Yi Wu · 2021
Later among the works it cites.
Maximum entropy population based training for zero-shot human-ai coordination
Rui Zhao, Jinming Song, Hu Haifeng, Yang Gao, Yi Wu, Zhongqian Sun, and Yang Wei · 2021
Later among the works it cites.
Continuously discovering novel strategies via reward-switching policy optimization
Zihan Zhou, Wei Fu, Bingliang Zhang, and Yi Wu · 2021
Later among the works it cites.
Do as i can, not as i say: Grounding language in robotic affordances
Michael Ahn, Anthony Brohan, Noah Brown, Yevgen Chebotar, Omar Cortes, Byron David, Chelsea Finn, Keerthana Gopalakrishnan, Karol Hausman, Alex Herzog, et al · 2022
Later among the works it cites.
The boltzmann policy distribution: Accounting for systematic suboptimality in human models
Cassidy Laidlaw and Anca Dragan · 2022
Later among the works it cites.
Co-gail: Learning diverse strategies for human-robot collaboration
Chen Wang, Claudia Pérez-D’Arpino, Danfei Xu, Li Fei-Fei, Karen Liu, and Silvio Savarese · 2022
Later among the works it cites.