Fetching the paper…
Reading the bibliography…
AI agents designed to collaborate with people benefit from models that enable them to anticipate human behavior.
On the utility of model learning in HRI
Rohan Choudhury, Gokul Swamy, Dylan Hadfield-Menell, and Anca D. Dragan · 1901
Earlier work this paper cites.
Judgment under uncertainty: Heuristics and biases
Daniel Kahneman, Stewart Paul Slovic, Paul Slovic, and Amos Tversky · 1982
Earlier work this paper cites.
Cognitive biases and their impact on strategic planning
James H Barnes Jr · 1984
Earlier work this paper cites.
A framework for behavioural cloning
Michael Bain and Claude Sammut · 1999
Earlier work this paper cites.
The cross-entropy method for combinatorial and continuous optimization
Reuven Y. Rubinstein · 1999
Earlier work this paper cites.
Psychsim: Modeling theory of mind with decision-theoretic agents
David V Pynadath and Stacy C Marsella · 2005
Earlier work this paper cites.
Learning Personalized Models of Human Behavior in Chess
Reid McIlroy-Young, Russell Wang, Siddhartha Sen, Jon Kleinberg, and Ashton Anderson · 2008
Earlier work this paper cites.
Watch-And-Help: A Challenge for Social Perception and Human-AI Collaboration
Xavier Puig, Tianmin Shu, Shuang Li, Zilin Wang, Joshua B. Tenenbaum, Sanja Fidler, and Antonio Torralba · 2010
Earlier work this paper cites.
Ad hoc autonomous agent teams: Collaboration without pre-coordination
Peter Stone, Gal A. Kaminka, Sarit Kraus, and Jeffrey S. Rosenschein · 2010
Earlier work this paper cites.
Unsupervised feature learning and deep learning: A review and new perspectives
Yoshua Bengio, Aaron C. Courville, and Pascal Vincent · 2012
Earlier work this paper cites.
A game-theoretic model and best-response learning method for ad hoc coordination in multiagent systems
Stefano V. Albrecht and Subramanian Ramamoorthy · 2013
Earlier work this paper cites.
Human-robot cross-training: Computational formulation, modeling and evaluation of a human team training strategy
Stefanos Nikolaidis and Julie A. Shah · 2013
Earlier work this paper cites.
Cnn features off-the-shelf: An astounding baseline for recognition
Ali Sharif Razavian, Hossein Azizpour, Josephine Sullivan, and Stefan Carlsson · 2014
Earlier work this paper cites.
Early Stopping is Nonparametric Variational Inference
Dougal Maclaurin, David Duvenaud, and Ryan P. Adams · 2015
Earlier work this paper cites.
Overcooked, 2016
Ghost Town Games · 2016
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
David Silver, Aja Huang, Christopher J. Maddison, Arthur Guez, Laurent Sifre, George van den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, Sander Dieleman, Dominik Grewe, John Nham, Nal Kalchbrenner, Ilya Sutskever, Timothy Lillicrap, Madeleine Leach, Koray Kavukcuoglu, Thore Graepel, and Demis Hassabis · 2016
Cited alongside, same era.
Solving for Best Responses and Equilibria in Extensive-Form Games with Reinforcement Learning Methods
Amy Greenwald, Jiacui Li, and Eric Sodomka · 2017
Cited alongside, same era.
Proximal Policy Optimization Algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
Preparing for the unknown: Learning a universal policy with online system identification
Wenhao Yu, C. Karen Liu, and Greg Turk · 2017
Cited alongside, same era.
Residual networks for computer go
Simplified action decoder for deep multi-agent reinforcement learning
Hengyuan Hu and Jakob N. Foerster · 2020
Later among the works it cites.
"other-play" for zero-shot coordination
Hengyuan Hu, Adam Lerer, Alexander Peysakhovich, and Jakob N. Foerster · 2020
Later among the works it cites.
Adversarially guided self-play for adopting social conventions
Mycal Tucker, Yilun Zhou, and J. Shah · 2020
Later among the works it cites.
Too many cooks: Coordinating multi-agent collaboration through inverse planning
Rose E. Wang, Sarah A. Wu, James A. Evans, Joshua B. Tenenbaum, David C. Parkes, and Max Kleiman-Weiner · 2020
Later among the works it cites.
Sim-to-real transfer in deep reinforcement learning for robotics: a survey
Wenshuai Zhao, Jorge Peña Queralta, and Tomi Westerlund · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tristan Cazenave · 2018
Cited alongside, same era.
Recasting Gradient-Based Meta-Learning as Hierarchical Bayes
Erin Grant, Chelsea Finn, Sergey Levine, Trevor Darrell, and Thomas Griffiths · 2018
Cited alongside, same era.
Universal language model fine-tuning for text classification
Jeremy Howard and Sebastian Ruder · 2018
Cited alongside, same era.
Openai five
OpenAI · 2018
Cited alongside, same era.
Sim-to-real: Learning agile locomotion for quadruped robots
Jie Tan, Tingnan Zhang, Erwin Coumans, Atil Iscen, Yunfei Bai, Danijar Hafner, Steven Bohez, and Vincent Vanhoucke · 2018
Cited alongside, same era.
Behavioral cloning from observation
Faraz Torabi, Garrett Warnell, and Peter Stone · 2018
Cited alongside, same era.
On the utility of learning about humans for human-ai coordination
Micah Carroll, Rohin Shah, Mark K. Ho, Tom Griffiths, Sanjit A. Seshia, Pieter Abbeel, and Anca D. Dragan · 2019
Cited alongside, same era.
A generalized framework for self-play training
Daniel Hernandez, Kevin Denamganaï, Yuan Gao, Pete York, Sam Devlin, Spyridon Samothrakis, and James Alfred Walker · 2019
Cited alongside, same era.
Later among the works it cites.
Creating multimodal interactive agents with imitation and self-supervised learning
DeepMind Interactive Agents Team Josh Abramson, Arun Ahuja, Arthur Brussee, Federico Carnevale, Mary Cassin, Felix Fischer, Petko Georgiev, Alex Goldin, Tim Harley, Felix Hill, Peter C. Humphreys, Alden Hung, Jessica Landon, Timothy P. Lillicrap, Hamza Merzic, Alistair Muldal, Adam Santoro, Guy Scully, Tamara von Glehn, Greg Wayne, Nathaniel Wong, Chen Yan, and Rui Zhu · 2021
Later among the works it cites.
On the Importance of Environments in Human-Robot Coordination
Matthew C. Fontaine, Ya-Chuan Hsu, Yulun Zhang, Bryon Tjanaka, and Stefanos Nikolaidis · 2021
Later among the works it cites.
Evaluating the robustness of collaborative agents
Paul Knott, Micah Carroll, Sam Devlin, Kamil Ciosek, Katja Hofmann, Anca D. Dragan, and Rohin Shah · 2021
Later among the works it cites.
Interaction Flexibility in Artificial Agents Teaming with Humans
Patrick Nalepka, Jordan Gregory-Dunsmore, James Simpson, Gaurav Patil, and Michael Richardson · 2021
Later among the works it cites.
Collaborating with humans without human data
DJ Strouse, Kevin R. McKee, Matthew M. Botvinick, Edward Hughes, and Richard Everett · 2021
Later among the works it cites.
Human-ai coordination via human-regularized search and learning
Hengyuan Hu, David J. Wu, Adam Lerer, Jakob Nicolaus Foerster, and Noam Brown · 2022
Closest in time.
Modeling strong and human-like gameplay with kl-regularized search
Athul Paul Jacob, David J. Wu, Gabriele Farina, Adam Lerer, Anton Bakhtin, Jacob Andreas, and Noam Brown · 2022
Closest in time.
Assisting unknown teammates in unknown tasks: Ad hoc teamwork under partial observability
Joao G. Ribeiro, Cassandro Martinho, Alberto Sardinha, and Francisco S. Melo · 2022
Closest in time.
Causalagents: A robustness benchmark for motion forecasting using causal relationships
Rebecca Roelofs, Liting Sun, Benjamin Caine, Khaled S. Refaat, Benjamin Sapp, Scott M. Ettinger, and Wei Chai · 2022
Closest in time.