Fetching the paper…
Reading the bibliography…
Models of human behavior for prediction and collaboration tend to fall into two categories: ones that learn from large amounts of data via imitation learning, and ones that assume human behavior to be noisily-optimal for some reward function.
On the Utility of Learning about Humans for Human-AI Coordination
Micah Carroll, Rohin Shah, Mark K. Ho, Thomas L. Griffiths, Sanjit A. Seshia, Pieter Abbeel, and Anca Dragan · 1910
Earlier work this paper cites.
Explorations in behavioral consistency: Properties of persons, situations, and behaviors
David C. Funder and C. Randall Colvin · 1939
Earlier work this paper cites.
Situational similarity and personality predict behavioral consistency
Ryne A. Sherman, Christopher S. Nave, and David C. Funder · 1939
Earlier work this paper cites.
Individual choice behavior
R. Duncan Luce · 1959
Earlier work this paper cites.
The Choice Axiom After Twenty Years
R. Duncan Luce · 1977
Earlier work this paper cites.
Long Short-Term Memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Algorithms for Inverse Reinforcement Learning
Andrew Y. Ng and Stuart J. Russell · 2000
Earlier work this paper cites.
Using advice to transfer knowledge acquired in one reinforcement learning task to another
Lisa Torrey, Trevor Walker, Jude Shavlik, and Richard Maclin · 2005
Earlier work this paper cites.
Planning-based prediction for pedestrians
Brian D. Ziebart, Nathan Ratliff, Garratt Gallagher, Christoph Mertz, Kevin Peterson, J. Andrew Bagnell, Martial Hebert, Anind K. Dey, and Siddhartha Srinivasa · 2009
Earlier work this paper cites.
Ad Hoc Autonomous Agent Teams: Collaboration without Pre-Coordination
Peter Stone, Gal Kaminka, Sarit Kraus, and Jeffrey Rosenschein · 2010
Earlier work this paper cites.
Human arm motion modeling and long-term prediction for safe and efficient Human-Robot-Interaction
Hao Ding, G. Reissig, Kurniawan Wijaya, D. Bortot, K. Bengler, and O. Stursberg · 2011
Earlier work this paper cites.
Intention-Aware Motion Planning
Tirthankar Bandyopadhyay, Kok Sung Won, Emilio Frazzoli, David Hsu, Wee Sun Lee, and Daniela Rus · 2013
Earlier work this paper cites.
A policy-blending formalism for shared control
Anca D Dragan and Siddhartha S Srinivasa · 2013
Earlier work this paper cites.
Anticipating human activities for reactive robotic response
H. Koppula and Ashutosh Saxena · 2013
Earlier work this paper cites.
Human-robot collaborative manipulation planning using early prediction of human motion
Jim Mainprice and D. Berenson · 2013
Earlier work this paper cites.
Generative Adversarial Networks
Ian J. Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization
Diederik P. Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Auto-Encoding Variational Bayes
Diederik P. Kingma and Max Welling · 2014
Cited alongside, same era.
Policy Transfer using Reward Shaping
Tim Brys, Anna Harutyunyan, Matthew E. Taylor, and Ann Nowé · 2015
Cited alongside, same era.
NICE: Non-linear Independent Components Estimation
Laurent Dinh, David Krueger, and Yoshua Bengio · 2015
Cited alongside, same era.
Shared Autonomy via Hindsight Optimization
Shervin Javdani, Siddhartha S. Srinivasa, and J. Andrew Bagnell · 2015
Cited alongside, same era.
Learning driving styles for autonomous vehicles from demonstration
M. Kuderer, S. Gulati, and W. Burgard · 2015
Cited alongside, same era.
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Later among the works it cites.
RLlib: Abstractions for Distributed Reinforcement Learning
Eric Liang, Richard Liaw, Philipp Moritz, Robert Nishihara, Roy Fox, Ken Goldberg, Joseph E. Gonzalez, Michael I. Jordan, and Ion Stoica · 2018
Later among the works it cites.
Multimodal Probabilistic Model-Based Planning for Human-Robot Interaction
E. Schmerling, Karen Leung, Wolf Vollprecht, and M. Pavone · 2018
Later among the works it cites.
Probabilistic Prediction of Interactive Driving Behavior via Hierarchical Inverse Reinforcement Learning
Liting Sun, W. Zhan, and M. Tomizuka · 2018
Later among the works it cites.
Modeling interaction via the principle of maximum causal entropy
Brian D. Ziebart, J. Andrew Bagnell, and Anind K. Dey · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Social LSTM: Human Trajectory Prediction in Crowded Spaces
Alexandre Alahi, Kratarth Goel, Vignesh Ramanathan, Alexandre Robicquet, Li Fei-Fei, and S. Savarese · 2016
Cited alongside, same era.
Neural Machine Translation by Jointly Learning to Align and Translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2016
Cited alongside, same era.
Generative Adversarial Imitation Learning
Jonathan Ho and Stefano Ermon · 2016
Cited alongside, same era.
Socially compliant mobile robot navigation via inverse reinforcement learning
Henrik Kretzschmar, Markus Spies, C. Sprunk, and W. Burgard · 2016
Cited alongside, same era.
Unsupervised Representation Learning with Deep Convolutional Generative Adversarial Networks
Alec Radford, Luke Metz, and Soumith Chintala · 2016
Cited alongside, same era.
Variational Inference: A Review for Statisticians
David M. Blei, Alp Kucukelbir, and Jon D. McAuliffe · 2017
Cited alongside, same era.
Deep reinforcement learning from human preferences
Paul Christiano, Jan Leike, Tom B. Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Cited alongside, same era.
Later among the works it cites.
MultiPath: Multiple Probabilistic Anchor Trajectory Hypotheses for Behavior Prediction
Yuning Chai, Benjamin Sapp, M. Bansal, and Dragomir Anguelov · 2019
Later among the works it cites.
Optimal rates of entropy estimation over Lipschitz balls
Yanjun Han, Jiantao Jiao, Tsachy Weissman, and Yihong Wu · 2019
Later among the works it cites.
PyTorch: An Imperative Style, High-Performance Deep Learning Library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala · 2019
Later among the works it cites.
Imitation Learning for Human Pose Prediction
Borui Wang, Ehsan Adeli, Hsu-kuang Chiu, De-An Huang, and Juan Carlos Niebles · 2019
Later among the works it cites.
“Other-Play” for Zero-Shot Coordination
Hengyuan Hu, Adam Lerer, Alex Peysakhovich, and Jakob Foerster · 2020
Later among the works it cites.
Scalable Multi-Task Imitation Learning with Autonomous Improvement
Avi Singh, Eric Jang, Alexander Irpan, Daniel Kappler, Murtaza Dalal, Sergey Levinev, Mohi Khansari, and Chelsea Finn · 2020
Later among the works it cites.
Off-Belief Learning
Hengyuan Hu, Adam Lerer, Brandon Cui, Luis Pineda, Noam Brown, and Jakob Foerster · 2021
Later among the works it cites.
Evaluating the Robustness of Collaborative Agents
Paul Knott, Micah Carroll, Sam Devlin, Kamil Ciosek, Katja Hofmann, A. D. Dragan, and Rohin Shah · 2021
Later among the works it cites.
Collaborating with Humans without Human Data
DJ Strouse, Kevin McKee, Matt Botvinick, Edward Hughes, and Richard Everett · 2021
Later among the works it cites.
On complementing end-to-end human behavior predictors with planning
Liting Sun, Xiaogang Jia, and A. Dragan · 2021
Later among the works it cites.
A New Formalism, Method and Open Issues for Zero-Shot Coordination
Johannes Treutlein, Michael Dennis, Caspar Oesterheld, and Jakob Foerster · 2021
Later among the works it cites.