Fetching the paper…
Reading the bibliography…
We study the problem of training a Reinforcement Learning (RL) agent that is collaborative with humans without using any human data.
Augmenting human intellect: A conceptual framework
D. C. Engelbart · 1962
Earlier work this paper cites.
Td-gammon, a self-teaching backgammon program, achieves master-level play
G. Tesauro · 1994
Earlier work this paper cites.
Planning, learning and coordination in multiagent decision processes
C. Boutilier · 1996
Earlier work this paper cites.
A framework for behavioural cloning
M. Bain and C. Sammut · 1999
Earlier work this paper cites.
Mathematical analysis
S. Ko · 2006
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
B. D. Ziebart, A. L. Maas, J. A. Bagnell, A. K. Dey, et al · 2008
Earlier work this paper cites.
A survey on transfer learning
S. J. Pan and Q. Yang · 2009
Earlier work this paper cites.
Robot trajectory optimization using approximate inference
M. Toussaint · 2009
Earlier work this paper cites.
Modeling purposeful adaptive behavior with the principle of maximum causal entropy
B. D. Ziebart · 2010
Earlier work this paper cites.
Machine learning: A probabilistic perspective. adaptive computation and machine learning, 2012
K. P. Murphy · 2012
Earlier work this paper cites.
On stochastic optimal control and reinforcement learning by approximate inference
K. Rawlik, M. Toussaint, and S. Vijayakumar · 2013
Earlier work this paper cites.
Taming the noise in reinforcement learning via soft updates
R. Fox, A. Pakman, and N. Tishby · 2015
Earlier work this paper cites.
Learning to communicate with deep multi-agent reinforcement learning
J. Foerster, I. A. Assael, N. De Freitas, and S. Whiteson · 2016
Earlier work this paper cites.
Overcooked, 2016
Ghost Town Games · 2016
Earlier work this paper cites.
Coordinate to cooperate or compete: abstract goals and joint intentions in social interaction
M. Kleiman-Weiner, M. K. Ho, J. L. Austerweil, M. L. Littman, and J. B. Tenenbaum · 2016
Earlier work this paper cites.
Gradient descent only converges to minimizers
J. D. Lee, M. Simchowitz, M. I. Jordan, and B. Recht · 2016
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu · 2016
Earlier work this paper cites.
Prioritized experience replay
T. Schaul, J. Quan, I. Antonoglou, and D. Silver · 2016
Earlier work this paper cites.
Using artificial intelligence to augment human intelligence
S. Carter and M. Nielsen · 2017
Earlier work this paper cites.
Reinforcement learning with deep energy-based policies
T. Haarnoja, H. Tang, P. Abbeel, and S. Levine · 2017
Earlier work this paper cites.
Population based training of neural networks
M. Jaderberg, V. Dalibard, S. Osindero, W. M. Czarnecki, J. Donahue, A. Razavi, O. Vinyals, T. Green, I. Dunning, K. Simonyan, et al · 2017
Cited alongside, same era.
Maintaining cooperation in complex social dilemmas using deep reinforcement learning
A. Lerer and A. Peysakhovich · 2017
Cited alongside, same era.
Multi-agent actor-critic for mixed cooperative-competitive environments
R. Lowe, Y. Wu, A. Tamar, J. Harb, P. Abbeel, and I. Mordatch · 2017
Cited alongside, same era.
Equivalence between policy gradients and soft q-learning
J. Schulman, X. Chen, and P. Abbeel · 2017
Cited alongside, same era.
Learning to walk via deep reinforcement learning
T. Haarnoja, S. Ha, A. Zhou, J. Tan, G. Tucker, and S. Levine · 2019
Later among the works it cites.
Human-level performance in 3d multiplayer games with population-based reinforcement learning
M. Jaderberg, W. M. Czarnecki, I. Dunning, L. Marris, G. Lever, A. G. Castañeda, C. Beattie, N. C. Rabinowitz, A. S. Morcos, A. Ruderman, N. Sonnerat, T. Green, L. Deason, J. Z. Leibo, D. Silver, D. Hassabis, K. Kavukcuoglu, and T. Graepel · 2019
Later among the works it cites.
Diversity-inducing policy gradient: Using maximum mean discrepancy to find a set of diverse policies
M. A. Masood and F. Doshi-Velez · 2019
Later among the works it cites.
OpenAI Five finals
OpenAI · 2019
Later among the works it cites.
Theory of minds: Understanding behavior in groups through inverse planning
M. Shum, M. Kleiman-Weiner, M. L. Littman, and J. B. Tenenbaum · 2019
Later among the works it cites.
AlphaStar: Mastering the Real-Time Strategy Game StarCraft II
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Cited alongside, same era.
Mastering chess and shogi by self-play with a general reinforcement learning algorithm
D. Silver, T. Hubert, J. Schrittwieser, I. Antonoglou, M. Lai, A. Guez, M. Lanctot, L. Sifre, D. Kumaran, T. Graepel, et al · 2017
Cited alongside, same era.
Domain randomization for transferring deep neural networks from simulation to the real world
J. Tobin, R. Fong, A. Ray, J. Schneider, W. Zaremba, and P. Abbeel · 2017
Cited alongside, same era.
Preparing for the unknown: Learning a universal policy with online system identification
W. Yu, J. Tan, C. K. Liu, and G. Turk · 2017
Cited alongside, same era.
Learning dexterous in-hand manipulation
M. Andrychowicz, B. Baker, M. Chociej, R. Jozefowicz, B. McGrew, J. Pachocki, A. Petron, M. Plappert, G. Powell, A. Ray, et al · 2018
Cited alongside, same era.
Counterfactual multi-agent policy gradients
J. Foerster, G. Farquhar, T. Afouras, N. Nardelli, and S. Whiteson · 2018
Cited alongside, same era.
Latent space policies for hierarchical reinforcement learning
T. Haarnoja, K. Hartikainen, P. Abbeel, and S. Levine · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine · 2018
Cited alongside, same era.
O. Vinyals, I. Babuschkin, J. Chung, M. Mathieu, M. Jaderberg, W. M. Czarnecki, A. Dudzik, A. Huang, P. Georgiev, R. Powell, T. Ewalds, D. Horgan, M. Kroiss, I. Danihelka, J. Agapiou, J. Oh, V. Dalibard, D. Choi, L. Sifre, Y. Sulsky, S. Vezhnevets, J. Molloy, T. Cai, D. Budden, T. Paine, C. Gulcehre, Z. Wang, T. Pfaff, T. Pohlen, Y. Wu, D. Yogatama, J. Cohen, K. McKinney, O. Smith, T. Schaul, T. Lillicrap, C. Apps, K. Kavukcuoglu, D. Hassabis, and D. Silver · 2019
Later among the works it cites.
Grandmaster level in starcraft ii using multi-agent reinforcement learning
O. Vinyals, I. Babuschkin, W. M. Czarnecki, M. Mathieu, A. Dudzik, J. Chung, D. H. Choi, R. Powell, T. Ewalds, P. Georgiev, et al · 2019
Later among the works it cites.
Maximum entropy-regularized multi-goal reinforcement learning
R. Zhao, X. Sun, and V. Tresp · 2019
Later among the works it cites.
L. Han, J. Xiong, P. Sun, X. Sun, M. Fang, Q. Guo, Q. Chen, T. Shi, H. Yu, and Z. Zhang · 2020
Later among the works it cites.
“other-play” for zero-shot coordination
H. Hu, A. Lerer, A. Peysakhovich, and J. Foerster · 2020
Later among the works it cites.
Effective diversity in population based reinforcement learning
J. Parker-Holder, A. Pacchiano, K. M. Choromanski, and S. J. Roberts · 2020
Later among the works it cites.
Discovering diverse multi-agent strategic behavior via reward randomization
Z. Tang, C. Yu, B. Chen, H. Xu, X. Wang, F. Fang, S. S. Du, Y. Wang, and Y. Wu · 2020
Later among the works it cites.
Adversarially guided self-play for adopting social conventions
M. Tucker, Y. Zhou, and J. Shah · 2020
Later among the works it cites.
Evaluating the robustness of collaborative agents
P. Knott, M. Carroll, S. Devlin, K. Ciosek, K. Hofmann, A. Dragan, and R. Shah · 2021
Closest in time.
Unifying behavioral and response diversity for open-ended learning in zero-sum games
X. Liu, H. Jia, Y. Wen, Y. Yang, Y. Hu, Y. Chen, C. Fan, and Z. Hu · 2021
Closest in time.
Trajectory diversity for zero-shot coordination
A. Lupu, B. Cui, H. Hu, and J. Foerster · 2021
Closest in time.
Modelling behavioural diversity for learning in open-ended games
N. Perez-Nieves, Y. Yang, O. Slumbers, D. H. Mguni, Y. Wen, and J. Wang · 2021
Closest in time.
Collaborating with humans without human data
D. Strouse, K. McKee, M. Botvinick, E. Hughes, and R. Everett · 2021
Closest in time.
Mutual information state intrinsic control
R. Zhao, Y. Gao, P. Abbeel, V. Tresp, and W. Xu · 2021
Closest in time.