Fetching the paper…
Reading the bibliography…
We consider the multi-agent reinforcement learning setting with imperfect information in which each agent is trying to maximize its own utility.
Stochastic games
Shapley, L. S · 1953
Earlier work this paper cites.
Does the chimpanzee have a theory of mind?
Premack, David and Woodruff, Guy · 1978
Earlier work this paper cites.
Folk psychology as simulation
Gordon, Robert M · 1986
Earlier work this paper cites.
Why the child’s theory of mind really is a theory
Gopnik, Alison and Wellman, Henry M · 1992
Earlier work this paper cites.
Long short-term memory
Hochreiter, Sepp and Schmidhuber, Jürgen · 1997
Earlier work this paper cites.
Mirror neurons and the simulation theory of mind-reading
Gallese, Vittorio and Goldman, Alvin · 1998
Earlier work this paper cites.
Learning agents for uncertain environments
Russell, Stuart · 1998
Earlier work this paper cites.
Using artificial neural networks to model opponents in texas hold’em
Davidson, Aaron · 1999
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
Ng, Andrew Y, Russell, Stuart J, et al · 2000
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
Abbeel, Pieter and Ng, Andrew Y · 2004
Earlier work this paper cites.
Agent-based modeling as a bridge between disciplines
Axelrod, Robert · 2006
Earlier work this paper cites.
Evolving explicit opponent models in game playing
Lockett, Alan J, Chen, Charles L, and Miikkulainen, Risto · 2007
Earlier work this paper cites.
Legibility and predictability of robot motion
Dragan, Anca D, Lee, Kenton CT, and Srinivasa, Siddhartha S · 2013
Cited alongside, same era.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Saxe, Andrew M, McClelland, James L, and Ganguli, Surya · 2013
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, Diederik and Ba, Jimmy · 2014
Cited alongside, same era.
Fast and accurate deep network learning by exponential linear units (elus)
Clevert, Djork-Arné, Unterthiner, Thomas, and Hochreiter, Sepp · 2015
Cited alongside, same era.
Mazebase: A sandbox for learning from games
Sukhbaatar, Sainbayar, Szlam, Arthur, Synnaeve, Gabriel, Chintala, Soumith, and Fergus, Rob · 2015
Learning multiagent communication with backpropagation
Sukhbaatar, Sainbayar, Fergus, Rob, et al · 2016
Later among the works it cites.
It takes two to tango: Towards theory of ai’s mind
Chandrasekaran, Arjun, Yadav, Deshraj, Chattopadhyay, Prithvijit, Prabhu, Viraj, and Parikh, Devi · 2017
Later among the works it cites.
Learning cooperative visual dialog agents with deep reinforcement learning
Das, Abhishek, Kottur, Satwik, Moura, José MF, Lee, Stefan, and Batra, Dhruv · 2017
Later among the works it cites.
Pragmatic-pedagogic value alignment
Fisac, Jaime F, Gates, Monica A, Hamrick, Jessica B, Liu, Chang, Hadfield-Menell, Dylan, Palaniappan, Malayandi, Malik, Dhruv, Sastry, S Shankar, Griffiths, Thomas L, and Dragan, Anca D · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Cooperative inverse reinforcement learning
Hadfield-Menell, Dylan, Russell, Stuart J, Abbeel, Pieter, and Dragan, Anca · 2016
Cited alongside, same era.
Opponent modeling in deep reinforcement learning
He, He, Boyd-Graber, Jordan, Kwok, Kevin, and Daumé III, Hal · 2016
Cited alongside, same era.
Categorical reparameterization with gumbel-softmax
Jang, Eric, Gu, Shixiang, and Poole, Ben · 2016
Cited alongside, same era.
Coordinate to cooperate or compete: abstract goals and joint intentions in social interaction
Kleiman-Weiner, Max, Ho, Mark K, Austerweil, Joseph L, Littman, Michael L, and Tenenbaum, Joshua B · 2016
Cited alongside, same era.
Multi-agent cooperation and the emergence of (natural) language
Lazaridou, Angeliki, Peysakhovich, Alexander, and Baroni, Marco · 2016
Cited alongside, same era.
The concrete distribution: A continuous relaxation of discrete random variables
Maddison, Chris J, Mnih, Andriy, and Teh, Yee Whye · 2016
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
Mnih, Volodymyr, Badia, Adria Puigdomenech, Mirza, Mehdi, Graves, Alex, Lillicrap, Timothy, Harley, Tim, Silver, David, and Kavukcuoglu, Koray · 2016
Cited alongside, same era.
Foerster, Jakob N, Chen, Richard Y, Al-Shedivat, Maruan, Whiteson, Shimon, Abbeel, Pieter, and Mordatch, Igor · 2017
Later among the works it cites.
Inverse reward design
Hadfield-Menell, Dylan, Milli, Smitha, Abbeel, Pieter, Russell, Stuart J, and Dragan, Anca · 2017
Later among the works it cites.
Multi-agent reinforcement learning in sequential social dilemmas
Leibo, Joel Z, Zambaldi, Vinicius, Lanctot, Marc, Marecki, Janusz, and Graepel, Thore · 2017
Later among the works it cites.
Maintaining cooperation in complex social dilemmas using deep reinforcement learning
Lerer, Adam and Peysakhovich, Alexander · 2017
Later among the works it cites.
Multi-agent actor-critic for mixed cooperative-competitive environments
Lowe, Ryan, Wu, Yi, Tamar, Aviv, Harb, Jean, Abbeel, Pieter, and Mordatch, Igor · 2017
Later among the works it cites.
Emergence of grounded compositional language in multi-agent populations
Mordatch, Igor and Abbeel, Pieter · 2017
Later among the works it cites.
Deep decentralized multi-task multi-agent rl under partial observability
Omidshafiei, Shayegan, Pazis, Jason, Amato, Christopher, How, Jonathan P, and Vian, John · 2017
Later among the works it cites.