Fetching the paper…
Reading the bibliography…
In multi-agent reinforcement learning, the problem of learning to act is particularly difficult because the policies of co-players may be heavily conditioned on information only observed by them.
Optimal control of markov processes with incomplete state information
Karl J Astrom · 1965
Earlier work this paper cites.
Does the chimpanzee have a theory of mind?
David Premack and Guy Woodruff · 1978
Earlier work this paper cites.
Markov games as a framework for multi-agent reinforcement learning
Michael L. Littman · 1994
Earlier work this paper cites.
On players’ models of other players: Theory and experimental evidence
Dale O Stahl and Paul W Wilson · 1995
Earlier work this paper cites.
The dynamics of reinforcement learning in cooperative multiagent systems
Caroline Claus and Craig Boutilier · 1998
Earlier work this paper cites.
Planning and acting in partially observable stochastic domains
Leslie Pack Kaelbling, Michael L. Littman, and Anthony R. Cassandra · 1998
Earlier work this paper cites.
Playing is believing: The role of beliefs in multi-agent learning
Yu-Han Chang and Leslie Pack Kaelbling · 2001
Earlier work this paper cites.
Coordination in multiagent reinforcement learning: A bayesian approach
Georgios Chalkiadakis and Craig Boutilier · 2003
Earlier work this paper cites.
Evolutionary game dynamics
Josef Hofbauer and Karl Sigmund · 2003
Earlier work this paper cites.
Dynamic programming for partially observable stochastic games
Eric A. Hansen, Daniel S. Bernstein, and Shlomo Zilberstein · 2004
Earlier work this paper cites.
A framework for sequential planning in multi-agent settings
Piotr J Gmytrasiewicz and Prashant Doshi · 2005
Earlier work this paper cites.
On the difficulty of achieving equilibrium in interactive pomdps
Prashant Doshi and Piotr Gmytrasiewicz · 2006
Earlier work this paper cites.
Multi-agent reinforcement learning algorithm to handle beliefs of other agents’ policies and embedded beliefs
Takaki Makino and Kazuyuki Aihara · 2006
Earlier work this paper cites.
Hierarchical dirichlet processes
Yee Whye Teh, Michael I Jordan, Matthew J Beal, and David M Blei · 2006
Earlier work this paper cites.
A comprehensive survey of multiagent reinforcement learning
L. Busoniu, R. Babuska, and B. De Schutter · 2008
Earlier work this paper cites.
Monte carlo sampling methods for approximating interactive pomdps
Prashant Doshi and Piotr J Gmytrasiewicz · 2009
Earlier work this paper cites.
A bayesian approach for learning and planning in partially observable markov decision processes
Stéphane Ross, Joelle Pineau, Brahim Chaib-draa, and Pierre Kreitmann · 2011
Earlier work this paper cites.
Theory of mind: Mechanisms, methods, and new directions
Lindsey Byom and Bilge Mutlu · 2013
Earlier work this paper cites.
Learning phrase representations using RNN encoder-decoder for statistical machine translation
Kyunghyun Cho, Bart van Merrienboer, Çaglar Gülçehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio · 2014
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P. Kingma and Max Welling · 2014
Cited alongside, same era.
From Classical To Epistemic Game Theory
Andrés Perea · 2014
Cited alongside, same era.
Stochastic backpropagation and approximate inference in deep generative models
Danilo Jimenez Rezende, Shakir Mohamed, and Daan Wierstra · 2014
Cited alongside, same era.
MADE: masked autoencoder for distribution estimation
Mathieu Germain, Karol Gregor, Iain Murray, and Hugo Larochelle · 2015
Cited alongside, same era.
On the identifiability of mixture models from grouped samples
Robert A. Vandermeulen and Clayton D. Scott · 2015
Cited alongside, same era.
Towards a neural statistician
Harrison Edwards and Amos J. Storkey · 2017
Cited alongside, same era.
Intrinsic social motivation via causal influence in multi-agent RL
Natasha Jaques, Angeliki Lazaridou, Edward Hughes, Çaglar Gülçehre, Pedro A. Ortega, DJ Strouse, Joel Z. Leibo, and Nando de Freitas · 2018
Later among the works it cites.
Neural belief states for partially observed domains
Pol Moreno, Jan Humplik, George Papamakarios, Bernardo Avila Pires, Lars Buesing, Nicolas Heess, and Theophane Weber · 2018
Later among the works it cites.
Neil C Rabinowitz, Frank Perbet, H Francis Song, Chiyuan Zhang, SM Eslami, and Matthew Botvinick · 2018
Later among the works it cites.
Generalized elbo with constrained optimization , geco
Danilo Jimenez Rezende and Fabio Viola · 2018
Later among the works it cites.
Theory of mind: The state of the art
Henry M. Wellman · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
beta-vae: Learning basic visual concepts with a constrained variational framework
Irina Higgins, Loïc Matthey, Arka Pal, Christopher Burgess, Xavier Glorot, Matthew Botvinick, Shakir Mohamed, and Alexander Lerchner · 2017
Cited alongside, same era.
Maintaining cooperation in complex social dilemmas using deep reinforcement learning
Adam Lerer and Alexander Peysakhovich · 2017
Cited alongside, same era.
Predictive-state decoders: Encoding the future into recurrent networks
Arun Venkatraman, Nicholas Rhinehart, Wen Sun, Lerrel Pinto, Martial Hebert, Byron Boots, Kris Kitani, and J Bagnell · 2017
Cited alongside, same era.
Learning hierarchical features from deep generative models
Shengjia Zhao, Jiaming Song, and Stefano Ermon · 2017
Cited alongside, same era.
Autonomous agents modelling other agents: A comprehensive survey and open problems
Stefano V. Albrecht and Peter Stone · 2018
Cited alongside, same era.
Depth-limited solving for imperfect-information games
Noam Brown, Tuomas Sandholm, and Brandon Amos · 2018
Cited alongside, same era.
Nolan Bard, Jakob N. Foerster, Sarath Chandar, Neil Burch, Marc Lanctot, H. Francis Song, Emilio Parisotto, Vincent Dumoulin, Subhodeep Moitra, Edward Hughes, Iain Dunning, Shibl Mourad, Hugo Larochelle, Marc G. Bellemare, and Michael Bowling · 2019
Later among the works it cites.
Autoregressive energy machines
Conor Durkan and Charlie Nash · 2019
Later among the works it cites.
Bayesian action decoder for deep multi-agent reinforcement learning
Jakob N. Foerster, H. Francis Song, Edward Hughes, Neil Burch, Iain Dunning, Shimon Whiteson, Matthew Botvinick, and Michael Bowling · 2019
Later among the works it cites.
Temporal difference variational auto-encoder
Karol Gregor, George Papamakarios, Frederic Besse, Lars Buesing, and Theophane Weber · 2019
Later among the works it cites.
Shaping belief states with generative environment models for rl
Karol Gregor, Danilo Jimenez Rezende, Frederic Besse, Yan Wu, Hamza Merzic, and Aaron van den Oord · 2019
Later among the works it cites.
Ipomdp-net: A deep neural network for partially observable multi-agent planning using interactive pomdps
Yanlin Han and Piotr Gmytrasiewicz · 2019
Later among the works it cites.
Meta reinforcement learning as task inference
Jan Humplik, Alexandre Galashov, Leonard Hasenclever, Pedro A Ortega, Yee Whye Teh, and Nicolas Heess · 2019
Later among the works it cites.
Generalized variational inference
Jeremias Knoblauch, Jack Jewson, and Theodoros Damoulas · 2019
Later among the works it cites.
Understanding posterior collapse in generative latent variable models
James Lucas, George Tucker, Roger B. Grosse, and Mohammad Norouzi · 2019
Later among the works it cites.
Probabilistic recursive reasoning for multi-agent reinforcement learning
Ying Wen, Yaodong Yang, Rui Luo, Jun Wang, and Wei Pan · 2019
Later among the works it cites.
Learning causal state representations of partially observable environments
Amy Zhang, Zachary C Lipton, Luis Pineda, Kamyar Azizzadenesheli, Anima Anandkumar, Laurent Itti, Joelle Pineau, and Tommaso Furlanello · 2019
Later among the works it cites.
Simplified action decoder for deep multi-agent reinforcement learning
Hengyuan Hu and Jakob N. Foerster · 2020
Later among the works it cites.
Options as responses: Grounding behavioural hierarchies in multi-agent reinforcement learning
Alexander Sasha Vezhnevets, Yuhuai Wu, Remi Leblond, and Joel Z Leibo · 2020
Later among the works it cites.