Fetching the paper…
Reading the bibliography…
This paper presents an algorithmic framework for learning robust policies in asymmetric imperfect-information games, where the joint reward could depend on the uncertain opponent type (a private information known only to the opponent itself and its ally).
Model-based learning of interaction strategies in multi-agent systems
David Carmel and Shaul Markovitch. 1998 · 1998
Earlier work this paper cites.
Planning and acting in partially observable stochastic domains
Leslie Pack Kaelbling, Michael L Littman, and Anthony R Cassandra. 1998 · 1998
Earlier work this paper cites.
Case-based plan recognition in computer games. In International Conference on Case-Based Reasoning . Springer, 161–170
Michael Fagan and Pádraig Cunningham. 2003 · 2003
Earlier work this paper cites.
Learning to play Bayesian games
Eddie Dekel, Drew Fudenberg, and David K Levine. 2004 · 2004
Earlier work this paper cites.
Dynamic programming for partially observable stochastic games. In AAAI , Vol. 4. 709–715
Eric A Hansen, Daniel S Bernstein, and Shlomo Zilberstein. 2004 · 2004
Earlier work this paper cites.
Monte-Carlo planning in large POMDPs. In Advances in neural information processing systems . 2164–2172
David Silver and Joel Veness. 2010 · 2010
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller. 2013 · 2013
Earlier work this paper cites.
DESPOT: Online POMDP planning with regularization. In Advances in neural information processing systems . 1772–1780
Adhiraj Somani, Nan Ye, David Hsu, and Wee Sun Lee. 2013 · 2013
Earlier work this paper cites.
Multiagent learning in the presence of memory-bounded agents
Doran Chakraborty and Peter Stone. 2014 · 2014
Earlier work this paper cites.
Deep recurrent Q-learning for partially observable mdps. In 2015 AAAI Fall Symposium Series
Matthew Hausknecht and Peter Stone. 2015 · 2015
Earlier work this paper cites.
Memory-Bounded Dynamic Programming for DEC-POMDPs.. In IJCAI . 2009–2015
Sven Seuken and Shlomo Zilberstein. 2007 · 2015
Cited alongside, same era.
Deep reinforcement learning from self-play in imperfect-information games
Johannes Heinrich and David Silver. 2016 · 2016
Cited alongside, same era.
A concise introduction to decentralized POMDPs . Vol. 1
Frans A Oliehoek, Christopher Amato, et al · 2016
Cited alongside, same era.
Plan Recognition as Planning Revisited.. In IJCAI . 3258–3264
Shirin Sohrabi, Anton V Riabov, and Octavian Udrea. 2016 · 2016
Cited alongside, same era.
AI Wolf Contest—Development of Game AI Using Collective Intelligence—
Fujio Toriumi, Hirotaka Osawa, Michimasa Inaba, Daisuke Katagami, Kosuke Shinoda, and Hitoshi Matsubara. 2016 · 2016
Cited alongside, same era.
Distral: Robust multitask reinforcement learning. In Advances in Neural Information Processing Systems . 4496–4506
Yee Teh, Victor Bapst, Wojciech M Czarnecki, John Quan, James Kirkpatrick, Raia Hadsell, Nicolas Heess, and Razvan Pascanu. 2017 · 2017
Later among the works it cites.
Ensemble adversarial training: Attacks and defenses
Florian Tramèr, Alexey Kurakin, Nicolas Papernot, Ian Goodfellow, Dan Boneh, and Patrick McDaniel. 2017 · 2017
Later among the works it cites.
Autonomous agents modelling other agents: A comprehensive survey and open problems
Stefano V Albrecht and Peter Stone. 2018 · 2018
Later among the works it cites.
Learning dexterous in-hand manipulation
Marcin Andrychowicz, Bowen Baker, Maciek Chociej, Rafal Jozefowicz, Bob McGrew, Jakub Pachocki, Arthur Petron, Matthias Plappert, Glenn Powell, Alex Ray, et al · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Trapit Bansal, Jakub Pachocki, Szymon Sidor, Ilya Sutskever, and Igor Mordatch. 2017 · 2017
Cited alongside, same era.
Population based training of neural networks
Max Jaderberg, Valentin Dalibard, Simon Osindero, Wojciech M Czarnecki, Jeff Donahue, Ali Razavi, Oriol Vinyals, Tim Green, Iain Dunning, Karen Simonyan, et al · 2017
Cited alongside, same era.
Multi-agent actor-critic for mixed cooperative-competitive environments. In Advances in Neural Information Processing Systems . 6379–6390
Ryan Lowe, Yi Wu, Aviv Tamar, Jean Harb, OpenAI Pieter Abbeel, and Igor Mordatch. 2017 · 2017
Cited alongside, same era.
Deepstack: Expert-level artificial intelligence in heads-up no-limit poker
Matej Moravčík, Martin Schmid, Neil Burch, Viliam Lisỳ, Dustin Morrill, Nolan Bard, Trevor Davis, Kevin Waugh, Michael Johanson, and Michael Bowling. 2017 · 2017
Cited alongside, same era.
Mastering the game of go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, et al · 2017
Cited alongside, same era.
Sarsop: Efficient point-based pomdp planning by approximating optimally reachable belief spaces
Hanna Kurniawati, David Hsu, and Wee Sun Lee. [n. d.]
Cited in the paper.
Collision avoidance for unmanned aircraft using Markov decision processes. In AIAA guidance, navigation, and control conference . 8040
Selim Temizer, Mykel Kochenderfer, Leslie Kaelbling, Tomas Lozano-Pérez, and James Kuchar. [n. d.]
Cited in the paper.
Max Jaderberg, Wojciech M Czarnecki, Iain Dunning, Luke Marris, Guy Lever, Antonio Garcia Castaneda, Charles Beattie, Neil C Rabinowitz, Ari S Morcos, Avraham Ruderman, et al · 2018
Later among the works it cites.
OpenAI Five
OpenAI. 2018 · 2018
Later among the works it cites.
Modeling others using oneself in multi-agent reinforcement learning
Roberta Raileanu, Emily Denton, Arthur Szlam, and Rob Fergus. 2018 · 2018
Later among the works it cites.
Collaborative evolutionary reinforcement learning
Shauharda Khadka, Somdeb Majumdar, Santiago Miret, Evren Tumer, Tarek Nassar, Zach Dwiel, Yinyin Liu, and Kagan Tumer. 2019 · 2019
Closest in time.
AlphaStar: Mastering the real-time strategy game StarCraft II
Oriol Vinyals, Igor Babuschkin, Junyoung Chung, Michael Mathieu, Max Jaderberg, Wojciech M Czarnecki, Andrew Dudzik, Aja Huang, Petko Georgiev, Richard Powell, et al · 2019
Closest in time.