Fetching the paper…
Reading the bibliography…
On April 13th, 2019, OpenAI Five became the first AI system to defeat the world champions at an esports game.
“Open-ended Learning in Symmetric Zero-sum Games”
David Balduzzi, Marta Garnelo, Yoram Bachrach, Wojciech. Czarnecki, Julien Pérolat, Max Jaderberg and Thore Graepel · 1901
Earlier work this paper cites.
“The MineRL Competition on Sample Efficient Reinforcement Learning using Human Priors”
William. Guss, Cayden Codel, Katja Hofmann, Brandon Houghton, Noburu Kuno, Stephanie Milani, Sharada Mohanty, Diego Liebana, Ruslan Salakhutdinov, Nichonolay Topin, Manuela Veloso and Phillip Wang · 1904
Earlier work this paper cites.
“Solving Rubik’s Cube with a Robot Hand”, 2019
OpenAI, Ilge Akkaya, Marcin Andrychowicz, Maciek Chociej, Mateusz Litwin, Bob McGrew, Arthur Petron, Alex Paino, Matthias Plappert, Glenn Powell, Raphael Ribas, Jonas Schneider, Nikolas Tezak, Jerry Tworek, Peter Welinder, Lilian Weng, Qiming Yuan, Wojciech Zaremba and Lei Zhang · 1910
Earlier work this paper cites.
“Asynchronous methods for deep reinforcement learning”
Volodymyr Mnih, Adria Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver and Koray Kavukcuoglu · 1937
Earlier work this paper cites.
“Asynchronous Methods for Deep Reinforcement Learning”
Volodymyr Mnih, Adria Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver and Koray Kavukcuoglu · 1937
Earlier work this paper cites.
“Iterative Solution of Games by Fictitious Play”
George. Brown · 1951
Earlier work this paper cites.
“An Efficient Gradient-Based Algorithm for On-Line Training of Recurrent Network Trajectories”
Ronald. Williams and Jing Peng · 1990
Earlier work this paper cites.
“Advances in Neural Information Processing Systems 2”
Scott. Fahlman and Christian Lebiere · 1990
Earlier work this paper cites.
“Function optimization using connectionist reinforcement learning algorithms”
Ronald Williams and Jing Peng · 1991
Earlier work this paper cites.
“A world championship caliber checkers program”
Jonathan Schaeffer, Joseph Culberson, Norman Treloar, Brent Knight, Paul Lu and Duane Szafron · 1992
Earlier work this paper cites.
“Q-learning”
Christopher Watkins and Peter Dayan · 1992
Earlier work this paper cites.
“TD-Gammon, a self-teaching backgammon program, achieves master-level play”
Gerald Tesauro · 1994
Earlier work this paper cites.
“Searching for solutions in games and artificial intelligence”, 1994
L. Allis · 1994
Earlier work this paper cites.
“Learning to forget: Continual prediction with LSTM”
Felix Gers, Jürgen Schmidhuber and Fred Cummins · 1999
Earlier work this paper cites.
“Policy invariance under reward transformations: Theory and application to reward shaping”
Andrew. Ng, Daishi Harada and Stuart Russell · 1999
Earlier work this paper cites.
“Actor-critic algorithms”
Vijay Konda and John Tsitsiklis · 2000
Earlier work this paper cites.
“Deep Blue”
Murray Campbell, A. Hoane Jr. and Feng-hsiung Hsu · 2002
Earlier work this paper cites.
“Adversarial Classification”
Nilesh Dalvi, Pedro Domingos, Mausam, Sumit Sanghai and Deepak Verma · 2004
Earlier work this paper cites.
“TrueSkill: a Bayesian skill rating system”
Ralf Herbrich, Tom Minka and Thore Graepel · 2007
Earlier work this paper cites.
“A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning”
Stephane Ross, Geoffrey Gordon and Drew Bagnell · 2011
Earlier work this paper cites.
“Playing atari with deep reinforcement learning”
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra and Martin Riedmiller · 2013
Earlier work this paper cites.
“Guided Policy Search”
Sergey Levine and Vladlen Koltun · 2013
Cited alongside, same era.
“Adam: A method for stochastic optimization”
Diederik Kingma and Jimmy Ba · 2014
Cited alongside, same era.
“A comparative study of visual and auditory reaction times on the basis of gender and physical activity levels of medical first year students”
Aditya Jain, Ramta Bansal, Avnish Kumar and KD Singh · 2015
Cited alongside, same era.
“Distilling the Knowledge in a Neural Network”
Geoffrey Hinton, Oriol Vinyals and Jeffrey Dean · 2015
Cited alongside, same era.
“Mastering the game of Go with deep neural networks and tree search”
David Silver, Aja Huang, Chris Maddison, Arthur Guez, Laurent Sifre, George Van, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam and Marc Lanctot · 2016
Cited alongside, same era.
“Scaling SGD Batch Size to 32K for ImageNet Training”, 2017
Yang You, Igor Gitman and Boris Ginsburg · 2017
Later among the works it cites.
“Growing a Brain: Fine-Tuning by Increasing Model Capacity”
Yu-Xiong Wang, Deva Ramanan and Martial Hebert · 2017
Later among the works it cites.
“Learning without Forgetting”
Z. Li and D. Hoiem · 2017
Later among the works it cites.
“Learning Dexterity” [Online; accessed 28-May-2019], https://openai.com/blog/learning-dexterity/ , 2018
OpenAI · 2018
Later among the works it cites.
“A general reinforcement learning algorithm that masters chess, shogi, and Go through self-play”
David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran and Thore Graepel · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“High-Dimensional Continuous Control Using Generalized Advantage Estimation”
John Schulman, Philipp Moritz, Sergey Levine, Michael. Jordan and Pieter Abbeel · 2016
Cited alongside, same era.
“Net2Net: Accelerating Learning via Knowledge Transfer”
Tianqi Chen, Ian. Goodfellow and Jonathon Shlens · 2016
Cited alongside, same era.
“Deep Reinforcement Learning from Self-Play in Imperfect-Information Games”
Johannes Heinrich and David Silver · 2016
Cited alongside, same era.
“Hierarchical deep reinforcement learning: Integrating temporal abstraction and intrinsic motivation”
Tejas Kulkarni, Karthik Narasimhan, Ardavan Saeedi and Josh Tenenbaum · 2016
Cited alongside, same era.
Andrei. Rusu, Neil. Rabinowitz, Guillaume Desjardins, Hubert Soyer, James Kirkpatrick, Koray Kavukcuoglu, Razvan Pascanu and Raia Hadsell · 2016
Cited alongside, same era.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang and Wojciech Zaremba · 2016
Cited alongside, same era.
“A Deep Reinforced Model for Abstractive Summarization”, 2017
Romain Paulus, Caiming Xiong and Richard Socher · 2017
Cited alongside, same era.
Dan Horgan, John Quan, David Budden, Gabriel Barth-Maron, Matteo Hessel, Hado van Hasselt and David Silver · 2018
Later among the works it cites.
“An empirical model of large-batch training”
Sam McCandlish, Jared Kaplan, Dario Amodei and OpenAI Team · 2018
Later among the works it cites.
“Quantifying Generalization in Reinforcement Learning”
Karl Cobbe, Oleg Klimov, Christopher Hesse, Taehoon Kim and John Schulman · 2018
Later among the works it cites.
Max Jaderberg, Wojciech Czarnecki, Iain Dunning, Luke Marris, Guy Lever, Antonio Castaneda, Charles Beattie, Neil Rabinowitz, Ari Morcos and Avraham Ruderman · 2018
Later among the works it cites.
“Exploration by random network distillation”
Yuri Burda, Harrison Edwards, Amos Storkey and Oleg Klimov · 2018
Later among the works it cites.
“Montezuma’s revenge solved by go-explore, a new algorithm for hard-exploration problems (sets records on pitfall too)”
Adrien Ecoffet, Joost Huizinga, Joel Lehman, Kenneth Stanley and Jeff Clune · 2018
Later among the works it cites.
“ImageNet Training in Minutes”
Yang You, Zhao Zhang, Cho-Jui Hsieh, James Demmel and Kurt Keutzer · 2018
Later among the works it cites.
“Impala: Scalable distributed deep-rl with importance weighted actor-learner architectures”
Lasse Espeholt, Hubert Soyer, Remi Munos, Karen Simonyan, Volodymir Mnih, Tom Ward, Yotam Doron, Vlad Firoiu, Tim Harley and Iain Dunning · 2018
Later among the works it cites.
“Mix&Match - Agent Curricula for Reinforcement Learning”
Wojciech Czarnecki, Siddhant. Jayakumar, Max Jaderberg, Leonard Hasenclever, Yee Teh, Simon Osindero, Nicolas Heess and Razvan Pascanu · 2018
Later among the works it cites.
“AI and Compute” [Online; accessed 9-Sept-2019], https://openai.com/blog/ai-and-compute/ , 2018
OpenAI · 2018
Later among the works it cites.
“Grandmaster level in StarCraft II using multi-agent reinforcement learning”
Oriol Vinyals, Igor Babuschkin, Wojciech Czarnecki, Michaël Mathieu, Andrew Dudzik, Junyoung Chung, David Choi, Richard Powell, Timo Ewalds and Petko Georgiev · 2019
Closest in time.
“Dota 2 — Wikipedia, The Free Encyclopedia” [Online; accessed 9-September-2019], https://en.wikipedia.org/w/index.php?title=Dota_2&oldid=913733447 , 2019
Wikipedia contributors · 2019
Closest in time.
“The International 2018 — Wikipedia, The Free Encyclopedia” [Online; accessed 9-September-2019], https://en.wikipedia.org/w/index.php?title=The_International_2018&oldid=912865272 , 2019
Wikipedia contributors · 2019
Closest in time.
“NVIDIA Collective Communications Library (NCCL)” [Online; accessed 9-September-2019], https://developer.nvidia.com/nccl , 2019
NVIDIA · 2019
Closest in time.
“Superhuman AI for multiplayer poker”
Noam Brown and Tuomas Sandholm · 2019
Closest in time.