Fetching the paper…
Reading the bibliography…
In this paper, we propose ELF, an Extensive, Lightweight and Flexible platform for fundamental reinforcement learning research.
Robocup simulation league: https://en.wikipedia.org/wiki/robocup_simulation_league
RoboCup Simulation League · 1995
Earlier work this paper cites.
Warzone 2100: https://wz2100.net/
Pumpkin Studios · 1999
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David A McAllester, Satinder P Singh, Yishay Mansour, et al · 1999
Earlier work this paper cites.
Battlecode, mit’s ai programming competition: https://www.battlecode.org/
BattleCode · 2000
Earlier work this paper cites.
On the development of a free rts game engine
Michael Buro and Timothy Furtak · 2005
Earlier work this paper cites.
Openra: http://www.openra.net/
OpenRA · 2007
Earlier work this paper cites.
Parallel monte-carlo tree search
Guillaume MJ-B Chaslot, Mark HM Winands, and H Jaap van Den Herik · 2008
Earlier work this paper cites.
Spring: https://springrts.com/
Stefan Johansson and Robin Westberg · 2008
Earlier work this paper cites.
Bullet physics engine
Erwin Coumans · 2010
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Marc G. Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling · 2012
Earlier work this paper cites.
A survey of monte carlo tree search methods
Cameron B Browne, Edward Powley, Daniel Whitehouse, Simon M Lucas, Peter I Cowling, Philipp Rohlfshagen, Stephen Tavener, Diego Perez, Spyridon Samothrakis, and Simon Colton · 2012
Earlier work this paper cites.
The combinatorial multi-armed bandit problem and its application to real-time strategy games
Santiago Ontanón · 2013
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Cited alongside, same era.
Massively parallel methods for deep reinforcement learning
Arun Nair, Praveen Srinivasan, Sam Blackwell, Cagdas Alcicek, Rory Fearon, Alessandro De Maria, Vedavyas Panneershelvam, Mustafa Suleyman, Charles Beattie, Stig Petersen, Shane Legg, Volodymyr Mnih, Koray Kavukcuoglu, and David Silver · 2015
Cited alongside, same era.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael I Jordan, and Philipp Moritz · 2015
Cited alongside, same era.
Vizdoom: A doom-based ai research platform for visual reinforcement learning
Michał Kempka, Marek Wydmuch, Grzegorz Runc, Jakub Toczek, and Wojciech Jaśkowski · 2016
Later among the works it cites.
Playing fps games with deep reinforcement learning
Guillaume Lample and Devendra Singh Chaplot · 2016
Later among the works it cites.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy P Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Later among the works it cites.
Mastering the game of go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al · 2016
Later among the works it cites.
Torchcraft: a library for machine learning research on real-time strategy games
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mazebase: A sandbox for learning from games
Sainbayar Sukhbaatar, Arthur Szlam, Gabriel Synnaeve, Soumith Chintala, and Rob Fergus · 2015
Cited alongside, same era.
Better computer go player with neural network and long-term prediction
Yuandong Tian and Yan Zhu · 2015
Cited alongside, same era.
Charles Beattie, Joel Z. Leibo, Denis Teplyashin, Tom Ward, Marcus Wainwright, Heinrich Küttler, Andrew Lefrancq, Simon Green, Víctor Valdés, Amir Sadik, Julian Schrittwieser, Keith Anderson, Sarah York, Max Cant, Adam Cain, Adrian Bolton, Stephen Gaffney, Helen King, Demis Hassabis, Shane Legg, and Stig Petersen · 2016
Cited alongside, same era.
Playing SNES in the retro learning environment
Nadav Bhonker, Shai Rozenberg, and Itay Hubara · 2016
Cited alongside, same era.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Cited alongside, same era.
The malmo platform for artificial intelligence experimentation
Matthew Johnson, Katja Hofmann, Tim Hutton, and David Bignell · 2016
Cited alongside, same era.
Rectifier nonlinearities improve neural network acoustic models
Andrew L Maas, Awni Y Hannun, and Andrew Y Ng
Cited in the paper.
Gabriel Synnaeve, Nantas Nardelli, Alex Auvolat, Soumith Chintala, Timothée Lacroix, Zeming Lin, Florian Richoux, and Nicolas Usunier · 2016
Later among the works it cites.
Reinforcement learning through asynchronous advantage actor-critic on a gpu
Mohammad Babaeizadeh, Iuri Frosio, Stephen Tyree, Jason Clemons, and Jan Kautz · 2017
Closest in time.
Learning to navigate in complex environments
Piotr Mirowski, Razvan Pascanu, Fabio Viola, Hubert Soyer, Andrew J. Ballard, Andrea Banino, Misha Denil, Ross Goroshin, Laurent Sifre, Koray Kavukcuoglu, Dharshan Kumaran, and Raia Hadsell · 2017
Closest in time.
Multiagent bidirectionally-coordinated nets for learning to play starcraft combat games
Peng Peng, Quan Yuan, Ying Wen, Yaodong Yang, Zhenkun Tang, Haitao Long, and Jun Wang · 2017
Closest in time.
Episodic exploration for deep deterministic policies: An application to starcraft micromanagement tasks
Nicolas Usunier, Gabriel Synnaeve, Zeming Lin, and Soumith Chintala · 2017
Closest in time.
Training agent for first-person shooter game with actor-critic curriculum learning
Yuxin Wu and Yuandong Tian · 2017
Closest in time.