Fetching the paper…
Reading the bibliography…
The development of deep reinforcement learning (DRL) has benefited from the emergency of a variety type of game environments where new challenging problems are proposed and new algorithms can be tested safely and quickly, such as Board games, RTS, FPS, and MOBA games.
Obstacle tower: A generalization challenge in vision, control, and planning
Juliani, A.; Khalifa, A.; Berges, V.-P.; Harper, J.; Teng, E.; Henry, H.; Crespi, A.; Togelius, J.; and Lange, D. 2019 · 1902
Earlier work this paper cites.
Dealing with non-stationarity in multi-agent deep reinforcement learning
Papoudakis, G.; Christianos, F.; Rahman, A.; and Albrecht, S. V. 2019 · 1906
Earlier work this paper cites.
Google research football: A novel reinforcement learning environment
Kurach, K.; Raichuk, A.; Stańczyk, P.; Zając, M.; Bachem, O.; Espeholt, L.; Riquelme, C.; Vincent, D.; Michalski, M.; Bousquet, O.; et al. 2019 · 1907
Earlier work this paper cites.
Making Efficient Use of Demonstrations to Solve Hard Exploration Problems
Paine, T. L.; Gulcehre, C.; Shahriari, B.; Denil, M.; Hoffman, M.; Soyer, H.; Tanburn, R.; Kapturowski, S.; Rabinowitz, N.; Williams, D.; et al. 2019 · 1909
Earlier work this paper cites.
Multi-agent reinforcement learning: A selective overview of theories and algorithms
Zhang, K.; Yang, Z.; and Başar, T. 2019 · 1911
Earlier work this paper cites.
Leveraging procedural generation to benchmark reinforcement learning
Cobbe, K.; Hesse, C.; Hilton, J.; and Schulman, J. 2019 · 1912
Earlier work this paper cites.
The dynamics of reinforcement learning in cooperative multiagent systems
Claus, C.; and Boutilier, C. 1998 · 1998
Earlier work this paper cites.
Agent57: Outperforming the atari human benchmark
Badia, A. P.; Piot, B.; Kapturowski, S.; Sprechmann, P.; Vitvitskyi, A.; Guo, D.; and Blundell, C. 2020 · 2003
Earlier work this paper cites.
Thinking While Moving: Deep Reinforcement Learning with Concurrent Control
Xiao, T.; Jang, E.; Kalashnikov, D.; Levine, S.; Ibarz, J.; Hausman, K.; and Herzog, A. 2020 · 2004
Earlier work this paper cites.
Macro-Action-Based Deep Multi-Agent Reinforcement Learning
Xiao, Y.; Hoffman, J.; and Amato, C. 2020 · 2004
Earlier work this paper cites.
Hysteretic q-learning: an algorithm for decentralized reinforcement learning in cooperative multi-agent teams
Matignon, L.; Laurent, G. J.; and Le Fort-Piat, N. 2007 · 2007
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Bellemare, M. G.; Naddaf, Y.; Veness, J.; and Bowling, M. 2013 · 2013
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Mnih, V.; Kavukcuoglu, K.; Silver, D.; Graves, A.; Antonoglou, I.; Wierstra, D.; and Riedmiller, M. 2013 · 2013
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V.; Kavukcuoglu, K.; Silver, D.; Rusu, A. A.; Veness, J.; Bellemare, M. G.; Graves, A.; Riedmiller, M.; Fidjeland, A. K.; Ostrovski, G.; et al. 2015 · 2015
Earlier work this paper cites.
Beattie, C.; Leibo, J. Z.; Teplyashin, D.; Ward, T.; Wainwright, M.; Küttler, H.; Lefrancq, A.; Green, S.; Valdés, V.; Sadik, A.; et al. 2016 · 2016
Earlier work this paper cites.
Pybullet, a python module for physics simulation for games, robotics and machine learning
Coumans, E.; and Bai, Y. 2016 · 2016
Earlier work this paper cites.
Deep reinforcement learning from self-play in imperfect-information games
Heinrich, J.; and Silver, D. 2016 · 2016
Cited alongside, same era.
Safe, multi-agent, reinforcement learning for autonomous driving
Shalev-Shwartz, S.; Shammah, S.; and Shashua, A. 2016 · 2016
Cited alongside, same era.
Mastering the game of Go with deep neural networks and tree search
Silver, D.; Huang, A.; Maddison, C. J.; Guez, A.; Sifre, L.; Van Den Driessche, G.; Schrittwieser, J.; Antonoglou, I.; Panneershelvam, V.; Lanctot, M.; et al. 2016 · 2016
Cited alongside, same era.
Training agent for first-person shooter game with actor-critic curriculum learning
Wu, Y.; and Tian, Y. 2016 · 2016
Cited alongside, same era.
Counterfactual multi-agent policy gradients
Foerster, J.; Farquhar, G.; Afouras, T.; Nardelli, N.; and Whiteson, S. 2017 · 2017
Deep reinforcement learning for recommender systems
Munemasa, I.; Tomomatsu, Y.; Hayashi, K.; and Takagi, T. 2018 · 2018
Later among the works it cites.
Credit assignment for collective multiagent RL with global rewards
Nguyen, D. T.; Kumar, A.; and Lau, H. C. 2018 · 2018
Later among the works it cites.
Gotta learn fast: A new benchmark for generalization in rl
Nichol, A.; Pfau, V.; Hesse, C.; Klimov, O.; and Schulman, J. 2018 · 2018
Later among the works it cites.
OpenAI Five
OpenAI. 2018 · 2018
Later among the works it cites.
QMIX: Monotonic value function factorisation for deep multi-agent reinforcement learning
Rashid, T.; Samvelyan, M.; De Witt, C. S.; Farquhar, G.; Foerster, J.; and Whiteson, S. 2018 · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Ai2-thor: An interactive 3d environment for visual ai
Kolve, E.; Mottaghi, R.; Han, W.; VanderBilt, E.; Weihs, L.; Herrasti, A.; Gordon, D.; Zhu, Y.; Gupta, A.; and Farhadi, A. 2017 · 2017
Cited alongside, same era.
Playing FPS games with deep reinforcement learning
Lample, G.; and Chaplot, D. S. 2017 · 2017
Cited alongside, same era.
Virtual to real reinforcement learning for autonomous driving
Pan, X.; You, Y.; Wang, Z.; and Lu, C. 2017 · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
Schulman, J.; Wolski, F.; Dhariwal, P.; Radford, A.; and Klimov, O. 2017 · 2017
Cited alongside, same era.
Mastering the game of go without human knowledge
Silver, D.; Schrittwieser, J.; Simonyan, K.; Antonoglou, I.; Huang, A.; Guez, A.; Hubert, T.; Baker, L.; Lai, M.; Bolton, A.; et al. 2017 · 2017
Cited alongside, same era.
Starcraft ii: A new challenge for reinforcement learning
Vinyals, O.; Ewalds, T.; Bartunov, S.; Georgiev, P.; Vezhnevets, A. S.; Yeo, M.; Makhzani, A.; Küttler, H.; Agapiou, J.; Schrittwieser, J.; et al. 2017 · 2017
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T.; Zhou, A.; Abbeel, P.; and Levine, S. 2018 · 2018
Cited alongside, same era.
Sutton, R. S.; and Barto, A. G. 2018 · 2018
Later among the works it cites.
Tassa, Y.; Doron, Y.; Muldal, A.; Erez, T.; Li, Y.; Casas, D. d. L.; Budden, D.; Abdolmaleki, A.; Merel, J.; Lefrancq, A.; et al. 2018 · 2018
Later among the works it cites.
Master-Slave Curriculum Design for Reinforcement Learning
Wu, Y.; Zhang, W.; and Song, K. 2018 · 2018
Later among the works it cites.
Modeling and planning with macro-actions in decentralized POMDPs
Amato, C.; Konidaris, G.; Kaelbling, L. P.; and How, J. P. 2019 · 2019
Later among the works it cites.
A survey and critique of multiagent deep reinforcement learning
Hernandez-Leal, P.; Kartal, B.; and Taylor, M. E. 2019 · 2019
Later among the works it cites.
Explicitly Coordinated Policy Iteration
Hu, Y.; Chen, Y.; Fan, C.; and Hao, J. 2019 · 2019
Later among the works it cites.
Playing Card-Based RTS Games with Deep Reinforcement Learning
Liu, T.; Zheng, Z.; Li, H.; Bian, K.; and Song, L. 2019 · 2019
Later among the works it cites.
Habitat: A platform for embodied ai research
Savva, M.; Kadian, A.; Maksymets, O.; Zhao, Y.; Wijmans, E.; Jain, B.; Straub, J.; Liu, J.; Koltun, V.; Malik, J.; et al. 2019 · 2019
Later among the works it cites.
AlphaStar: Mastering the real-time strategy game StarCraft II
Vinyals, O.; Babuschkin, I.; Chung, J.; Mathieu, M.; Jaderberg, M.; Czarnecki, W. M.; Dudzik, A.; Huang, A.; Georgiev, P.; Powell, R.; et al. 2019 · 2019
Later among the works it cites.
Value-Decomposition Networks For Cooperative Multi-Agent Learning Based On Team Reward
Sunehag, P.; Lever, G.; Gruslys, A.; Czarnecki, W. M.; Zambaldi, V. F.; Jaderberg, M.; Lanctot, M.; Sonnerat, N.; Leibo, J. Z.; Tuyls, K.; et al. 2018 · 2087
Closest in time.