Fetching the paper…
Reading the bibliography…
Reinforcement learning (RL) has been successful in training agents in various learning environments, including video-games.
S. Thrun and A. Schwartz, “Issues in using function approximation for reinforcement learning,” in Proceedings of the 1993 Connectionist Models Summer School
1993
Earlier work this paper cites.
A. Y. Ng, D. Harada, and S. Russell, “Policy invariance under reward transformations: Theory and application to reward shaping,” in ICML
1999
Earlier work this paper cites.
2004
Earlier work this paper cites.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” arXiv:1412.6980
2014
Earlier work this paper cites.
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
M. Kempka, M. Wydmuch, G. Runc, J. Toczek, and W. Jaśkowski, “Vizdoom: A doom-based ai research platform for visual reinforcement learning,” in CIG
2016
Earlier work this paper cites.
M. Johnson, K. Hofmann, T. Hutton, and D. Bignell, “The malmo platform for artificial intelligence experimentation.,” in IJCAI
2016
Earlier work this paper cites.
S. Gu, T. Lillicrap, I. Sutskever, and S. Levine, “Continuous deep q-learning with model-based acceleration,” in ICML
2016
Earlier work this paper cites.
G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba, “Openai gym,” 2016
2016
Earlier work this paper cites.
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu, “Asynchronous methods for deep reinforcement learning,” in ICML
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
Y. Duan, X. Chen, R. Houthooft, J. Schulman, and P. Abbeel, “Benchmarking deep reinforcement learning for continuous control,” in ICML
2016
Earlier work this paper cites.
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
P.-W. Chou, D. Maturana, and S. Scherer, “Improving stochastic policy gradients in continuous control with deep reinforcement learning using the beta distribution,” in ICML
2017
Cited alongside, same era.
A. Dosovitskiy and V. Koltun, “Learning to act by predicting the future,” in ICLR
2017
Cited alongside, same era.
Y. Wu and Y. Tian, “Training agent for first-person shooter game with actor-critic curriculum learning,” in ICLR
2017
Cited alongside, same era.
2017
Cited alongside, same era.
MIT press, 2018
R. S. Sutton and A. G. Barto, Reinforcement learning: An introduction · 2018
Cited alongside, same era.
E. Liang, R. Liaw, R. Nishihara, P. Moritz, R. Fox, K. Goldberg, J. E. Gonzalez, M. I. Jordan, and I. Stoica, “RLlib: Abstractions for distributed reinforcement learning,” in ICML
2018
Later among the works it cites.
A. Raffin, “Rl baselines zoo.” https://github.com/araffin/rl-baselines-zoo , 2018
2018
Later among the works it cites.
O. Vinyals, I. Babuschkin, W. M. Czarnecki, M. Mathieu, A. Dudzik, J. Chung, D. H. Choi, R. Powell, T. Ewalds, P. Georgiev, et al
2019
Later among the works it cites.
OpenAI et al., “Dota 2 with large scale deep reinforcement learning,” arXiv:1807.01281
2019
Later among the works it cites.
2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2018
Cited alongside, same era.
M. Wydmuch, M. Kempka, and W. Jaśkowski, “Vizdoom competitions: Playing doom from pixels,” IEEE Transactions on Games
2018
Cited alongside, same era.
J. Harmer, L. Gisslén, J. del Val, H. Holst, J. Bergdahl, T. Olsson, K. Sjöö, and M. Nordin, “Imitation learning with concurrent actions in 3d games,” in CIG
2018
Cited alongside, same era.
Y. Gao, H. Xu, J. Lin, F. Yu, S. Levine, and T. Darrell, “Reinforcement learning from imperfect demonstrations,” in ICML
2018
Cited alongside, same era.
T. Zahavy, M. Haroush, N. Merlis, D. J. Mankowitz, and S. Mannor, “Learn what not to learn: Action elimination with deep reinforcement learning,” in NIPS
2018
Cited alongside, same era.
A. Tavakoli, F. Pardo, and P. Kormushev, “Action branching architectures for deep reinforcement learning,” in AAAI
2018
Cited alongside, same era.
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine, “Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor,” in ICML
2018
Cited alongside, same era.
Later among the works it cites.
O. Delalleau, M. Peter, E. Alonso, and A. Logut, “Discrete and continuous action representation for practical rl in video games,” in AAAI Workshop on Reinforcement Learning in Games
2019
Later among the works it cites.
A. Kanervisto and V. Hautamäki, “Torille: Learning environment for hand-to-hand combat,” in CoG
2019
Later among the works it cites.
2019
Later among the works it cites.
A. Nichol, “Competing in the obstacle tower challenge.” https://blog.aqnichol.com/2019/07/24/competing-in-the-obstacle-tower-challenge , 2019
2019
Later among the works it cites.
V. Zambaldi, D. Raposo, A. Santoro, V. Bapst, Y. Li, I. Babuschkin, K. Tuyls, D. Reichert, T. Lillicrap, E. Lockhart, et al
2019
Later among the works it cites.
O. Vinyals, I. Babuschkin, J. Chung, M. Mathieu, M. Jaderberg, W. Czarnecki, A. Dudzik, A. Huang, P. Georgiev, R. Powell, et al
2019
Later among the works it cites.
D. Ye, Z. Liu, M. Sun, B. Shi, P. Zhao, H. Wu, H. Yu, S. Yang, X. Wu, Q. Guo, et al
2020
Closest in time.
2020
Closest in time.