Fetching the paper…
Reading the bibliography…
Advantage Actor-critic (A2C) and Proximal Policy Optimization (PPO) are popular deep reinforcement learning algorithms used for game AI in recent years.
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu, “Asynchronous methods for deep reinforcement learning,” in International conference on machine learning . PMLR, 2016, pp. 1928–1937
1937
Earlier work this paper cites.
2013
Earlier work this paper cites.
2016
Earlier work this paper cites.
2017
Earlier work this paper cites.
E. Liang, R. Liaw, R. Nishihara, P. Moritz, R. Fox, K. Goldberg, J. Gonzalez, M. Jordan, and I. Stoica, “Rllib: Abstractions for distributed reinforcement learning,” in International Conference on Machine Learning . PMLR, 2018, pp. 3053–3062
2018
Earlier work this paper cites.
R. S. Sutton and A. G. Barto, Reinforcement learning: An introduction . MIT press, 2018
2018
Earlier work this paper cites.
N. Justesen, L. M. Uth, C. Jakobsen, P. D. Moore, J. Togelius, and S. Risi, “Blood bowl: A new board game challenge and competition for ai,” in 2019 IEEE Conference on Games (CoG) , 2019, pp. 1–8
2019
Cited alongside, same era.
L. Engstrom, A. Ilyas, S. Santurkar, D. Tsipras, F. Janoos, L. Rudolph, and A. Madry, “Implementation matters in deep rl: A case study on ppo and trpo,” in International Conference on Learning Representations , 2019
2019
Cited alongside, same era.
mglss, “Why is the log probability replaced with the importance sampling in the loss function?” Artificial Intelligence Stack Exchange, 2019. [Online]. Available: https://ai.stackexchange.com/a/13216/31987
2019
Cited alongside, same era.
K. Kurach, A. Raichuk, P. Stanczyk, M. Zajac, O. Bachem, L. Espeholt, C. Riquelme, D. Vincent, M. Michalski, O. Bousquet, and S. Gelly, “Google research football: A novel reinforcement learning environment,” in AAAI , 2020
2020
Cited alongside, same era.
C. D’Eramo, D. Tateo, A. Bonarini, M. Restelli, and J. Peters, “Mushroomrl: Simplifying reinforcement learning research,” Journal of Machine Learning Research , 2020
2020
Later among the works it cites.
A. Raffin, A. Hill, A. Gleave, A. Kanervisto, M. Ernestus, and N. Dormann, “Stable-baselines3: Reliable reinforcement learning implementations,” Journal of Machine Learning Research , vol. 22, no. 268, pp. 1–8, 2021. [Online]. Available: http://jmlr.org/papers/v22/20-1364.html
2021
Later among the works it cites.
Y. Fujita, P. Nagarajan, T. Kataoka, and T. Ishikawa, “Chainerrl: A deep reinforcement learning library,” Journal of Machine Learning Research , vol. 22, no. 77, pp. 1–14, 2021
2021
Later among the works it cites.
M. Andrychowicz, A. Raichuk, P. Stańczyk, M. Orsini, S. Girgin, R. Marinier, L. Hussenot, M. Geist, O. Pietquin, M. Michalski, S. Gelly, and O. Bachem, “What matters for on-policy deep actor-critic methods? a large-scale study,” in International Conference on Learning Representations , 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. T. Kristensen and P. Burelli, “Strategies for using proximal policy optimization in mobile puzzle games,” International Conference on the Foundations of Digital Games , 2020
2020
Cited alongside, same era.
2021
Later among the works it cites.
S. Huang, R. F. J. Dossa, A. Raffin, A. Kanervisto, and W. Wang, “The 37 implementation details of proximal policy optimization,” in ICLR Blog Track , 2022, https://iclr-blog-track.github.io/2022/03/25/ppo-implementation-details/. [Online]. Available: https://iclr-blog-track.github.io/2022/03/25/ppo-implementation-details/
2022
Closest in time.