Fetching the paper…
Reading the bibliography…
You have an environment, a model, and a reinforcement learning library that are designed to work together but don't.
The nethack learning environment
H. Küttler, N. Nardelli, A. H. Miller, R. Raileanu, M. Selvatici, E. Grefenstette, and T. Rocktäschel · 2006
Earlier work this paper cites.
Pettingzoo: Gym for multi-agent reinforcement learning
J. K. Terry, B. Black, A. Hari, L. S. Santos, C. Dieffendahl, N. L. Williams, Y. Lokesh, C. Horsch, and P. Ravi · 2009
Earlier work this paper cites.
Playing atari with deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. Graves, I. Antonoglou, D. Wierstra, and M. A. Riedmiller · 2013
Earlier work this paper cites.
G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba · 2016
Earlier work this paper cites.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Earlier work this paper cites.
IMPALA: scalable distributed deep-rl with importance weighted actor-learner architectures
L. Espeholt, H. Soyer, R. Munos, K. Simonyan, V. Mnih, T. Ward, Y. Doron, V. Firoiu, T. Harley, I. Dunning, S. Legg, and K. Kavukcuoglu · 2018
Cited alongside, same era.
dm_env: A python interface for reinforcement learning environments, 2019
A. Muldal, Y. Doron, J. Aslanides, T. Harley, T. Ward, and S. Liu · 2019
Cited alongside, same era.
Sample factory: Egocentric 3d control from pixels at 100000 FPS with asynchronous reinforcement learning
A. Petrenko, Z. Huang, T. Kumar, G. S. Sukhatme, and V. Koltun · 2020
Cited alongside, same era.
Cleanrl: High-quality single-file implementations of deep reinforcement learning algorithms
S. Huang, R. F. J. Dossa, C. Ye, and J. Braga · 2021
Cited alongside, same era.
Tianshou: A highly modularized deep reinforcement learning library
J. Weng, H. Chen, D. Yan, K. You, A. Duburcq, M. Zhang, Y. Su, H. Su, and J. Zhu
Stable-baselines3: Reliable reinforcement learning implementations
A. Raffin, A. Hill, A. Gleave, A. Kanervisto, M. Ernestus, and N. Dormann · 2021
Later among the works it cites.
The neural mmo platform for massively multiagent research
J. Suarez, Y. Du, C. Zhu, I. Mordatch, and P. Isola · 2021
Later among the works it cites.
Torchrl: A data-driven decision-making library for pytorch, 2023
A. Bou, M. Bettini, S. Dittert, V. Kumar, S. Sodhani, X. Yang, G. D. Fabritiis, and V. Moens · 2023
Later among the works it cites.
Gymnasium, Mar. 2023
M. Towers, J. K. Terry, A. Kwiatkowski, J. U. Balis, G. d. Cola, T. Deleu, M. Goulão, A. Kallinteris, A. KG, M. Krimmel, R. Perez-Vicente, A. Pierré, S. Schulhoff, J. J. Tai, A. T. J. Shen, and O. G. Younis · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited in the paper.
Envpool: A highly parallel reinforcement learning environment execution engine, 2022b
J. Weng, M. Lin, S. Huang, B. Liu, D. Makoviichuk, V. Makoviychuk, Z. Liu, Y. Song, T. Luo, Y. Jiang, Z. Xu, and S. Yan
Cited in the paper.