Fetching the paper…
Reading the bibliography…
There has been significant progress in developing reinforcement learning (RL) training systems.
Mujoco: A physics engine for model-based control
E. Todorov, T. Erez, and Y. Tassa · 2012
Earlier work this paper cites.
The Arcade Learning Environment: An evaluation platform for general agents
M. G. Bellemare, Y. Naddaf, J. Veness, and M. Bowling · 2013
Earlier work this paper cites.
A simple c++11 thread pool implementation
J. Progsch · 2014
Earlier work this paper cites.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, S. Petersen, C. Beattie, A. Sadik, I. Antonoglou, H. King, D. Kumaran, D. Wierstra, S. Legg, and D. Hassabis · 2015
Earlier work this paper cites.
G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba · 2016
Earlier work this paper cites.
ViZDoom: A Doom-based AI research platform for visual reinforcement learning
M. Kempka, M. Wydmuch, G. Runc, J. Toczek, and W. Jaśkowski · 2016
Earlier work this paper cites.
Mastering the game of Go with deep neural networks and tree search
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, et al · 2016
Earlier work this paper cites.
pybind11 – seamless operability between C++11 and Python
W. Jakob, J. Rhinelander, and D. Moldovan · 2017
Earlier work this paper cites.
In-Datacenter performance analysis of a Tensor Processing Unit
N. P. Jouppi, C. Young, N. Patil, D. A. Patterson, G. Agrawal, R. Bajwa, S. Bates, S. Bhatia, N. Boden, A. Borchers, R. Boyle, et al · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Earlier work this paper cites.
Mastering Chess and Shogi by self-play with a general reinforcement learning algorithm
D. Silver, T. Hubert, J. Schrittwieser, I. Antonoglou, M. Lai, A. Guez, M. Lanctot, L. Sifre, D. Kumaran, T. Graepel, et al · 2017
Earlier work this paper cites.
Mastering the game of Go without human knowledge
D. Silver, J. Schrittwieser, K. Simonyan, I. Antonoglou, A. Huang, A. Guez, T. Hubert, L. Baker, M. Lai, A. Bolton, et al · 2017
Earlier work this paper cites.
JAX: composable transformations of Python+NumPy programs, 2018
J. Bradbury, R. Frostig, P. Hawkins, M. J. Johnson, C. Leary, D. Maclaurin, G. Necula, A. Paszke, J. VanderPlas, S. Wanderman-Milne, and Q. Zhang · 2018
Cited alongside, same era.
Minimalistic gridworld environment for openai gym
M. Chevalier-Boisvert, L. Willems, and S. Pal · 2018
Cited alongside, same era.
IMPALA: Scalable distributed deep-rl with importance weighted actor-learner architectures
L. Espeholt, H. Soyer, R. Munos, K. Simonyan, V. Mnih, T. Ward, Y. Doron, V. Firoiu, T. Harley, I. Dunning, S. Legg, and K. Kavukcuoglu · 2018
Cited alongside, same era.
RLlib: Abstractions for distributed reinforcement learning
E. Liang, R. Liaw, R. Nishihara, P. Moritz, R. Fox, K. Goldberg, J. Gonzalez, M. Jordan, and I. Stoica · 2018
Cited alongside, same era.
Ray: A distributed framework for emerging AI applications
P. Moritz, R. Nishihara, S. Wang, A. Tumanov, R. Liaw, E. Liang, M. Elibol, Z. Yang, W. Paul, M. I. Jordan, et al · 2018
Cited alongside, same era.
Mastering Atari, Go, Chess and Shogi by planning with a learned model
J. Schrittwieser, I. Antonoglou, T. Hubert, K. Simonyan, L. Sifre, S. Schmitt, A. Guez, E. Lockhart, D. Hassabis, T. Graepel, et al · 2020
Later among the works it cites.
dm_control: Software and tasks for continuous control
S. Tunyasuvunakool, A. Muldal, Y. Doron, S. Liu, S. Bohez, J. Merel, T. Erez, T. P. Lillicrap, N. Heess, and Y. Tassa · 2020
Later among the works it cites.
Extending jax with custom c++ and cuda code
D. Foreman-Mackey · 2021
Later among the works it cites.
Brax–a differentiable physics engine for large scale rigid body simulation
C. D. Freeman, E. Frey, A. Raichuk, S. Girgin, I. Mordatch, and O. Bachem · 2021
Later among the works it cites.
Podracer architectures for scalable reinforcement learning
M. Hessel, M. Kroiss, A. Clark, I. Kemaev, J. Quan, T. Keck, F. Viola, and H. van Hasselt · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Reinforcement learning: An introduction
R. S. Sutton and A. G. Barto · 2018
Cited alongside, same era.
Dota 2 with large scale deep reinforcement learning
C. Berner, G. Brockman, B. Chan, V. Cheung, P. Dębiak, C. Dennison, D. Farhi, Q. Fischer, S. Hashme, C. Hesse, et al · 2019
Cited alongside, same era.
Grandmaster level in StarCraft II using multi-agent reinforcement learning
O. Vinyals, I. Babuschkin, W. M. Czarnecki, M. Mathieu, A. Dudzik, J. Chung, D. H. Choi, R. Powell, T. Ewalds, P. Georgiev, et al · 2019
Cited alongside, same era.
Seed RL: Scalable and efficient deep-rl with accelerated central inference
L. Espeholt, R. Marinier, P. Stanczyk, K. Wang, and M. Michalski · 2020
Cited alongside, same era.
Array programming with NumPy
C. R. Harris, K. J. Millman, S. J. van der Walt, R. Gommers, P. Virtanen, D. Cournapeau, E. Wieser, J. Taylor, S. Berg, N. J. Smith, R. Kern, M. Picus, S. Hoyer, M. H. van Kerkwijk, M. Brett, A. Haldane, J. F. del Río, M. Wiebe, P. Peterson, P. Gérard-Marchant, K. Sheppard, T. Reddy, W. Weckesser, H. Abbasi, C. Gohlke, and T. E. Oliphant · 2020
Cited alongside, same era.
Acme: A research framework for distributed reinforcement learning
M. Hoffman, B. Shahriari, J. Aslanides, G. Barth-Maron, F. Behbahani, T. Norman, A. Abdolmaleki, A. Cassirer, F. Yang, K. Baumli, S. Henderson, A. Novikov, S. G. Colmenarejo, S. Cabi, C. Gulcehre, T. L. Paine, A. Cowie, Z. Wang, B. Piot, and N. de Freitas · 2020
Cited alongside, same era.
Sample factory: Egocentric 3D control from pixels at 100000 FPS with asynchronous reinforcement learning
A. Petrenko, Z. Huang, T. Kumar, G. Sukhatme, and V. Koltun · 2020
Cited alongside, same era.
S. Huang, R. F. J. Dossa, C. Ye, and J. Braga · 2021
Later among the works it cites.
Warpdrive: Extremely fast end-to-end deep multi-agent reinforcement learning on a GPU
T. Lan, S. Srinivasa, and S. Zheng · 2021
Later among the works it cites.
Isaac Gym: High performance GPU-based physics simulation for robot learning
V. Makoviychuk, L. Wawrzyniak, Y. Guo, M. Lu, K. Storey, M. Macklin, D. Hoeller, N. Rudin, A. Allshire, A. Handa, et al · 2021
Later among the works it cites.
Stable-Baselines3: Reliable reinforcement learning implementations
A. Raffin, A. Hill, A. Gleave, A. Kanervisto, M. Ernestus, and N. Dormann · 2021
Later among the works it cites.
The 37 implementation details of proximal policy optimization
S. Huang, R. F. J. Dossa, A. Raffin, A. Kanervisto, and W. Wang · 2022
Closest in time.
rl-games: A high-performance framework for reinforcement learning
D. Makoviichuk and V. Makoviychuk · 2022
Closest in time.
Tianshou: A highly modularized deep reinforcement learning library
J. Weng, H. Chen, D. Yan, K. You, A. Duburcq, M. Zhang, Y. Su, H. Su, and J. Zhu · 2022
Closest in time.