Fetching the paper…
Reading the bibliography…
We present a review of popular simulation engines and frameworks used in reinforcement learning (RL) research, aiming to guide researchers in selecting tools for creating simulated physical environments for RL and training setups.
N. Koenig and A. Howard, “Design and use paradigms for gazebo, an open-source multi-robot simulator,” vol. 3, pp. 2149–2154 vol.3, 2004
2004
Earlier work this paper cites.
O. Michel, “Webotstm: Professional mobile robot simulation,”
2004
Earlier work this paper cites.
S. Legg and M. Hutter, “Universal intelligence: A definition of machine intelligence,” 2007
2007
Earlier work this paper cites.
E. Todorov, T. Erez, and Y. Tassa, “Mujoco: A physics engine for model-based control,” pp. 5026–5033, 2012
2012
Earlier work this paper cites.
S. Ivaldi, V. Padois, and F. Nori, “Tools for dynamics simulation of robots: a survey based on user feedback,” 2014
2014
Earlier work this paper cites.
M. Abadi, A. Agarwal, P. Barham, E. Brevdo, Z. Chen, C. Citro, G. S. Corrado, A. Davis, J. Dean, M. Devin, S. Ghemawat, I. Goodfellow, A. Harp, G. Irving, M. Isard, Y. Jia, R. Jozefowicz, L. Kaiser, M. Kudlur, J. Levenberg, D. Mané, R. Monga, S. Moore, D. Murray, C. Olah, M. Schuster, J. Shlens, B. Steiner, I. Sutskever, K. Talwar, P. Tucker, V. Vanhoucke, V. Vasudevan, F. Viégas, O. Vinyals, P. Warden, M. Wattenberg, M. Wicke, Y. Yu, and X. Zheng, “TensorFlow: Large-scale machine learning on heterogeneous systems,” 2015, software available from tensorflow.org. [Online]. Available:
2015
Earlier work this paper cites.
T. Erez, Y. Tassa, and E. Todorov, “Simulation tools for model-based robotics: Comparison of bullet, havok, mujoco, ode and physx,” 05 2015
2015
Earlier work this paper cites.
C. Beattie, J. Z. Leibo, D. Teplyashin, T. Ward, M. Wainwright, H. Küttler, A. Lefrancq, S. Green, V. Valdés, A. Sadik, J. Schrittwieser, K. Anderson, S. York, M. Cant, A. Cain, A. Bolton, S. Gaffney, H. King, D. Hassabis, S. Legg, and S. Petersen, “Deepmind lab,” 2016
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
——, “Pybullet, a python module for physics simulation for games, robotics and machine learning.” 2016
2016
Earlier work this paper cites.
M. Johnson, K. Hofmann, T. Hutton, D. Bignell, and K. Hofmann, “The malmo platform for artificial intelligence experimentation,” July 2016. [Online]. Available:
2016
Earlier work this paper cites.
M. Kempka, M. Wydmuch, G. Runc, J. Toczek, and W. Jaśkowski, “Vizdoom: A doom-based ai research platform for visual reinforcement learning,” 2016
2016
Earlier work this paper cites.
A. Tasora, R. Serban, H. Mazhar, A. Pazouki, D. Melanz, J. Fleischmann, M. Taylor, H. Sugiyama, and D. Negrut, “Chrono: An open source multi-physics dynamics engine,” pp. 19–49, 06 2016
2016
Earlier work this paper cites.
2017
Earlier work this paper cites.
D. Silver, T. Hubert, J. Schrittwieser, I. Antonoglou, M. Lai, A. Guez, M. Lanctot, L. Sifre, D. Kumaran, T. Graepel, T. Lillicrap, K. Simonyan, and D. Hassabis, “Mastering chess and shogi by self-play with a general reinforcement learning algorithm,” 2017
2017
Earlier work this paper cites.
Y. Tassa, Y. Doron, A. Muldal, T. Erez, Y. Li, D. de Las Casas, D. Budden, A. Abdolmaleki, J. Merel, A. Lefrancq, T. Lillicrap, and M. Riedmiller, “Deepmind control suite,” 2018
2018
Earlier work this paper cites.
A. Juliani, A. Khalifa, V.-P. Berges, J. Harper, E. Teng, H. Henry, A. Crespi, J. Togelius, and D. Lange, “Obstacle tower: A generalization challenge in vision, control, and planning,” 2019
2019
Earlier work this paper cites.
N. G. Lopez, Y. L. E. Nuin, E. B. Moral, L. U. S. Juan, A. S. Rueda, V. M. Vilches, and R. Kojcev, “gym-gazebo2, a toolkit for reinforcement learning using ros 2 and gazebo,” 2019
2019
Earlier work this paper cites.
OpenAI, M. Andrychowicz, B. Baker, M. Chociej, R. Jozefowicz, B. McGrew, J. Pachocki, A. Petron, M. Plappert, G. Powell, A. Ray, J. Schneider, S. Sidor, J. Tobin, P. Welinder, L. Weng, and W. Zaremba, “Learning dexterous in-hand manipulation,” 2019
2019
Earlier work this paper cites.
OpenAI, C. Berner, G. Brockman, B. Chan, V. Cheung, P. Dębiak, C. Dennison, D. Farhi, Q. Fischer, S. Hashme, C. Hesse, R. Józefowicz, S. Gray, C. Olsson, J. Pachocki, M. Petrov, H. P. d. O. Pinto, J. Raiman, T. Salimans, J. Schlatter, J. Schneider, S. Sidor, I. Sutskever, J. Tang, F. Wolski, and S. Zhang, “Dota 2 with large scale deep reinforcement learning,” 2019
2019
Earlier work this paper cites.
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga
2019
Cited alongside, same era.
O. Vinyals, I. Babuschkin, W. Czarnecki, M. Mathieu, A. Dudzik, J. Chung, D. Choi, R. Powell, T. Ewalds, P. Georgiev, J. Oh, D. Horgan, M. Kroiss, I. Danihelka, A. Huang, L. Sifre, T. Cai, J. Agapiou, M. Jaderberg, and D. Silver, “Grandmaster level in starcraft ii using multi-agent reinforcement learning,”
2019
Cited alongside, same era.
B. Baker, I. Kanitscheider, T. Markov, Y. Wu, G. Powell, B. McGrew, and I. Mordatch, “Emergent tool use from multi-agent autocurricula,” 2020
2020
Cited alongside, same era.
S. Benatti, A. Tasora, and D. Mangoni, “Training a four legged robot via deep reinforcement learning and multibody simulation,” pp. 391–398, 2020
2020
Cited alongside, same era.
Open-Ended-Learning-Team, A. Stooke, A. Mahajan, C. Barros, C. Deck, J. Bauer, J. Sygnowski, M. Trebacz, M. Jaderberg, M. Mathieu, N. McAleese, N. Bradley-Schmieg, N. Wong, N. Porcel, R. Raileanu, S. Hughes-Fitt, V. Dalibard, and W. M. Czarnecki, “Open-ended learning leads to generally capable agents,” 2021
2021
Later among the works it cites.
J. Panerati, H. Zheng, S. Zhou, J. Xu, A. Prorok, and A. P. Schoellig, “Learning to fly—a gym environment with pybullet physics for reinforcement learning of multi-agent quadcopter control,” in
2021
Later among the works it cites.
S. Pateria, B. Subagdja, A.-H. Tan, and C. Quek, “End-to-end hierarchical reinforcement learning with integrated subgoal discovery,”
2021
Later among the works it cites.
J. Terry, B. Black, N. Grammel, M. Jayakumar, A. Hari, R. Sullivan, L. S. Santos, C. Dieffendahl, C. Horsch, R. Perez-Vicente
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2020
Cited alongside, same era.
A. Juliani, V.-P. Berges, E. Teng, A. Cohen, J. Harper, C. Elion, C. Goy, Y. Gao, H. Henry, M. Mattar, and D. Lange, “Unity: A general platform for intelligent agents,” 2020
2020
Cited alongside, same era.
M. Kirtas, K. Tsampazis, N. Passalis, and A. Tefas, “Deepbots: A webots-based deep reinforcement learning framework for robotics,” in
2020
Cited alongside, same era.
A. P. Mohammed and M. Valdenegro-Toro, “Can reinforcement learning for continuous control generalize across physics engines?” 2020
2020
Cited alongside, same era.
NVIDIA. (2020) Nvidia physx. [Online]. Available:
2020
Cited alongside, same era.
J. Schrittwieser, I. Antonoglou, T. Hubert, K. Simonyan, L. Sifre, S. Schmitt, A. Guez, E. Lockhart, D. Hassabis, T. Graepel, T. Lillicrap, and D. Silver, “Mastering atari, go, chess and shogi by planning with a learned model,”
2020
Cited alongside, same era.
T. Ward, A. Bolt, N. Hemmings, S. Carter, M. Sanchez, R. Barreira, S. Noury, K. Anderson, J. Lemmon, J. Coe, P. Trochim, T. Handley, and A. Bolton, “Using unity to help solve intelligence,” 2020
2020
Cited alongside, same era.
Y. Zhou, S. Manuel, P. Morales, S. Li, J. Peña, and R. Allen, “Towards a distributed framework for multi-agent reinforcement learning research,”
2020
Cited alongside, same era.
J. Albrecht, A. Fetterman, B. Fogelman, E. Kitanidis, B. Wróblewski, N. Seo, M. Rosenthal, M. Knutins, Z. Polizzi, J. Simon, and K. Qiu, “Avalon: A benchmark for rl generalization using procedurally generated worlds,” in
2022
Later among the works it cites.
S. Benatti, A. Young, A. Elmquist, J. Taves, R. Serban, D. Mangoni, A. Tasora, and D. Negrut, “Pychrono and gym-chrono: A deep reinforcement learning framework leveraging multibody dynamics to control autonomous vehicles and robots,” pp. 573–584, 01 2022
2022
Later among the works it cites.
S. Benatti, A. Young, A. Elmquist, J. Taves, A. Tasora, R. Serban, and D. Negrut, “End-to-end learning for off-road terrain navigation using the chrono open-source simulation platform,”
2022
Later among the works it cites.
M. Bettini, R. Kortvelesy, J. Blumenkamp, and A. Prorok, “Vmas: A vectorized multi-agent simulator for collective robot learning,” 2022
2022
Later among the works it cites.
J. Chen, F. Deng, Y. Gao, J. Hu, X. Guo, G. Liang, and T. L. Lam, “Multirobolearn: An open-source framework for multi-robot deep reinforcement learning,” 2022
2022
Later among the works it cites.
Y. Chen, Y. Yang, T. Wu, S. Wang, X. Feng, J. Jiang, Z. Lu, S. M. McAleer, H. Dong, and S.-C. Zhu, “Towards human-level bimanual dexterous manipulation with reinforcement learning,” in
2022
Later among the works it cites.
M. Feng, W. Zhou, Y. Yang, and H. Li, “Joint-predictive representations for multi-agent reinforcement learning,” 2022
2022
Later among the works it cites.
L. M. Schmidt, J. Brosig, A. Plinge, B. M. Eskofier, and C. Mutschler, “An introduction to multi-agent reinforcement learning and review of its application to autonomous mobility,” in
2022
Later among the works it cites.
J. Weng, M. Lin, S. Huang, B. Liu, D. Makoviichuk, V. Makoviychuk, Z. Liu, Y. Song, T. Luo, Y. Jiang
2022
Later among the works it cites.
A. Young, J. Taves, A. Elmquist, S. Benatti, A. Tasora, R. Serban, and D. Negrut, “Enabling artificial intelligence studies in off-road mobility through physics-based simulation of multiagent scenarios,”
2022
Later among the works it cites.
E. Coumans and Y. Bai, “Pybulletquickstartguide - github,”
2023
Later among the works it cites.
DeepMind-Adaptive-Agents-Team, J. Bauer, K. Baumli, S. Baveja, F. Behbahani, A. Bhoopchand, N. Bradley-Schmieg, M. Chang, N. Clay, A. Collister, V. Dasagi, L. Gonzalez, K. Gregor, E. Hughes, S. Kashem, M. Loks-Thompson, H. Openshaw, J. Parker-Holder, S. Pathak, N. Perez-Nieves, N. Rakicevic, T. Rocktäschel, Y. Schroecker, J. Sygnowski, K. Tuyls, S. York, A. Zacherl, and L. Zhang, “Human-timescale adaptation in an open-ended task space,” 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
J. S. Park, J. C. O’Brien, C. J. Cai, M. R. Morris, P. Liang, and M. S. Bernstein, “Generative agents: Interactive simulacra of human behavior,” 2023
2023
Later among the works it cites.
M. L. Trang, “Multi-task reinforcement learning: From single-agent to multi-agent systems.” Ph.D. dissertation, Virginia Tech, 2023
2023
Later among the works it cites.
Unity, “Learning environment examples - unity,”
2023
Later among the works it cites.