Fetching the paper…
Reading the bibliography…
We present the Minigrid and Miniworld libraries which provide a suite of goal-oriented 2D and 3D environments.
Partially observable markov decision processes for artificial intelligence
L. P. Kaelbling, M. L. Littman, and A. R. Cassandra · 1995
Earlier work this paper cites.
Reinforcement Learning: An Introduction
R. S. Sutton and A. G. Barto · 1998
Earlier work this paper cites.
MuJoCo: A physics engine for model-based control
E. Todorov, T. Erez, and Y. Tassa · 2012
Earlier work this paper cites.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. A. Riedmiller, A. Fidjeland, G. Ostrovski, S. Petersen, C. Beattie, A. Sadik, I. Antonoglou, H. King, D. Kumaran, D. Wierstra, S. Legg, and D. Hassabis · 2015
Earlier work this paper cites.
Mazebase: A sandbox for learning from games
S. Sukhbaatar, A. Szlam, G. Synnaeve, S. Chintala, and R. Fergus · 2015
Earlier work this paper cites.
C. Beattie, J. Z. Leibo, D. Teplyashin, T. Ward, M. Wainwright, H. Küttler, A. Lefrancq, S. Green, V. Valdés, A. Sadik, J. Schrittwieser, K. Anderson, S. York, M. Cant, A. Cain, A. Bolton, S. Gaffney, H. King, D. Hassabis, S. Legg, and S. Petersen · 2016
Earlier work this paper cites.
G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba · 2016
Earlier work this paper cites.
ViZDoom: A Doom-based AI research platform for visual reinforcement learning
M. Kempka, M. Wydmuch, G. Runc, J. Toczek, and W. Jaskowski · 2016
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. van den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, S. Dieleman, D. Grewe, J. Nham, N. Kalchbrenner, I. Sutskever, T. P. Lillicrap, M. Leach, K. Kavukcuoglu, T. Graepel, and D. Hassabis · 2016
Earlier work this paper cites.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Earlier work this paper cites.
Mastering chess and shogi by self-play with a general reinforcement learning algorithm
D. Silver, T. Hubert, J. Schrittwieser, I. Antonoglou, M. Lai, A. Guez, M. Lanctot, L. Sifre, D. Kumaran, T. Graepel, T. P. Lillicrap, K. Simonyan, and D. Hassabis · 2017
Cited alongside, same era.
BabyAI: A platform to study the sample efficiency of grounded language learning
M. Chevalier-Boisvert, D. Bahdanau, S. Lahlou, L. Willems, C. Saharia, T. H. Nguyen, and Y. Bengio · 2019
Cited alongside, same era.
Generalization in reinforcement learning with selective noise injection and information bottleneck
M. Igl, K. Ciosek, Y. Li, S. Tschiatschek, C. Zhang, S. Devlin, and K. Hofmann · 2019
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, A. Desmaison, A. Köpf, E. Z. Yang, Z. DeVito, M. Raison, A. Tejani, S. Chilamkurthy, B. Steiner, L. Fang, J. Bai, and S. Chintala · 2019
Cited alongside, same era.
Emergent complexity and zero-shot transfer via unsupervised environment design
Decoupling exploration and exploitation for meta-reinforcement learning without sacrifices
E. Z. Liu, A. Raghunathan, P. Liang, and C. Finn · 2021
Later among the works it cites.
Isaac Gym: High performance GPU based physics simulation for robot learning
V. Makoviychuk, L. Wawrzyniak, Y. Guo, M. Lu, K. Storey, M. Macklin, D. Hoeller, N. Rudin, A. Allshire, A. Handa, and G. State · 2021
Later among the works it cites.
Stable-baselines3: Reliable reinforcement learning implementations
A. Raffin, A. Hill, A. Gleave, A. Kanervisto, M. Ernestus, and N. Dormann · 2021
Later among the works it cites.
State entropy maximization with random encoders for efficient exploration
Y. Seo, L. Chen, J. Shin, H. Lee, P. Abbeel, and K. Lee · 2021
Later among the works it cites.
Safe policy optimization with local generalized linear function approximations
A. Wachi, Y. Wei, and Y. Sui · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M. Dennis, N. Jaques, E. Vinitsky, A. M. Bayen, S. Russell, A. Critch, and S. Levine · 2020
Cited alongside, same era.
Relay Policy Learning: Solving long-horizon tasks via imitation and reinforcement learning
A. Gupta, V. Kumar, C. Lynch, S. Levine, and K. Hausman · 2020
Cited alongside, same era.
Information-theoretic task selection for meta-reinforcement learning
R. L. Gutierrez and M. Leonetti · 2020
Cited alongside, same era.
Pre-trained word embeddings for goal-conditional transfer learning in reinforcement learning
M. Hutsebaut-Buysse, K. Mets, and S. Latré · 2020
Cited alongside, same era.
dm_control: Software and tasks for continuous control
S. Tunyasuvunakool, A. Muldal, Y. Doron, S. Liu, S. Bohez, J. Merel, T. Erez, T. Lillicrap, N. Heess, and Y. Tassa · 2020
Cited alongside, same era.
Griddly: A platform for AI research in games
C. Bamford · 2021
Cited alongside, same era.
T. Zhang, H. Xu, X. Wang, Y. Wu, K. Keutzer, J. E. Gonzalez, and Y. Tian · 2021
Later among the works it cites.
How to stay curious while avoiding noisy tvs using aleatoric uncertainty estimation
A. N. Mavor-Parker, K. A. Young, C. Barry, and L. D. Griffin · 2022
Later among the works it cites.
Evolving curricula with regret-based environment design
J. Parker-Holder, M. Jiang, M. Dennis, M. Samvelyan, J. N. Foerster, E. Grefenstette, and T. Rocktäschel · 2022
Later among the works it cites.
Real-world robot learning with masked visual pre-training
I. Radosavovic, T. Xiao, S. James, P. Abbeel, J. Malik, and T. Darrell · 2022
Later among the works it cites.