Fetching the paper…
Reading the bibliography…
We introduce Procgen Benchmark, a suite of 16 procedurally generated game-like environments designed to benchmark both sample efficiency and generalization in reinforcement learning.
On the shortest spanning subtree of a graph and the traveling salesman problem
J. B. Kruskal · 1956
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
X. Glorot and Y. Bengio · 2010
Earlier work this paper cites.
Cellular automata for real-time generation of infinite cave levels
L. Johnson, G. N. Yannakakis, and J. Togelius · 2010
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
M. G. Bellemare, Y. Naddaf, J. Veness, and M. Bowling · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Earlier work this paper cites.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. A. Riedmiller, A. Fidjeland, G. Ostrovski, S. Petersen, C. Beattie, A. Sadik, I. Antonoglou, H. King, D. Kumaran, D. Wierstra, S. Legg, and D. Hassabis · 2015
Earlier work this paper cites.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Earlier work this paper cites.
IMPALA: scalable distributed deep-rl with importance weighted actor-learner architectures
L. Espeholt, H. Soyer, R. Munos, K. Simonyan, V. Mnih, T. Ward, Y. Doron, V. Firoiu, T. Harley, I. Dunning, S. Legg, and K. Kavukcuoglu · 2018
Cited alongside, same era.
Generalization and regularization in DQN
J. Farebrother, M. C. Machado, and M. Bowling · 2018
Cited alongside, same era.
Rainbow: Combining improvements in deep reinforcement learning
M. Hessel, J. Modayil, H. Van Hasselt, T. Schaul, G. Ostrovski, W. Dabney, D. Horgan, B. Piot, M. Azar, and D. Silver · 2018
Cited alongside, same era.
Illuminating generalization in deep reinforcement learning through procedural level generation
N. Justesen, R. R. Torrado, P. Bontrager, A. Khalifa, J. Togelius, and S. Risi · 2018
Cited alongside, same era.
Revisiting the arcade learning environment: Evaluation protocols and open problems for general agents
Gym retro
V. Pfau, A. Nichol, C. Hesse, L. Schiavo, J. Schulman, and O. Klimov · 2018
Later among the works it cites.
Safety gym
J. Achiam, A. Ray, and D. Amodei · 2019
Closest in time.
The animal-ai environment: Training and testing animal-like artificial cognition
B. Beyret, J. Hern’andez-Orallo, L. Cheke, M. Halina, M. Shanahan, and M. Crosby · 2019
Closest in time.
Quantifying generalization in reinforcement learning
K. Cobbe, O. Klimov, C. Hesse, T. Kim, and J. Schulman · 2019
Closest in time.
Obstacle tower: A generalization challenge in vision, control, and planning
A. Juliani, A. Khalifa, V.-P. Berges, J. Harper, H. Henry, A. Crespi, J. Togelius, and D. Lange · 2019
Closest in time.
A simple randomization technique for generalization in deep reinforcement learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M. C. Machado, M. G. Bellemare, E. Talvitie, J. Veness, M. J. Hausknecht, and M. Bowling · 2018
Cited alongside, same era.
Gotta learn fast: A new benchmark for generalization in RL
A. Nichol, V. Pfau, C. Hesse, O. Klimov, and J. Schulman · 2018
Cited alongside, same era.
Assessing generalization in deep reinforcement learning
C. Packer, K. Gao, J. Kos, P. Krähenbühl, V. Koltun, and D. Song · 2018
Cited alongside, same era.
D. Perez-Liebana, J. Liu, A. Khalifa, R. D. Gaina, J. Togelius, and S. M. Lucas · 2018
Cited alongside, same era.
A dissection of overfitting and generalization in continuous reinforcement learning
A. Zhang, N. Ballas, and J. Pineau
Cited in the paper.
Natural environment benchmarks for reinforcement learning
A. Zhang, Y. Wu, and J. Pineau
Cited in the paper.
A study on overfitting in deep reinforcement learning
C. Zhang, O. Vinyals, R. Munos, and S. Bengio
Cited in the paper.
K. Lee, K. Lee, J. Shin, and H. Lee · 2019
Closest in time.
Behaviour suite for reinforcement learning
I. Osband, Y. Doron, M. Hessel, J. Aslanides, E. Sezener, A. Saraiva, K. McKinney, T. Lattimore, C. Szepesvári, S. Singh, B. Van Roy, R. Sutton, D. Silver, and H. van Hasselt · 2019
Closest in time.
Meta-world: A benchmark and evaluation for multi-task and meta-reinforcement learning, 2019
T. Yu, D. Quillen, Z. He, R. Julian, K. Hausman, C. Finn, and S. Levine · 2019
Closest in time.