Fetching the paper…
Reading the bibliography…
The NeurIPS 2020 Procgen Competition was designed as a centralized benchmark with clearly defined tasks for measuring Sample Efficiency and Generalization in Reinforcement Learning.
Cellular automata for real-time generation of infinite cave levels
Lawrence Johnson, Georgios N Yannakakis, and Julian Togelius · 2010
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
M. G. Bellemare, Y. Naddaf, J. Veness, and M. Bowling · 2013
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Kavukcuoglu Koray, Silver David, Andrei A. Rusu, Joel Veness1, Marc G. Bellemare, Alex Graves, Riedmiller Martin, Andreas K. Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran1, Daan Wierstra, Shane Legg, and Demis Hassabis · 2015
Earlier work this paper cites.
The malmo platform for artificial intelligence experimentation
Matthew Johnson, Katja Hofmann, Tim Hutton, and David Bignell · 2016
Earlier work this paper cites.
High-dimensional continuous control using generalized advantage estimation
John Schulman, Philipp Moritz, Sergey Levine, Michael Jordan, and Pieter Abbeel · 2016
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
The reactor: A fast and sample-efficient actor-critic agent for reinforcement learning
Audrunas Gruslys, Will Dabney, Mohammad Gheshlaghi Azar, Bilal Piot, Marc Bellemare, and Remi Munos · 2017
Earlier work this paper cites.
Learning to run challenge: Synthesizing physiologically accurate motion using deep reinforcement learning
Łukasz Kidziński, Sharada P Mohanty, Carmichael F Ong, Jennifer L Hicks, Sean F Carroll, Sergey Levine, Marcel Salathé, and Scott L Delp · 2018
Earlier work this paper cites.
Gotta learn fast: A new benchmark for generalization in rl
Alex Nichol, Vicki Pfau, Christopher Hesse, Oleg Klimov, and John Schulman · 2018
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Earlier work this paper cites.
Super-convergence: Very fast training of neural networks using large learning rates, 2018
Leslie N. Smith and Nicholay Topin · 2018
Cited alongside, same era.
Exploration by random network distillation, 2018
Yuri Burda, Harrison Edwards, Amos Storkey, and Oleg Klimov · 2018
Cited alongside, same era.
IMPALA: scalable distributed deep-rl with importance weighted actor-learner architectures
Lasse Espeholt, Hubert Soyer, Rémi Munos, Karen Simonyan, Volodymyr Mnih, Tom Ward, Yotam Doron, Vlad Firoiu, Tim Harley, Iain Dunning, Shane Legg, and Koray Kavukcuoglu · 2018
Cited alongside, same era.
IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures
Lasse Espeholt, Hubert Soyer, Remi Munos, Karen Simonyan, Volodymyr Mnih, Tom Ward, Boron Yotam, Firoiu Vlad, Harley Tim, Iain Dunning, Shane Legg, and Koray Kavukcuoglu · 2018
Cited alongside, same era.
Implicit quantile networks for distributional reinforcement learning
Will Dabney, Georg Ostrovski, David Silver, and Remi Munos · 2018
Deepmdp: Learning continuous latent space models for representation learning, 2019
Carles Gelada, Saurabh Kumar, Jacob Buckman, Ofir Nachum, and Marc G. Bellemare · 2019
Later among the works it cites.
Leveraging procedural generation to benchmark reinforcement learning, 2020
Karl Cobbe, Christopher Hesse, Jacob Hilton, and John Schulman · 2020
Later among the works it cites.
Jelly bean world: A testbed for never-ending learning
Emmanouil Antonios Platanios, Abulhair Saparov, and Tom Mitchell · 2020
Later among the works it cites.
Karl Cobbe, Jacob Hilton, Oleg Klimov, and John Schulman · 2020
Later among the works it cites.
P3o: Policy-on policy-off policy optimization
Rasool Fakoor, Pratik Chaudhari, and Alexander J. Smola · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
CBAM: Convolutional block attention module
Sanghyun Woo, Jongchan Park, Joon Young Lee, and In So Kweon · 2018
Cited alongside, same era.
The minerl competition on sample efficient reinforcement learning using human priors
William H Guss, Cayden Codel, Katja Hofmann, Brandon Houghton, Noboru Kuno, Stephanie Milani, Sharada Mohanty, Diego Perez Liebana, Ruslan Salakhutdinov, Nicholay Topin, et al · 2019
Cited alongside, same era.
MineRL: A large-scale dataset of Minecraft demonstrations
William H. Guss, Brandon Houghton, Nicholay Topin, Phillip Wang, Cayden Codel, Manuela Veloso, and Ruslan Salakhutdinov · 2019
Cited alongside, same era.
Making convolutional networks shift-invariant again
Richard Zhang · 2019
Cited alongside, same era.
Noisy networks for exploration, 2019
Meire Fortunato, Mohammad Gheshlaghi Azar, Bilal Piot, Jacob Menick, Ian Osband, Alex Graves, Vlad Mnih, Remi Munos, Demis Hassabis, Olivier Pietquin, Charles Blundell, and Shane Legg · 2019
Cited alongside, same era.
https://azure.microsoft.com/en-us/services/batch/
Microsoft Azure Batch
Cited in the paper.
Michael Laskin, Kimin Lee, Adam Stooke, Lerrel Pinto, Pieter Abbeel, and Aravind Srinivas · 2020
Later among the works it cites.
Working memory graphs
Ricky Loynd, Roland Fernandez, Asli Celikyilmaz, Adith Swaminathan, and Matthew Hausknecht · 2020
Later among the works it cites.
Squeeze-and-Excitation Networks
Jie Hu, Li Shen, Samuel Albanie, Gang Sun, and Enhua Wu · 2020
Later among the works it cites.
Curl: Contrastive unsupervised representations for reinforcement learning, 2020
Aravind Srinivas, Michael Laskin, and Pieter Abbeel · 2020
Later among the works it cites.