Fetching the paper…
Reading the bibliography…
Deep reinforcement learning (RL) algorithms are predominantly evaluated by comparing their relative performance on a large suite of tasks.
On a test of whether one of two random variables is stochastically larger than the other
Henry B Mann and Donald R Whitney · 1947
Earlier work this paper cites.
The generalization ofstudent’s’ problem when several different population variances are involved
Bernard L Welch · 1947
Earlier work this paper cites.
A survey of sampling from contaminated distributions
John W Tukey · 1960
Earlier work this paper cites.
Bootstrap methods: another look at the jackknife
Bradley Efron · 1979
Earlier work this paper cites.
Better bootstrap confidence intervals
Bradley Efron · 1987
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Richard S Sutton · 1988
Earlier work this paper cites.
Stochastic dominance and expected utility: Survey and analysis
Haim Levy · 1992
Earlier work this paper cites.
Generalization in reinforcement learning: Successful examples using sparse coarse coding
Richard S Sutton · 1996
Earlier work this paper cites.
Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning
Richard S Sutton, Doina Precup, and Satinder Singh · 1999
Earlier work this paper cites.
Benchmarking optimization software with performance profiles
Elizabeth D Dolan and Jorge J Moré · 2002
Earlier work this paper cites.
Tree-based batch mode reinforcement learning
Damien Ernst, Pierre Geurts, and Louis Wehenkel · 2005
Earlier work this paper cites.
Why most published research findings are false
John PA Ioannidis · 2005
Earlier work this paper cites.
Learning tetris using the noisy cross-entropy method
István Szita and András Lörincz · 2006
Earlier work this paper cites.
How to assess and report the performance of a stochastic algorithm on a benchmark problem: mean or best result on a number of runs?
Mauro Birattari and Marco Dorigo · 2007
Earlier work this paper cites.
Karl Cobbe, Jacob Hilton, Oleg Klimov, and John Schulman · 2009
Earlier work this paper cites.
Resampling fewer than n observations: gains, losses, and remedies for losses
Peter J Bickel, Friedrich Götze, and Willem R van Zwet · 2012
Earlier work this paper cites.
Editors’ introduction to the special section on replicability in psychological science: A crisis of confidence?
Harold Pashler and Eric-Jan Wagenmakers · 2012
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Marc G Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling · 2013
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Earlier work this paper cites.
1,500 scientists lift the lid on reproducibility
Monya Baker · 2016
Earlier work this paper cites.
What does research reproducibility mean?
Steven N Goodman, Daniele Fanelli, and John PA Ioannidis · 2016
Earlier work this paper cites.
Statistical tests, p values, confidence intervals, and power: a guide to misinterpretations
Sander Greenland, Stephen J Senn, Kenneth J Rothman, John B Carlin, Charles Poole, Steven N Goodman, and Douglas G Altman · 2016
Earlier work this paper cites.
Reinforcement learning with unsupervised auxiliary tasks
Max Jaderberg, Volodymyr Mnih, Wojciech Marian Czarnecki, Tom Schaul, Joel Z Leibo, David Silver, and Koray Kavukcuoglu · 2016
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al · 2016
Earlier work this paper cites.
A distributional perspective on reinforcement learning
Marc G Bellemare, Will Dabney, and Rémi Munos · 2017
Earlier work this paper cites.
Reproducibility of benchmarked deep reinforcement learning tasks for continuous control
Riashat Islam, Peter Henderson, Maziar Gomrokchi, and Doina Precup · 2017
Earlier work this paper cites.
Are gans created equal? a large-scale study
Mario Lucic, Karol Kurach, Marcin Michalski, Sylvain Gelly, and Olivier Bousquet · 2017
Earlier work this paper cites.
Reporting score distributions makes a difference: Performance study of lstm-networks for sequence tagging
Nils Reimers and Iryna Gurevych · 2017
Earlier work this paper cites.
Cognitive psychology for deep neural networks: A shape bias case study
Samuel Ritter, David GT Barrett, Adam Santoro, and Matt M Botvinick · 2017
Earlier work this paper cites.
Evolution strategies as a scalable alternative to reinforcement learning
Tim Salimans, Jonathan Ho, Xi Chen, Szymon Sidor, and Ilya Sutskever · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
JAX: composable transformations of Python+NumPy programs, 2018
James Bradbury, Roy Frostig, Peter Hawkins, Matthew James Johnson, Chris Leary, Dougal Maclaurin, George Necula, Adam Paszke, Jake VanderPlas, Skye Wanderman-Milne, and Qiao Zhang · 2018
Earlier work this paper cites.
Dopamine: A research framework for deep reinforcement learning
Pablo Samuel Castro, Subhodeep Moitra, Carles Gelada, Saurabh Kumar, and Marc G Bellemare · 2018
Earlier work this paper cites.
How many random seeds? statistical power analysis in deep reinforcement learning experiments
Cédric Colas, Olivier Sigaud, and Pierre-Yves Oudeyer · 2018
Earlier work this paper cites.
Impala: Scalable distributed deep-rl with importance weighted actor-learner architectures
Lasse Espeholt, Hubert Soyer, Remi Munos, Karen Simonyan, Vlad Mnih, Tom Ward, Yotam Doron, Vlad Firoiu, Tim Harley, Iain Dunning, et al · 2018
Cited alongside, same era.
Statistical rituals: The replication delusion and how we got there
Gerd Gigerenzer · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Cited alongside, same era.
Deep reinforcement learning that matters
Peter Henderson, Riashat Islam, Philip Bachman, Joelle Pineau, Doina Precup, and David Meger · 2018
Cited alongside, same era.
Rainbow: Combining improvements in deep reinforcement learning
Matteo Hessel, Joseph Modayil, Hado Van Hasselt, Tom Schaul, Georg Ostrovski, Will Dabney, Dan Horgan, Bilal Piot, Mohammad Azar, and David Silver · 2018
Cited alongside, same era.
Fast task inference with variational intrinsic successor features
Steven Hansen, Will Dabney, Andre Barreto, David Warde-Farley, Tom Van de Wiele, and Volodymyr Mnih · 2020
Later among the works it cites.
Evaluating the performance of reinforcement learning algorithms
Scott Jordan, Yash Chandak, Daniel Cohen, Mengxue Zhang, and Philip Thomas · 2020
Later among the works it cites.
Do recent advancements in model-based deep reinforcement learning really improve data efficiency?
Kacper Kielak · 2020
Later among the works it cites.
A survey on reproducibility by evaluating deep reinforcement learning algorithms on real-world robots
Nicolai A Lynnerup, Laura Nolling, Rasmus Hasle, and John Hallam · 2020
Later among the works it cites.
Joelle Pineau, Philippe Vincent-Lamarre, Koustuv Sinha, Vincent Larivière, Alina Beygelzimer, Florence d’Alché Buc, Emily Fox, and Hugo Larochelle · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Alex Irpan · 2018
Cited alongside, same era.
Revisiting the arcade learning environment: Evaluation protocols and open problems for general agents
Marlos C Machado, Marc G Bellemare, Erik Talvitie, Joel Veness, Matthew Hausknecht, and Michael Bowling · 2018
Cited alongside, same era.
Simple random search provides a competitive approach to reinforcement learning
Horia Mania, Aurelia Guy, and Benjamin Recht · 2018
Cited alongside, same era.
On the state of the art of evaluation in neural language models
Gábor Melis, Chris Dyer, and Phil Blunsom · 2018
Cited alongside, same era.
Deterministic implementations for reproducibility in deep reinforcement learning
Prabhat Nagarajan, Garrett Warnell, and Peter Stone · 2018
Cited alongside, same era.
Benchmarking Machine Learning with Performance Profiles, 03 2018
Ben Recht · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
Richard S. Sutton and Andrew G. Barto · 2018
Cited alongside, same era.
Later among the works it cites.
Automatic data augmentation for generalization in deep reinforcement learning
Roberta Raileanu, Max Goldstein, Denis Yarats, Ilya Kostrikov, and Rob Fergus · 2020
Later among the works it cites.
Smaller world models for reinforcement learning
Jan Robine, Tobias Uelwer, and Stefan Harmeling · 2020
Later among the works it cites.
In praise of confidence intervals
David Romer · 2020
Later among the works it cites.
Mastering atari, go, chess and shogi by planning with a learned model
Julian Schrittwieser, Ioannis Antonoglou, Thomas Hubert, Karen Simonyan, Laurent Sifre, Simon Schmitt, Arthur Guez, Edward Lockhart, Demis Hassabis, Thore Graepel, et al · 2020
Later among the works it cites.
D2rl: Deep dense architectures in reinforcement learning
Samarth Sinha, Homanga Bharadhwaj, Aravind Srinivas, and Animesh Garg · 2020
Later among the works it cites.
Curl: Contrastive unsupervised representations for reinforcement learning
Aravind Srinivas, Michael Laskin, and Pieter Abbeel · 2020
Later among the works it cites.
Munchausen reinforcement learning
Nino Vieillard, Olivier Pietquin, and Matthieu Geist · 2020
Later among the works it cites.
Improving generalization in reinforcement learning with mixture regularization
Kaixin Wang, Bingyi Kang, Jie Shao, and Jiashi Feng · 2020
Later among the works it cites.
Masked contrastive representation learning for reinforcement learning
Jinhua Zhu, Yingce Xia, Lijun Wu, Jiajun Deng, Wengang Zhou, Tao Qin, and Houqiang Li · 2020
Later among the works it cites.
Accounting for variance in machine learning benchmarks
Xavier Bouthillier, Pierre Delaunay, Mirko Bronzi, Assya Trofimov, Brennan Nichyporuk, Justin Szeto, Nazanin Mohammadi Sepahvand, Edward Raff, Kanika Madan, Vikram Voleti, et al · 2021
Closest in time.
Revisiting rainbow: Promoting more insightful and inclusive deep reinforcement learning research
Johan Samir Obando Ceron and Pablo Samuel Castro · 2021
Closest in time.
Mostafa Dehghani, Yi Tay, Alexey A Gritsenko, Zhe Zhao, Neil Houlsby, Fernando Diaz, Donald Metzler, and Oriol Vinyals · 2021
Closest in time.
Bootstrap Confidence Intervals, 01 2021
Nathaniel E Helwig · 2021
Closest in time.
Muesli: Combining improvements in policy optimization
Matteo Hessel, Ivo Danihelka, Fabio Viola, Arthur Guez, Simon Schmitt, Laurent Sifre, Theophane Weber, David Silver, and Hado van Hasselt · 2021
Closest in time.
Prioritized level replay
Minqi Jiang, Ed Grefenstette, and Tim Rocktäschel · 2021
Closest in time.
Image augmentation is all you need: Regularizing deep reinforcement learning from pixels
Ilya Kostrikov*, Denis Yarats*, and Rob Fergus · 2021
Closest in time.
Q-value weighted regression: Reinforcement learning with limited data
Piotr Kozakowski, Lukasz Kaiser, Henryk Michalewski, Afroz Mohiuddin, and Katarzyna Kańska · 2021
Closest in time.
Sunrise: A simple unified framework for ensemble learning in deep reinforcement learning
Kimin Lee, Michael Laskin, Aravind Srinivas, and Pieter Abbeel · 2021
Closest in time.
Jimmy Lin, Daniel Campos, Nick Craswell, Bhaskar Mitra, and Emine Yilmaz · 2021
Closest in time.
Return-based contrastive representation learning for reinforcement learning
Guoqing Liu, Chuheng Zhang, Li Zhao, Tao Qin, Jinhua Zhu, Li Jian, Nenghai Yu, and Tie-Yan Liu · 2021
Closest in time.
Ensemble and auxiliary tasks for data-efficient deep reinforcement learning
Muhammad Rizki Maulana and Wee Sun Lee · 2021
Closest in time.
Decoupling value and policy for generalization in reinforcement learning
Roberta Raileanu and Rob Fergus · 2021
Closest in time.
Synthetic returns for long-term credit assignment
David Raposo, Sam Ritter, Adam Santoro, Greg Wayne, Theophane Weber, Matt Botvinick, Hado van Hasselt, and Francis Song · 2021
Closest in time.
Seerl: Sample efficient ensemble reinforcement learning
Rohan Saphal, Balaraman Ravindran, Dheevatsa Mudigere, Sasikant Avancha, and Bharat Kaul · 2021
Closest in time.
Return-based scaling: Yet another normalisation trick for deep rl
Tom Schaul, Georg Ostrovski, Iurii Kemaev, and Diana Borsa · 2021
Closest in time.
Data-efficient reinforcement learning with self-predictive representations
Max Schwarzer, Ankesh Anand, Rishab Goel, R Devon Hjelm, Aaron Courville, and Philip Bachman · 2021
Closest in time.
The multiberts: Bert reproductions for robustness analysis
Thibault Sellam, Steve Yadlowsky, Jason Wei, Naomi Saphra, Alexander D’Amour, Tal Linzen, Jasmijn Bastings, Iulia Turc, Jacob Eisenstein, Dipanjan Das, et al · 2021
Closest in time.
State entropy maximization with random encoders for efficient exploration
Younggyo Seo, Lili Chen, Jinwoo Shin, Honglak Lee, Pieter Abbeel, and Kimin Lee · 2021
Closest in time.
How i failed machine learning in medical imaging–shortcomings and recommendations
Gaël Varoquaux and Veronika Cheplygina · 2021
Closest in time.
Randomness in neural network training: Characterizing the impact of tooling
Donglin Zhuang, Xingyao Zhang, Shuaiwen Leon Song, and Sara Hooker · 2021
Closest in time.
Leveraging procedural generation to benchmark reinforcement learning
Karl Cobbe, Chris Hesse, Jacob Hilton, and John Schulman · 2056
Closest in time.