Fetching the paper…
Reading the bibliography…
Empirical design in reinforcement learning is no small task.
The probable error of a mean
Student · 1908
Earlier work this paper cites.
Multiple comparisons among means
Olive Jean Dunn · 1961
Earlier work this paper cites.
How not to lie with statistics: The correct way to summarize benchmark results
Philip J. Fleming and John J. Wallace · 1986
Earlier work this paper cites.
Efficient Memory-Based Learning for Robot Control
Andrew William Moore · 1990
Earlier work this paper cites.
Residual Algorithms: Reinforcement Learning with Function Approximation
Leemon Baird · 1995
Earlier work this paper cites.
Testing heuristics: We have it all wrong
John N. Hooker · 1995
Earlier work this paper cites.
Temporal difference learning and TD-Gammon
Gerald Tesauro · 1995
Earlier work this paper cites.
Generalization in reinforcement learning: Successful examples using sparse coarse coding
Richard S. Sutton · 1996
Earlier work this paper cites.
An analysis of temporal-difference learning with function approximation
J.N. Tsitsiklis and B. Van Roy · 1997
Earlier work this paper cites.
Calibration of ρ \rho values for testing precise null hypotheses
Thomas Sellke, M. J. Bayarri, and James O. Berger · 2001
Earlier work this paper cites.
Morris water maze: Procedures for assessing spatial and related forms of learning and memory
Charles V. Vorhees and Michael T. Williams · 2006
Earlier work this paper cites.
An analysis of model-based interval estimation for Markov decision processes
Alexander L. Strehl and Michael L. Littman · 2008
Earlier work this paper cites.
The many faces of optimism: A unifying approach
István Szita and András Lorincz · 2008
Earlier work this paper cites.
Fast gradient-descent methods for temporal-difference learning with linear function approximation
Richard S. Sutton, Hamid Reza Maei, Doina Precup, Shalabh Bhatnagar, David Silver, Csaba Szepesvári, and Eric Wiewiora · 2009
Earlier work this paper cites.
Generalized domains for empirical evaluations in reinforcement learning
Shimon Whiteson, Brian Tanner, Matthew E. Taylor, and Peter Stone · 2009
Earlier work this paper cites.
Evaluating Learning Algorithms: A Classification Perspective
Nathalie Japkowicz and Mohak Shah · 2011
Earlier work this paper cites.
The Fixed Points of Off-Policy TD
J Z Kolter · 2011
Earlier work this paper cites.
Protecting against evaluation overfitting in empirical reinforcement learning
Shimon Whiteson, Brian Tanner, Matthew E. Taylor, and Peter Stone · 2011
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Cited alongside, same era.
The Arcade Learning Environment: An Evaluation Platform for General Agents
M. G. Bellemare, Y. Naddaf, J. Veness, and M. Bowling · 2013
Cited alongside, same era.
Playing Atari with Deep Reinforcement Learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Cited alongside, same era.
Adaptive Step-Sizes for Reinforcement Learning
William C Dabney · 2014
Cited alongside, same era.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Cited alongside, same era.
Benchmarking deep reinforcement learning for continuous control
Revisiting the Arcade Learning Environment: Evaluation Protocols and Open Problems for General Agents
Marlos C. Machado, Marc G. Bellemare, Erik Talvitie, Joel Veness, Matthew Hausknecht, and Michael Bowling · 2018
Later among the works it cites.
Reinforcement Learning: An Introduction
Richard S. Sutton and Andrew G. Barto · 2018
Later among the works it cites.
Implementation Matters In Deep Policy Gradients: A Case Study On PPO and TRPO
Logan Engstrom, Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras, Firdaus Janoos, Larry Rudolph, and Aleksander Ma · 2019
Later among the works it cites.
Soft Actor-Critic Algorithms and Applications, 2019
Tuomas Haarnoja, Aurick Zhou, Kristian Hartikainen, George Tucker, Sehoon Ha, Jie Tan, Vikash Kumar, Henry Zhu, Abhishek Gupta, Pieter Abbeel, and Sergey Levine · 2019
Later among the works it cites.
Understanding multi-step deep reinforcement learning: A systematic study of the DQN target
J. Fernando Hernandez-Garcia and Richard S. Sutton · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yan Duan, Xi Chen, Rein Houthooft, John Schulman, and Pieter Abbeel · 2016
Cited alongside, same era.
An Emphatic Approach to the Problem of Off-policy Temporal-Difference Learning
Richard S Sutton, A Rupam Mahmood, and Martha White · 2016
Cited alongside, same era.
The ASA Statement on p-Values: Context, Process, and Purpose
Ronald L. Wasserstein and Nicole A. Lazar · 2016
Cited alongside, same era.
The reproducibility of research and the misinterpretation of p-values
David Colquhoun · 2017
Cited alongside, same era.
Reproducibility of Benchmarked Deep Reinforcement Learning Tasks for Continuous Control
Riashat Islam, Peter Henderson, Maziar Gomrokchi, and Doina Precup · 2017
Cited alongside, same era.
Incremental off-policy reinforcement learning algorithms
Ashique Mahmood · 2017
Cited alongside, same era.
Stronger generalization bounds for deep nets via a compression approach
Sanjeev Arora, Rong Ge, Behnam Neyshabur, and Yi Zhang · 2018
Cited alongside, same era.
DeepMellow: Removing the Need for a Target Network in Deep Q-Learning
Seungchan Kim, Kavosh Asadi, Michael Littman, and George Konidaris · 2019
Later among the works it cites.
Gradient Temporal-Difference Learning with Regularized Corrections
Sina Ghiassian, Andrew Patterson, Shivam Garg, Dhawal Gupta, Adam White, and Martha White · 2020
Later among the works it cites.
Representations for Stable Off-Policy Reinforcement Learning
Dibya Ghosh and Marc G. Bellemare · 2020
Later among the works it cites.
Evaluating the Performance of Reinforcement Learning Algorithms
Scott M. Jordan, Yash Chandak, Daniel Cohen, Mengxue Zhang, and Philip S. Thomas · 2020
Later among the works it cites.
On bonus-based exploration methods in the arcade learning environment
Adrien Ali Taïga, William Fedus, Marlos C. Machado, Aaron Courville, and Marc G. Bellemare · 2020
Later among the works it cites.
Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning
Tianhe Yu, Deirdre Quillen, Zhanpeng He, Ryan Julian, Karol Hausman, Chelsea Finn, and Sergey Levine · 2020
Later among the works it cites.
Deep reinforcement learning at the edge of the statistical precipice
Rishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville, and Marc Bellemare · 2021
Later among the works it cites.
AutoML: A survey of the state-of-the-art
Xin He, Kaiyong Zhao, and Xiaowen Chu · 2021
Later among the works it cites.
Automated reinforcement learning (autorl): A survey and open problems
Jack Parker-Holder, Raghu Rajan, Xingyou Song, André Biedenkapp, Yingjie Miao, Theresa Eimer, Baohe Zhang, Vu Nguyen, Roberto Calandra, Aleksandra Faust, et al · 2022
Later among the works it cites.
Hyperparameters in reinforcement learning and how to tune them
Theresa Eimer, Marius Lindauer, and Roberta Raileanu · 2023
Closest in time.
A method for evaluating hyperparameter sensitivity in reinforcement learning
Jacob Adkins, Michael Bowling, and Adam White · 2024
Closest in time.
Position: Benchmarking is limited in reinforcement learning research
Scott M Jordan, Adam White, Bruno Castro Da Silva, Martha White, and Philip S Thomas · 2024
Closest in time.
On the consistency of hyper-parameter selection in value-based deep reinforcement learning
Johan Obando-Ceron, João GM Araújo, Aaron Courville, and Pablo Samuel Castro · 2024
Closest in time.