Fetching the paper…
Reading the bibliography…
Consistently checking the statistical significance of experimental results is the first mandatory step towards reproducible science.
Student, “The probable error of a mean,” Biometrika , pp. 1–25, 1908
1908
Earlier work this paper cites.
1936
Earlier work this paper cites.
F. Wilcoxon, “Individual comparisons by ranking methods,” Biometrics bulletin , vol. 1, no. 6, pp. 80–83, 1945
1945
Earlier work this paper cites.
B. L. Welch, “The generalization of student’s’ problem when several different population variances are involved,” Biometrika , vol. 34, no. 1/2, pp. 28–35, 1947
1947
Earlier work this paper cites.
N. Smirnov, “Table for estimating the goodness of fit of empirical distributions,” The annals of mathematical statistics , vol. 19, no. 2, pp. 279–281, 1948
1948
Earlier work this paper cites.
W. J. Conover and R. L. Iman, “Rank transformations as a bridge between parametric and nonparametric statistics,” The American Statistician , vol. 35, no. 3, pp. 124–129, 1981
1981
Cited alongside, same era.
B. Efron and R. J. Tibshirani, An introduction to the bootstrap . CRC press, 1994
1994
Cited alongside, same era.
G. K. Kanji, 100 statistical tests . Sage, 2006
2006
Cited alongside, same era.
J. C. De Winter, “Using the student’s t-test with extremely small sample sizes.” Practical Assessment, Research & Evaluation , vol. 18, no. 10, 2013
2013
Cited alongside, same era.
2016
Cited alongside, same era.
R. Islam, P. Henderson, M. Gomrokchi, and D. Precup, “Reproducibility of benchmarked deep reinforcement learning tasks for continuous control,” in Proceedings of the ICML 2017 workshop on Reproducibility in Machine Learning (RML) , 2017
2017
Later among the works it cites.
2017
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…