Fetching the paper…
Reading the bibliography…
Strong empirical evidence that one machine-learning algorithm A outperforms another one B ideally calls for multiple trials optimizing the learning pipeline over sources of variation such as data sampling, data augmentation, parameter initialization, and hyperparameters choices.
A Step Toward Quantifying Independently Reproducible Machine Learning Research
Raff, E · 1909
Earlier work this paper cites.
On the use and interpretation of certain test criteria for purposes of statistical inference: Part i
Neyman, J. and Pearson, E. S · 1928
Earlier work this paper cites.
Bootstrap methods: Another look at the jackknife
Efron, B · 1979
Earlier work this paper cites.
The jackknife, the bootstrap, and other resampling plans , volume 38
Efron, B · 1982
Earlier work this paper cites.
Sample size determination for some common nonparametric tests
Noether, G. E · 1987
Earlier work this paper cites.
Amino acid substitution matrices from protein blocks
Henikoff, S. and Henikoff, J. G · 1992
Earlier work this paper cites.
An introduction to the bootstrap
Efron, B. and Tibshirani, R. J · 1994
Earlier work this paper cites.
Approximate statistical tests for comparing supervised classification learning algorithms
Dietterich, T. G · 1998
Earlier work this paper cites.
Inference for the generalization error
Nadeau, C. and Bengio, Y · 2000
Earlier work this paper cites.
Analyzing bagging
Bühlmann, P., Yu, B., et al · 2002
Earlier work this paper cites.
Multiple hypothesis testing in microarray experiments
Dudoit, S., Shaffer, J. P., and Boldrick, J. C · 2003
Earlier work this paper cites.
A Metric Learning Reality Check
Musgrave, K., Belongie, S., and Lim, S.-N · 2003
Earlier work this paper cites.
Evaluating the replicability of significance tests for comparing learning algorithms
Bouckaert, R. R. and Frank, E · 2004
Earlier work this paper cites.
The design and analysis of benchmark experiments
Hothorn, T., Leisch, F., Zeileis, A., and Hornik, K · 2005
Earlier work this paper cites.
On some pitfalls in automatic evaluation and significance testing for mt
Riezler, S. and Maxwell, J. T · 2005
Earlier work this paper cites.
Bootstrap diagnostics and remedies
Canty, A. J., Davison, A. C., Hinkley, D. V., and Ventura, V · 2006
Earlier work this paper cites.
Statistical comparisons of classifiers over multiple data sets
Demšar, J · 2006
Earlier work this paper cites.
Netmhcpan, a method for quantitative predictions of peptide binding to any hla-a and-b locus protein of known sequence
Nielsen, M., Lundegaard, C., Blicher, T., Lamberth, K., Harndahl, M., Justesen, S., Røder, G., Peters, B., Sette, A., Lund, O., et al · 2007
Earlier work this paper cites.
80 million tiny images: A large data set for nonparametric object and scene recognition
Torralba, A., Fergus, R., and Freeman, W. T · 2008
Earlier work this paper cites.
The fifth PASCAL recognizing textual entailment challenge
Bentivogli, L., Dagan, I., Dang, H. T., Giampiccolo, D., and Magnini, B · 2009
Earlier work this paper cites.
ImageNet: A Large-Scale Hierarchical Image Database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A., Hinton, G., et al · 2009
Cited alongside, same era.
Understanding the difficulty of training deep feedforward neural networks
Glorot, X. and Bengio, Y · 2010
Cited alongside, same era.
The PASCAL Visual Object Classes Challenge 2012 (VOC2012) Results
Everingham, M., Van Gool, L., Williams, C. K. I., Winn, J., and Zisserman, A · 2012
Cited alongside, same era.
Research Reproducibility as a Survival Analysis
Raff, E · 2012
Cited alongside, same era.
An empirical investigation of statistical significance in nlp
Taylor Berg-Kirkpatrick, D. B. and Klein, D · 2012
Cited alongside, same era.
Recursive deep models for semantic compositionality over a sentiment treebank
Reporting score distributions makes a difference: Performance study of LSTM-networks for sequence tagging
Reimers, N. and Gurevych, I · 2017
Later among the works it cites.
Attention is all you need, 2017
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Later among the works it cites.
Bert: Pre-training of deep bidirectional transformers for language understanding, 2018
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Later among the works it cites.
Deep reinforcement learning that matters
Henderson, P., Islam, R., Bachman, P., Pineau, J., Precup, D., and Meger, D · 2018
Later among the works it cites.
Progressive neural architecture search
Liu, C., Zoph, B., Neumann, M., Shlens, J., Hua, W., Li, L.-J., Fei-Fei, L., Yuille, A., Huang, J., and Murphy, K · 2018
Later among the works it cites.
Are gans created equal? a large-scale study
Lucic, M., Kurach, K., Michalski, M., Gelly, S., and Bousquet, O · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Socher, R., Perelygin, A., Wu, J., Chuang, J., Manning, C. D., Ng, A., and Potts, C · 2013
Cited alongside, same era.
What’s in a p-value in nlp?
Anders Sogaard, Anders Johannsen, B. P. D. H. and Alonso, H. M · 2014
Cited alongside, same era.
An efficient approach for assessing hyperparameter importance
Hutter, F., Hoos, H., and Leyton-Brown, K · 2014
Cited alongside, same era.
Auto-encoding variational bayes
Kingma, D. P. and Welling, M · 2014
Cited alongside, same era.
Fully convolutional networks for semantic segmentation, 2014
Long, J., Shelhamer, E., and Darrell, T · 2014
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
Simonyan, K. and Zisserman, A · 2014
Cited alongside, same era.
Dropout: A simple way to prevent neural networks from overfitting
Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R · 2014
Cited alongside, same era.
Later among the works it cites.
Exploring the limits of weakly supervised pretraining
Mahajan, D., Girshick, R., Ramanathan, V., He, K., Paluri, M., Li, Y., Bharambe, A., and van der Maaten, L · 2018
Later among the works it cites.
On the state of the art of evaluation in neural language models
Melis, G., Dyer, C., and Blunsom, P · 2018
Later among the works it cites.
Mhcflurry: open-source class i mhc binding affinity prediction
O’Donnell, T. J., Rubinsteyn, A., Bonsack, M., Riemer, A. B., Laserson, U., and Hammerbacher, J · 2018
Later among the works it cites.
Unreproducible research is reproducible
Bouthillier, X., Laurent, C., and Vincent, P · 2019
Later among the works it cites.
Are we really making much progress? A Worrying Analysis of Recent Neural Recommendation Approaches
Dacrema, M. F., Cremonesi, P., and Jannach, D · 2019
Later among the works it cites.
We need to talk about standard splits
Gorman, K. and Bedrick, S · 2019
Later among the works it cites.
Confidence intervals for the mann–whitney test
Perme, M. P. and Manevski, D · 2019
Later among the works it cites.
The immune epitope database (iedb): 2018 update
Vita, R., Mahajan, S., Overton, J. A., Dhanda, S. K., Martini, S., Cantrell, J. R., Wheeler, D. K., Sette, A., and Peters, B · 2019
Later among the works it cites.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R · 2019
Later among the works it cites.
Huggingface’s transformers: State-of-the-art natural language processing
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., and Brew, J · 2019
Later among the works it cites.
Self-training with noisy student improves imagenet classification
Xie, Q., Hovy, E., Luong, M.-T., and Le, Q. V · 2019
Later among the works it cites.
What is the State of Neural Network Pruning?
Blalock, D., Gonzalez Ortiz, J. J., Frankle, J., and Guttag, J · 2020
Later among the works it cites.
Survey of machine-learning experimental methods at NeurIPS2019 and ICLR2020
Bouthillier, X. and Varoquaux, G · 2020
Later among the works it cites.
Mlperf inference benchmark
Reddi, V. J., Cheng, C., Kanter, D., Mattson, P., Schmuelling, G., Wu, C.-J., Anderson, B., Breughe, M., Charlebois, M., Chou, W., et al · 2020
Later among the works it cites.