Fetching the paper…
Reading the bibliography…
Medical imaging is an important research field with many opportunities for improving patients' health.
The file drawer problem and tolerance for null results
R. Rosenthal · 1979
Earlier work this paper cites.
Artificial intelligence in medicine, 1987
W. B. Schwartz, R. S. Patil, and P. Szolovits · 1987
Earlier work this paper cites.
HARKing: hypothesizing after the results are known
N. L. Kerr · 1998
Earlier work this paper cites.
Points to consider on switching between superiority and non-inferiority
E. A. for the Evaluation of Medicinal Products · 2001
Earlier work this paper cites.
Non-inferiority trials: design concepts and issues–the encounters of academic consultants in statistics
R. B. D’Agostino Sr, J. M. Massaro, and L. M. Sullivan · 2003
Earlier work this paper cites.
Why most published research findings are false
J. P. A. Ioannidis · 2005
Earlier work this paper cites.
Ways toward an early diagnosis in Alzheimer’s disease: the Alzheimer’s Disease Neuroimaging Initiative (ADNI)
S. G. Mueller, M. W. Weiner, L. J. Thal, R. C. Petersen, C. R. Jack, W. Jagust, J. Q. Trojanowski, A. W. Toga, and L. Beckett · 2005
Earlier work this paper cites.
Statistical comparisons of classifiers over multiple data sets
J. Demšar · 2006
Earlier work this paper cites.
Methodology of superiority vs. equivalence trials and non-inferiority trials
E. Christensen · 2007
Earlier work this paper cites.
On the appropriateness of statistical tests in machine learning
J. Demšar · 2008
Earlier work this paper cites.
A comparison of AUC estimators in small-sample studies
A. Airola, T. Pahikkala, W. Waegeman, B. De Baets, and T. Salakoski · 2009
Earlier work this paper cites.
Image similarity and tissue overlaps as surrogates for image registration accuracy: widely used but unreliable
T. Rohlfing · 2011
Earlier work this paper cites.
Through the looking glass: understanding non-inferiority
J. Schumi and J. T. Wittes · 2011
Earlier work this paper cites.
Machine learning that matters
K. L. Wagstaff · 2012
Earlier work this paper cites.
Do we need hundreds of classifiers to solve real world classification problems?
M. Fernández-Delgado, E. Cernadas, S. Barro, D. Amorim, and D. Amorim Fernández-Delgado · 2014
Earlier work this paper cites.
Increasing disparities between resource inputs and outcomes, as measured by certain health deliverables, in biomedical research
A. Bowen and A. Casadevall · 2015
Earlier work this paper cites.
Failure: Why science is so successful
S. Firestein · 2015
Earlier work this paper cites.
Performance evaluation in machine learning
N. Japkowicz and M. Shah · 2015
Earlier work this paper cites.
Dealing with the evaluation of supervised classification algorithms
G. Santafe, I. Inza, and J. A. Lozano · 2015
Earlier work this paper cites.
Hidden technical debt in machine learning systems
D. Sculley, G. Holt, D. Golovin, E. Davydov, T. Phillips, D. Ebner, V. Chaudhary, M. Young, J.-F. Crespo, and D. Dennison · 2015
Earlier work this paper cites.
Should we really use post-hoc tests based on mean-ranks?
A. Benavoli, G. Corani, and F. Mangili · 2016
Earlier work this paper cites.
How to develop, validate, and compare clinical prediction models involving radiological parameters: study design and statistical methods
K. Han, K. Song, and B. W. Choi · 2016
Earlier work this paper cites.
Prediction models need appropriate internal, internal–external, and external validation
E. W. Steyerberg and F. E. Harrell · 2016
Earlier work this paper cites.
Single subject prediction of brain disorders in neuroimaging: Promises and pitfalls
M. R. Arbabshirani, S. Plis, J. Sui, and V. D. Calhoun · 2017
Earlier work this paper cites.
Confidence curves: an alternative to null hypothesis significance testing for the comparison of classifiers
D. Berrar · 2017
Earlier work this paper cites.
Machine learning and microsimulation techniques on the prognosis of dementia: A systematic literature review
A. L. Dallora, S. Eivazzadeh, E. Mendes, J. Berglund, and P. Anderberg · 2017
Earlier work this paper cites.
Ethical and legal implications of the methodological crisis in neuroimaging
P. Kellmeyer · 2017
Earlier work this paper cites.
A survey on deep learning in medical image analysis
G. Litjens, T. Kooi, B. E. Bejnordi, A. A. A. Setio, F. Ciompi, M. Ghafoorian, J. A. van der Laak, B. Van Ginneken, and C. I. Sánchez · 2017
Earlier work this paper cites.
The need to approximate the use-case in clinical machine learning
S. Saeb, L. Lonini, A. Jayaraman, D. C. Mohr, and K. P. Kording · 2017
Earlier work this paper cites.
Building better biomarkers: brain models in translational neuroimaging
C.-W. Woo, L. J. Chang, M. A. Lindquist, and T. D. Wager · 2017
Earlier work this paper cites.
How good is my test data? introducing safety analysis for computer vision
O. Zendel, M. Murschitz, M. Humenberger, and W. Herzner · 2017
Earlier work this paper cites.
Learning to unlearn: building immunity to dataset bias in medical imaging studies
A. Ashraf, S. Khan, N. Bhagwat, M. Chakravarty, and B. Taati · 2018
Earlier work this paper cites.
Negative results in computer vision: A perspective
A. Borji · 2018
Cited alongside, same era.
Datasheets for datasets
T. Gebru, J. Morgenstern, B. Vecchione, J. W. Vaughan, H. M. Wallach, H. D. III, and K. Crawford · 2018
Cited alongside, same era.
Statistical rituals: The replication delusion and how we got there
G. Gigerenzer · 2018
Cited alongside, same era.
State of the art: Reproducibility in artificial intelligence
O. E. Gundersen and S. Kjensmo · 2018
Cited alongside, same era.
Dimensions metrics API reference & getting started
A. Mori and M. Taylor · 2018
Cited alongside, same era.
Realistic evaluation of semi-supervised learning algorithms
A. Oliver, A. Odena, C. Raffel, E. D. Cubuk, and I. J. Goodfellow · 2018
Cited alongside, same era.
Risk of training diagnostic algorithms on data with demographic bias
S. Abbasi-Sureshjani, R. Raumanns, B. E. Michels, G. Schouten, and V. Cheplygina · 2020
Later among the works it cites.
Predicting the progression of mild cognitive impairment using machine learning: a systematic, quantitative and critical review
M. Ansart, S. Epelbaum, G. Bassignana, A. Bône, S. Bottani, T. Cattai, R. Couronne, J. Faouzi, I. Koval, M. Louis, et al · 2020
Later among the works it cites.
Evaluating progress on machine learning for longitudinal electronic healthcare data
D. Bellamy, L. Celi, and A. L. Beam · 2020
Later among the works it cites.
L. Beyer, O. J. Hénaff, A. Kolesnikov, X. Zhai, and A. v. d. Oord · 2020
Later among the works it cites.
Threats of a replication crisis in empirical computer science
A. Cockburn, P. Dragicevic, L. Besançon, and C. Gutwin · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Methodologic guide for evaluating clinical performance and effect of artificial intelligence technology for medical diagnosis and prediction
S. H. Park and K. Han · 2018
Cited alongside, same era.
Failing loudly: an empirical study of methods for detecting dataset shift
S. Rabanser, S. Günnemann, and Z. C. Lipton · 2018
Cited alongside, same era.
Improving the robustness of convolutional networks to appearance variability in biomedical images
T. Tasdizen, M. Sajjadi, M. Javanmardi, and N. Ramesh · 2018
Cited alongside, same era.
A practical taxonomy of reproducibility for machine learning research
R. Tatman, J. VanderPlas, and S. Dane · 2018
Cited alongside, same era.
Cross-validation failure: small sample sizes lead to large error bars
G. Varoquaux · 2018
Cited alongside, same era.
M. Voets, K. Møllersen, and L. A. Bongo · 2018
Cited alongside, same era.
Later among the works it cites.
Towards the systematic reporting of the energy and carbon footprints of machine learning
P. Henderson, J. Hu, J. Romoff, E. Brunskill, D. Jurafsky, and J. Pineau · 2020
Later among the works it cites.
I tried a bunch of things: The dangers of unexpected overfitting in classification of brain data
M. Hosseini, M. Powell, J. Collins, C. Callahan-Flintoft, W. Jones, H. Bowman, and B. Wyble · 2020
Later among the works it cites.
Best practices for authors of healthcare-related artificial intelligence manuscripts
S. Kakarmath, A. Esteva, R. Arnaout, H. Harvey, S. Kumar, E. Muse, F. Dong, L. Wedlund, and J. Kvedar · 2020
Later among the works it cites.
Gender imbalance in medical imaging datasets produces biased classifiers for computer-aided diagnosis
A. J. Larrazabal, N. Nieto, V. Peterson, D. H. Milone, and E. Ferrante · 2020
Later among the works it cites.
A metric learning reality check
K. Musgrave, S. Belongie, and S.-N. Lim · 2020
Later among the works it cites.
Minimum information about clinical artificial intelligence modeling: the MI-CLAIM checklist
B. Norgeot, G. Quer, B. K. Beaulieu-Jones, A. Torkamani, R. Dias, M. Gianfrancesco, R. Arnaout, I. S. Kohane, S. Saria, E. Topol, et al · 2020
Later among the works it cites.
Exploring large-scale public medical image datasets
L. Oakden-Rayner · 2020
Later among the works it cites.
Hidden stratification causes clinically meaningful failures in machine learning for medical imaging
L. Oakden-Rayner, J. Dunnmon, G. Carneiro, and C. Ré · 2020
Later among the works it cites.
A survey of crowdsourcing in medical image analysis
S. N. Ørting, A. Doyle, A. van Hilten, M. Hirth, O. Inel, C. R. Madan, P. Mavridis, H. Spiers, and V. Cheplygina · 2020
Later among the works it cites.
Establishment of best practices for evidence for prediction: a review
R. A. Poldrack, G. Huckins, and G. Varoquaux · 2020
Later among the works it cites.
What your radiologist might be missing: using machine learning to identify mislabeled instances of X-ray images
T. Rädsch, S. Eckhardt, F. Leiser, K. D. Pandl, S. Thiebes, and A. Sunyaev · 2020
Later among the works it cites.
Sample size determination for biomedical big data with limited labels
A. N. Richter and T. M. Khoshgoftaar · 2020
Later among the works it cites.
A path for translation of machine learning products into healthcare delivery
M. P. Sendak, J. D’Arcy, S. Kashyap, M. Gao, M. Nichols, K. Corey, W. Ratliff, and S. Balu · 2020
Later among the works it cites.
Evaluating machine accuracy on imagenet
V. Shankar, R. Roelofs, H. Mania, A. Fang, B. Recht, and L. Schmidt · 2020
Later among the works it cites.
Sample size evolution in neuroimaging research: an evaluation of highly-cited studies (1990-2012) and of latest practices (2017-2018) in high-impact journals
D. Szucs and J. P. Ioannidis · 2020
Later among the works it cites.
On the value of out-of-distribution testing: an example of Goodhart’s Law
D. Teney, K. Kafle, R. Shrestha, E. Abbasnejad, C. Kanan, and A. v. d. Hengel · 2020
Later among the works it cites.
The problem with metrics is a fundamental problem for AI
R. Thomas and D. Uminsky · 2020
Later among the works it cites.
Meta-research: Dataset decay and the problem of sequential analyses on open datasets
W. H. Thompson, J. Wright, P. G. Bissett, and R. A. Poldrack · 2020
Later among the works it cites.
Convolutional neural networks for classification of Alzheimer’s disease: overview and reproducible evaluation
J. Wen, E. Thibeau-Sutre, M. Diaz-Melo, J. Samper-González, A. Routier, S. Bottani, D. Dormont, S. Durrleman, N. Burgos, O. Colliot, et al · 2020
Later among the works it cites.
#bropenscience is broken science
K. Whitaker and O. Guest · 2020
Later among the works it cites.
Time to reality check the promises of machine learning-powered precision medicine
J. Wilkinson, K. F. Arnold, E. J. Murray, M. van Smeden, K. Carr, R. Sippy, M. de Kamps, A. Beam, S. Konigorski, C. Lippert, et al · 2020
Later among the works it cites.
Preparing medical imaging data for machine learning
M. J. Willemink, W. A. Koszek, C. Hardell, J. Wu, D. Fleischmann, H. Harvey, L. R. Folio, R. M. Summers, D. L. Rubin, and M. P. Lungren · 2020
Later among the works it cites.
A review of deep learning in medical imaging: Image traits, technology trends, case studies with progress highlights, and future promises
S. K. Zhou, H. Greenspan, C. Davatzikos, J. S. Duncan, B. van Ginneken, A. Madabhushi, J. L. Prince, D. Rueckert, and R. M. Summers · 2020
Later among the works it cites.
Accounting for variance in machine learning benchmarks
X. Bouthillier, P. Delaunay, M. Bronzi, A. Trofimov, B. Nichyporuk, J. Szeto, N. Mohammadi Sepahvand, E. Raff, K. Madan, V. Voleti, S. E. Kahou, V. Michalski, T. Arbel, C. Pal, G. Varoquaux, and P. Vincent · 2021
Closest in time.
Common pitfalls and recommendations for using machine learning to detect and prognosticate for COVID-19 using chest radiographs and CT scans
M. Roberts, D. Driggs, M. Thorpe, J. Gilbey, M. Yeung, S. Ursprung, A. I. Aviles-Rivero, C. Etmann, C. McCague, L. Beer, et al · 2021
Closest in time.
Detect and correct bias in multi-site neuroimaging datasets
C. Wachinger, A. Rieckmann, S. Pölsterl, A. D. N. Initiative, et al · 2021
Closest in time.