Fetching the paper…
Reading the bibliography…
While the importance of automatic image analysis is continuously increasing, recent meta-research revealed major flaws with respect to algorithm validation.
The distribution of the flora in the alpine zone. 1
Paul Jaccard · 1912
Earlier work this paper cites.
Measures of the amount of ecologic association between species
Lee R Dice · 1945
Earlier work this paper cites.
Verification of forecasts expressed in terms of probability
Glenn W Brier et al · 1950
Earlier work this paper cites.
The hungarian method for the assignment problem
Harold W Kuhn · 1955
Earlier work this paper cites.
Sequential operations in digital picture processing
Azriel Rosenfeld and John L Pfaltz · 1966
Earlier work this paper cites.
Delphi process: a methodology used for the elicitation of opinions of experts
Bernice B Brown · 1968
Earlier work this paper cites.
Comparing the areas under two or more correlated receiver operating characteristic curves: a nonparametric approach
Elizabeth R DeLong, David M DeLong, and Daniel L Clarke-Pearson · 1988
Earlier work this paper cites.
Interjudge agreement and the maximum value of kappa
Uchila N Umesh, Robert A Peterson, and Matthew H Sauber · 1989
Earlier work this paper cites.
Comparing images using the hausdorff distance
Daniel P Huttenlocher, Gregory A. Klanderman, and William J Rucklidge · 1993
Earlier work this paper cites.
Likelihood ratios: a real improvement for clinical decision making?
Bruno Dujardin, Jef Van den Ende, Alfons Van Gompel, Jean-Pierre Unger, and Patrick Van der Stuyft · 1994
Earlier work this paper cites.
Defining the tumour and target volumes for radiotherapy
Neil G Burnet, Simon J Thomas, Kate E Burton, and Sarah J Jefferies · 2004
Earlier work this paper cites.
Video object relevance metrics for overall segmentation quality evaluation
Paulo Correia and Fernando Pereira · 2006
Earlier work this paper cites.
The relationship between precision-recall and roc curves
Jesse Davis and Mark Goadrich · 2006
Earlier work this paper cites.
Use and misuse of the receiver operating characteristic curve in risk prediction
Nancy R Cook · 2007
Earlier work this paper cites.
Strictly proper scoring rules, prediction, and estimation
Tilmann Gneiting and Adrian E Raftery · 2007
Earlier work this paper cites.
3d segmentation in the clinic: A grand challenge
Bram Van Ginneken, Tobias Heimann, and Martin Styner · 2007
Earlier work this paper cites.
Fuzzy pulmonary vessel segmentation in contrast enhanced ct data
Jens N Kaftan, Atilla P Kiraly, Annemarie Bakai, Marco Das, Carol L Novak, and Til Aach · 2008
Earlier work this paper cites.
Measures of diagnostic accuracy: basic definitions
Ana-Maria Šimundić · 2009
Earlier work this paper cites.
Classification of imbalanced data: a review
Yanmin Sun, Andrew K. C. Wong, and Mohamed S. Kamel · 2009
Earlier work this paper cites.
Assessing the performance of prediction models: a framework for some traditional and novel measures
Ewout W Steyerberg, Andrew J Vickers, Nancy R Cook, Thomas Gerds, Mithat Gonen, Nancy Obuchowski, Michael J Pencina, and Michael W Kattan · 2010
Earlier work this paper cites.
Comparing and combining algorithms for computer-aided detection of pulmonary nodules in computed tomography scans: the anode09 study
Bram Van Ginneken, Samuel G Armato III, Bartjan de Hoop, Saskia van Amelsvoort-van de Vorst, Thomas Duindam, Meindert Niemeijer, Keelin Murphy, Arnold Schilham, Alessandra Retico, Maria Evelina Fantacci, et al · 2010
Earlier work this paper cites.
A function for quality evaluation of retinal vessel segmentations
Manuel Emilio Gegúndez-Arias, Arturo Aquino, José Manuel Bravo, and Diego Marín · 2011
Earlier work this paper cites.
Discriminative segmentation-based evaluation through shape dissimilarity
Ender Konukoglu, Ben Glocker, Dong Hye Ye, Antonio Criminisi, and Kilian M Pohl · 2012
Earlier work this paper cites.
On use of partial area under the roc curve for evaluation of diagnostic performance
Hua Ma, Andriy I Bandos, Howard E Rockette, and David Gur · 2013
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
How to evaluate foreground maps?
Ran Margolin, Lihi Zelnik-Manor, and Ayellet Tal · 2014
Earlier work this paper cites.
The multimodal brain tumor image segmentation benchmark (brats)
Bjoern H Menze, Andras Jakab, Stefan Bauer, Jayashree Kalpathy-Cramer, Keyvan Farahani, Justin Kirby, Yuliya Burren, Nicole Porz, Johannes Slotboom, Roland Wiest, et al · 2014
Earlier work this paper cites.
A formal method for selecting evaluation metrics for image segmentation
Abdel Aziz Taha, Allan Hanbury, and Oscar A Jimenez del Toro · 2014
Earlier work this paper cites.
Is the area under an roc curve a valid measure of the performance of a screening or diagnostic test?
NJ Wald and JP Bestwick · 2014
Earlier work this paper cites.
Performance evaluation of image segmentation algorithms on microscopic image data
Miroslav Beneš and Barbara Zitová · 2015
Earlier work this paper cites.
The cityscapes dataset
Marius Cordts, Mohamed Omran, Sebastian Ramos, Timo Scharwächter, Markus Enzweiler, Rodrigo Benenson, Uwe Franke, Stefan Roth, and Bernt Schiele · 2015
Earlier work this paper cites.
The pascal visual object classes challenge: A retrospective
Mark Everingham, SM Ali Eslami, Luc Van Gool, Christopher KI Williams, John Winn, and Andrew Zisserman · 2015
Cited alongside, same era.
The hci stereo metrics: Geometry-aware performance analysis of stereo algorithms
Katrin Honauer, Lena Maier-Hein, and Daniel Kondermann · 2015
Cited alongside, same era.
A review on evaluation metrics for data classification evaluations
Mohammad Hossin and Md Nasir Sulaiman · 2015
Cited alongside, same era.
Cell tracking accuracy measurement based on comparison of acyclic oriented graphs
Pavel Matula, Martin Maška, Dmitry V Sorokin, Petr Matula, Carlos Ortiz-de Solórzano, and Michal Kozubek · 2015
Cited alongside, same era.
Diagnostic tests: how to estimate the positive predictive value
Annette M Molinaro · 2015
Cited alongside, same era.
Obtaining well calibrated probabilities using bayesian binning
Evaluating white matter lesion segmentations with refined sørensen-dice analysis
Aaron Carass, Snehashis Roy, Adrian Gherman, Jacob C Reinhold, Andrew Jesson, Tal Arbel, Oskar Maier, Heinz Handels, Mohsen Ghafoorian, Bram Platel, et al · 2020
Later among the works it cites.
The advantages of the matthews correlation coefficient (mcc) over f1 score and accuracy in binary classification evaluation
Davide Chicco and Giuseppe Jurman · 2020
Later among the works it cites.
Metrics for multi-class classification: an overview
Margherita Grandini, Enrico Bagli, and Giorgio Visani · 2020
Later among the works it cites.
Patchperpix for instance segmentation
Peter Hirsch, Lisa Mais, and Dagmar Kainmueller · 2020
Later among the works it cites.
Retina u-net: Embarrassingly simple exploitation of segmentation supervision for medical object detection
Paul F Jaeger, Simon AA Kohl, Sebastian Bickelhaupt, Fabian Isensee, Tristan Anselm Kuder, Heinz-Peter Schlemmer, and Klaus H Maier-Hein · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mahdi Pakdaman Naeini, Gregory Cooper, and Milos Hauskrecht · 2015
Cited alongside, same era.
Metrics for evaluating 3d medical image segmentation: analysis, selection, and tool
Abdel Aziz Taha and Allan Hanbury · 2015
Cited alongside, same era.
Net benefit approaches to the evaluation of prediction models, molecular markers, and diagnostic tests
Andrew J Vickers, Ben Van Calster, and Ewout W Steyerberg · 2016
Cited alongside, same era.
Deep watershed transform for instance segmentation
Min Bai and Raquel Urtasun · 2017
Cited alongside, same era.
Fairness in machine learning
Solon Barocas, Moritz Hardt, and Arvind Narayanan · 2017
Cited alongside, same era.
Semantic instance segmentation with a discriminative loss function
Bert De Brabandere, Davy Neven, and Luc Van Gool · 2017
Cited alongside, same era.
Generalised dice overlap as a deep learning loss function for highly unbalanced segmentations
Carole H Sudre, Wenqi Li, Tom Vercauteren, Sebastien Ourselin, and M Jorge Cardoso · 2017
Cited alongside, same era.
Later among the works it cites.
Challenges and Opportunities of End-to-End Learning in Medical Image Classification
Paul Ferdinand Jäger · 2020
Later among the works it cites.
Instance segmentation of biological images using harmonic embeddings
Victor Kulikov and Victor Lempitsky · 2020
Later among the works it cites.
Anatomically consistent cnn-based segmentation of organs-at-risk in cranial radiotherapy
Pawel Mlynarski, Hervé Delingette, Hamza Alghamdi, Pierre-Yves Bondiau, and Nicholas Ayache · 2020
Later among the works it cites.
Problems and opportunities in training deep learning software systems: An analysis of variance
Hung Viet Pham, Shangshu Qian, Jiannan Wang, Thibaud Lutellier, Jonathan Rosenthal, Lin Tan, Yaoliang Yu, and Nachiappan Nagappan · 2020
Later among the works it cites.
Evaluation: from precision, recall and f-measure to roc, informedness, markedness and correlation
David MW Powers · 2020
Later among the works it cites.
Diagnostic testing accuracy: Sensitivity, specificity, predictive values and likelihood ratios
Jacob Shreffler and Martin R Huecker · 2020
Later among the works it cites.
Evaluation of measures for assessing time-saving of automatic organ-at-risk segmentation in radiotherapy
Femke Vaassen, Colien Hazelaar, Ana Vaniqui, Mark Gooding, Brent van der Heyden, Richard Canters, and Wouter van Elmpt · 2020
Later among the works it cites.
Scipy 1.0: fundamental algorithms for scientific computing in python
Pauli Virtanen, Ralf Gommers, Travis E Oliphant, Matt Haberland, Tyler Reddy, David Cournapeau, Evgeni Burovski, Pearu Peterson, Warren Weckesser, Jonathan Bright, et al · 2020
Later among the works it cites.
Deep semantic segmentation of natural and medical images: a review
Saeid Asgari Taghanaki, Kumar Abhishek, Joseph Paul Cohen, Julien Cohen-Adad, and Ghassan Hamarneh · 2021
Closest in time.
Boundary iou: Improving object-centric image segmentation evaluation
Bowen Cheng, Ross Girshick, Piotr Dollár, Alexander C Berg, and Alexander Kirillov · 2021
Closest in time.
The matthews correlation coefficient (mcc) is more reliable than balanced accuracy, bookmaker informedness, and markedness in two-class confusion matrix evaluation
Davide Chicco, Niklas Tötsch, and Giuseppe Jurman · 2021
Closest in time.
On evaluation metrics for medical applications of artificial intelligence
Steven Hicks, Inga Strüke, Vajira Thambawita, Malek Hammou, Pål Halvorsen, Michael Riegler, and Sravanthi Parasa · 2021
Closest in time.
Designing deep learning studies in cancer diagnostics
Andreas Kleppe, Ole-Johan Skrede, Sepp De Raedt, Knut Liestøl, David J Kerr, and Håvard E Danielsen · 2021
Closest in time.
Florian Kofler, Ivan Ezhov, Fabian Isensee, Christoph Berger, Maximilian Korner, Johannes Paetzold, Hongwei Li, Suprosanna Shit, Richard McKinley, Spyridon Bakas, et al · 2021
Closest in time.
Comparison of metrics for the evaluation of medical segmentations using prostate mri dataset
Ying-Hwey Nai, Bernice W Teo, Nadya L Tan, Sophie O’Doherty, Mary C Stephenson, Yee Liang Thian, Edmund Chiong, and Anthonin Reilhac · 2021
Closest in time.
Clinically applicable segmentation of head and neck anatomy for radiotherapy: deep learning algorithm development and validation study
Stanislav Nikolov, Sam Blackwell, Alexei Zverovitch, Ruheena Mendes, Michelle Livne, Jeffrey De Fauw, Yojan Patel, Clemens Meyer, Harry Askham, Bernadino Romera-Paredes, et al · 2021
Closest in time.
Comparative validation of multi-instance instrument segmentation in endoscopy: results of the robust-mis 2019 challenge
Tobias Roß, Annika Reinke, Peter M Full, Martin Wagner, Hannes Kenngott, Martin Apitz, Hellena Hempe, Diana Mindroc-Filimon, Patrick Scholz, Thuy Nuong Tran, et al · 2021
Closest in time.
Anindo Saha, Joeran Bosma, Jasper Linmans, Matin Hosseinzadeh, and Henkjan Huisman · 2021
Closest in time.
cldice-a novel topology-preserving loss function for tubular structure segmentation
Suprosanna Shit, Johannes C Paetzold, Anjany Sekuboyina, Ivan Ezhov, Alexander Unger, Andrey Zhylka, Josien PW Pluim, Ulrich Bauer, and Bjoern H Menze · 2021
Closest in time.
Methods and open-source toolkit for analyzing and visualizing challenge results
Manuel Wiesenfarth, Annika Reinke, Bennett A Landman, Matthias Eisenmann, Laura Aguilera Saiz, M Jorge Cardoso, Lena Maier-Hein, and Annette Kopp-Schneider · 2021
Closest in time.
Aston Zhang, Zachary C Lipton, Mu Li, and Alexander J Smola · 2021
Closest in time.
Multicenter comparison of measures for quantitative evaluation of contouring in radiotherapy
Mark J Gooding, Djamal Boukerroui, Eliana Vasquez Osorio, René Monshouwer, and Ellen Brunenberg · 2022
Closest in time.
Sebastian Gruber and Florian Buettner · 2022
Closest in time.
Area under the curve may hide poor generalisation to external datasets
A Kleppe · 2022
Closest in time.
Metrics reloaded: Pitfalls and recommendations for image analysis validation
Lena Maier-Hein, Annika Reinke, Evangelia Christodoulou, Ben Glocker, Patrick Godau, Fabian Isensee, Jens Kleesiek, Michal Kozubek, Mauricio Reyes, Michael A Riegler, et al · 2022
Closest in time.