Fetching the paper…
Reading the bibliography…
Increasing evidence shows that flaws in machine learning (ML) algorithm validation are an underestimated global problem.
The distribution of the flora in the alpine zone. 1
Paul Jaccard · 1912
Earlier work this paper cites.
Measures of the amount of ecologic association between species
Lee R Dice · 1945
Earlier work this paper cites.
Verification of forecasts expressed in terms of probability
Glenn W Brier et al · 1950
Earlier work this paper cites.
The hungarian method for the assignment problem
Harold W Kuhn · 1955
Earlier work this paper cites.
A coefficient of agreement for nominal scales
Jacob Cohen · 1960
Earlier work this paper cites.
Delphi process: a methodology used for the elicitation of opinions of experts
Bernice B Brown · 1968
Earlier work this paper cites.
Objective criteria for the evaluation of clustering methods
William M Rand · 1971
Earlier work this paper cites.
Comparison of the predicted and observed secondary structure of t4 phage lysozyme
Brian W Matthews · 1975
Earlier work this paper cites.
Therapeutic decision making: a cost-benefit analysis
Stephen G Pauker and Jerome P Kassirer · 1975
Earlier work this paper cites.
Information retrieval: theory and practice
C Van Rijsbergen · 1979
Earlier work this paper cites.
The meaning and use of the area under a receiver operating characteristic (roc) curve
James A Hanley and Barbara J McNeil · 1982
Earlier work this paper cites.
The comparison and evaluation of forecasters
Morris H DeGroot and Stephen E Fienberg · 1983
Earlier work this paper cites.
Longitudinal data analysis using generalized linear models
Kung-Yee Liang and Scott L Zeger · 1986
Earlier work this paper cites.
Muc-4 evaluation metrics
Nancy Chinchor · 1992
Earlier work this paper cites.
Comparing images using the hausdorff distance
Daniel P Huttenlocher, Gregory A. Klanderman, and William J Rucklidge · 1993
Earlier work this paper cites.
The Mathematics of Information Coding, Extraction and Distribution , volume 107
George Cybenko, Dianne P O’Leary, and Jorma Rissanen · 1998
Earlier work this paper cites.
Quality management systems: Fundamentals and vocabulary
BSEN ISO 9000 · 2000
Earlier work this paper cites.
Moving beyond sensitivity and specificity: using likelihood ratios to help interpret diagnostic tests
John Attia · 2003
Earlier work this paper cites.
Towards complete and accurate reporting of studies of diagnostic accuracy: the stard initiative
Patrick M Bossuyt, Johannes B Reitsma, David E Bruns, Constantine A Gatsonis, Paul P Glasziou, Les M Irwig, Jeroen G Lijmer, David Moher, Drummond Rennie, Henrica CW De Vet, et al · 2003
Earlier work this paper cites.
Comparing clusterings by the variation of information
Marina Meilă · 2003
Earlier work this paper cites.
Evaluating predictive uncertainty challenge
Joaquin Quinonero-Candela, Carl Edward Rasmussen, Fabian Sinz, Olivier Bousquet, and Bernhard Schölkopf · 2005
Earlier work this paper cites.
Pattern recognition and machine learning , volume 4
Christopher M Bishop and Nasser M Nasrabadi · 2006
Earlier work this paper cites.
Application-independent evaluation of speaker detection
Niko Brümmer and Johan Du Preez · 2006
Earlier work this paper cites.
Video object relevance metrics for overall segmentation quality evaluation
Paulo Correia and Fernando Pereira · 2006
Earlier work this paper cites.
The relationship between precision-recall and roc curves
Jesse Davis and Mark Goadrich · 2006
Earlier work this paper cites.
The 2005 pascal visual object classes challenge
Mark Everingham, Andrew Zisserman, Christopher KI Williams, Luc Van Gool, Moray Allan, Christopher M Bishop, Olivier Chapelle, Navneet Dalal, Thomas Deselaers, Gyuri Dorkó, et al · 2006
Earlier work this paper cites.
Model for defining and reporting reference-based validation protocols in medical image processing
Pierre Jannin, Christophe Grova, and Calvin R Maurer · 2006
Earlier work this paper cites.
Correlated label propagation with application to multi-label learning
Feng Kang, Rong Jin, and Rahul Sukthankar · 2006
Earlier work this paper cites.
Comparing image detection algorithms using resampling
Frank W Samuelson and Nicholas Petrick · 2006
Earlier work this paper cites.
Decision curve analysis: a novel method for evaluating prediction models
Andrew J Vickers and Elena B Elkin · 2006
Earlier work this paper cites.
Strictly proper scoring rules, prediction, and estimation
Tilmann Gneiting and Adrian E Raftery · 2007
Earlier work this paper cites.
An introduction to application-independent evaluation of speaker recognition systems
David A van Leeuwen and Niko Brümmer · 2007
Earlier work this paper cites.
Area under the free-response roc curve (froc) and a related summary index
Andriy I Bandos, Howard E Rockette, Tao Song, and David Gur · 2009
Earlier work this paper cites.
The elements of statistical learning: data mining, inference, and prediction , volume 2
Trevor Hastie, Robert Tibshirani, Jerome H Friedman, and Jerome H Friedman · 2009
Earlier work this paper cites.
The pascal visual object classes (voc) challenge
Mark Everingham, Luc Van Gool, Christopher KI Williams, John Winn, and Andrew Zisserman · 2010
Earlier work this paper cites.
Pattern recognition in histopathological images: An icpr 2010 contest
Metin N Gurcan, Anant Madabhushi, and Nasir Rajpoot · 2010
Earlier work this paper cites.
Consort 2010 statement: updated guidelines for reporting parallel group randomized trials
Kenneth F Schulz, Douglas G Altman, David Moher, and CONSORT Group* · 2010
Earlier work this paper cites.
Assessing the performance of prediction models: a framework for some traditional and novel measures
Ewout W Steyerberg, Andrew J Vickers, Nancy R Cook, Thomas Gerds, Mithat Gonen, Nancy Obuchowski, Michael J Pencina, and Michael W Kattan · 2010
Earlier work this paper cites.
Comparing and combining algorithms for computer-aided detection of pulmonary nodules in computed tomography scans: the anode09 study
Bram Van Ginneken, Samuel G Armato III, Bartjan de Hoop, Saskia van Amelsvoort-van de Vorst, Thomas Duindam, Meindert Niemeijer, Keelin Murphy, Arnold Schilham, Alessandra Retico, Maria Evelina Fantacci, et al · 2010
Earlier work this paper cites.
The lung image database consortium (lidc) and image database resource initiative (idri): a completed reference database of lung nodules on ct scans
Samuel G Armato III, Geoffrey McLennan, Luc Bidaut, Michael F McNitt-Gray, Charles R Meyer, Anthony P Reeves, Binsheng Zhao, Denise R Aberle, Claudia I Henschke, Eric A Hoffman, et al · 2011
Earlier work this paper cites.
Guidelines for reporting reliability and agreement studies (grras) were proposed
Jan Kottner, Laurent Audigé, Stig Brorson, Allan Donner, Byron J Gajewski, Asbjørn Hróbjartsson, Chris Roberts, Mohamed Shoukri, and David L Streiner · 2011
Earlier work this paper cites.
Discriminative segmentation-based evaluation through shape dissimilarity
Ender Konukoglu, Ben Glocker, Dong Hye Ye, Antonio Criminisi, and Kilian M Pohl · 2012
Earlier work this paper cites.
Annotated high-throughput microscopy image sets for validation
Vebjorn Ljosa, Katherine L Sokolnicki, and Anne E Carpenter · 2012
Earlier work this paper cites.
Improved assessment of multiple sclerosis lesion segmentation agreement via detection and outline error estimates
David S Wack, Michael G Dwyer, Niels Bergsland, Carol Di Perri, Laura Ranza, Sara Hussein, Deepa Ramasamy, Guy Poloni, and Robert Zivadinov · 2012
Earlier work this paper cites.
Some paradoxical results for the quadratically weighted kappa
Matthijs J Warrens · 2012
Earlier work this paper cites.
The cancer imaging archive (tcia): maintaining and operating a public information repository
Kenneth Clark, Bruce Vendt, Kirk Smith, John Freymann, Justin Kirby, Paul Koppel, Stephen Moore, Stanley Phillips, David Maffitt, Michael Pringle, et al · 2013
Earlier work this paper cites.
Tractometer: towards validation of tractography pipelines
Marc-Alexandre Côté, Gabriel Girard, Arnaud Boré, Eleftherios Garyfallidis, Jean-Christophe Houde, and Maxime Descoteaux · 2013
Earlier work this paper cites.
Objective comparison of particle tracking methods
Nicolas Chenouard, Ihor Smal, Fabrice De Chaumont, Martin Maška, Ivo F Sbalzarini, Yuanhao Gong, Janick Cardinale, Craig Carthel, Stefano Coraluppi, Mark Winter, et al · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
How to evaluate foreground maps?
Ran Margolin, Lihi Zelnik-Manor, and Ayellet Tal · 2014
Earlier work this paper cites.
A benchmark for comparison of cell tracking algorithms
Martin Maška, Vladimír Ulman, David Svoboda, Pavel Matula, Petr Matula, Cristina Ederra, Ainhoa Urbiola, Tomás España, Subramanian Venkatesan, Deepak MW Balak, et al · 2014
Earlier work this paper cites.
Data from lidc-idri [data set]
S. G. Armato III, G. McLennan, L. Bidaut, M. F. McNitt-Gray, C. R. Meyer, A. P. Reeves, B. Zhao, D. R. Aberle, C. I. Henschke, E. A. Hoffman, et al · 2015
Earlier work this paper cites.
Performance evaluation of image segmentation algorithms on microscopic image data
Miroslav Beneš and Barbara Zitová · 2015
Earlier work this paper cites.
The pascal visual object classes challenge: A retrospective
Mark Everingham, SM Ali Eslami, Luc Van Gool, Christopher KI Williams, John Winn, and Andrew Zisserman · 2015
Earlier work this paper cites.
The hci stereo metrics: Geometry-aware performance analysis of stereo algorithms
Katrin Honauer, Lena Maier-Hein, and Daniel Kondermann · 2015
Earlier work this paper cites.
Cell tracking accuracy measurement based on comparison of acyclic oriented graphs
Pavel Matula, Martin Maška, Dmitry V Sorokin, Petr Matula, Carlos Ortiz-de Solórzano, and Michal Kozubek · 2015
Earlier work this paper cites.
Transparent reporting of a multivariable prediction model for individual prognosis or diagnosis (tripod): explanation and elaboration
Karel GM Moons, Douglas G Altman, Johannes B Reitsma, John PA Ioannidis, Petra Macaskill, Ewout W Steyerberg, Andrew J Vickers, David F Ransohoff, and Gary S Collins · 2015
Earlier work this paper cites.
Obtaining well calibrated probabilities using bayesian binning
Mahdi Pakdaman Naeini, Gregory Cooper, and Milos Hauskrecht · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun · 2015
Cited alongside, same era.
Quantitative evaluation of software packages for single-molecule localization microscopy
Daniel Sage, Hagai Kirshner, Thomas Pengo, Nico Stuurman, Junhong Min, Suliana Manley, and Michael Unser · 2015
Cited alongside, same era.
Metrics for evaluating 3d medical image segmentation: analysis, selection, and tool
Abdel Aziz Taha and Allan Hanbury · 2015
Cited alongside, same era.
A spline-based tool to assess and visualize the calibration of multiclass risk predictions
Kirsten Van Hoorde, Sabine Van Huffel, Dirk Timmerman, Tom Bourne, and Ben Van Calster · 2015
Cited alongside, same era.
An overview of current evaluation methods used in medical image segmentation
Varduhi Yeghiazaryan and Irina Voiculescu · 2015
Cited alongside, same era.
Deep-learning-assisted detection and segmentation of rib fractures from ct scans: Development and validation of fracnet
Liang Jin, Jiancheng Yang, Kaiming Kuang, Bingbing Ni, Yiyi Gao, Yingli Sun, Pan Gao, Weiling Ma, Mingyu Tan, Hui Kang, et al · 2020
Later among the works it cites.
Multivariate confidence calibration for object detection
Fabian Kuppers, Jan Kronenberger, Amirhossein Shantia, and Anselm Haselhoff · 2020
Later among the works it cites.
Patchperpix for instance segmentation
Lisa Mais, Peter Hirsch, and Dagmar Kainmueller · 2020
Later among the works it cites.
Confidence calibration and predictive uncertainty estimation for deep medical image segmentation
Alireza Mehrtash, William M Wells, Clare M Tempany, Purang Abolmaesumi, and Tina Kapur · 2020
Later among the works it cites.
Robust classification of cell cycle phase and biological feature extraction by image-based deep learning
Yukiko Nagao, Mika Sakamoto, Takumi Chinen, Yasushi Okada, and Daisuke Takao · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
External validation of clinical prediction models using big datasets from e-health records or ipd meta-analysis: opportunities and challenges
Richard D Riley, Joie Ensor, Kym IE Snell, Thomas PA Debray, Doug G Altman, Karel GM Moons, and Gary S Collins · 2016
Cited alongside, same era.
A calibration hierarchy for risk models was defined: from utopia to empirical data
Ben Van Calster, Daan Nieboer, Yvonne Vergouwe, Bavo De Cock, Michael J Pencina, and Ewout W Steyerberg · 2016
Cited alongside, same era.
Net benefit approaches to the evaluation of prediction models, molecular markers, and diagnostic tests
Andrew J Vickers, Ben Van Calster, and Ewout W Steyerberg · 2016
Cited alongside, same era.
Sononet: real-time detection and localisation of fetal standard scan planes in freehand ultrasound
Christian F Baumgartner, Konstantinos Kamnitsas, Jacqueline Matthew, Tara P Fletcher, Sandra Smith, Lisa M Koch, Bernhard Kainz, and Daniel Rueckert · 2017
Cited alongside, same era.
Diagnostic assessment of deep learning algorithms for detection of lymph node metastases in women with breast cancer
Babak Ehteshami Bejnordi, Mitko Veta, Paul Johannes Van Diest, Bram Van Ginneken, Nico Karssemeijer, Geert Litjens, Jeroen AWM Van Der Laak, Meyke Hermsen, Quirine F Manson, Maschenka Balkenhol, et al · 2017
Cited alongside, same era.
The Declaration - Montreal Responsible AI, 2017
Université de Montréal · 2017
Cited alongside, same era.
On Calibration of Modern Neural Networks
Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q Weinberger · 2017
Cited alongside, same era.
Tractography reproducibility challenge with empirical data (traced): the 2017 ismrm diffusion study group challenge
Vishwesh Nath, Kurt G Schilling, Prasanna Parvathaneni, Yuankai Huo, Justin A Blaber, Allison E Hainline, Muhamed Barakovic, David Romascano, Jonathan Rafael-Patino, Matteo Frigo, et al · 2020
Later among the works it cites.
Hidden stratification causes clinically meaningful failures in machine learning for medical imaging
Luke Oakden-Rayner, Jared Dunnmon, Gustavo Carneiro, and Christopher Ré · 2020
Later among the works it cites.
Developing specific reporting guidelines for diagnostic accuracy studies assessing AI interventions: The STARD-AI steering group
Viknesh Sounderajah, Hutan Ashrafian, Ravi Aggarwal, Jeffrey De Fauw, Alastair K Denniston, Felix Greaves, Alan Karthikesalingam, Dominic King, Xiaoxuan Liu, Sheraz R Markar, Matthew D F McInnes, Trishan Panch, Jonathan Pearson-Stuttard, Daniel S W Ting, Robert M Golub, David Moher, Patrick M Bossuyt, and Ara Darzi · 2020
Later among the works it cites.
Classification assessment methods
Alaa Tharwat · 2020
Later among the works it cites.
Evaluation of measures for assessing time-saving of automatic organ-at-risk segmentation in radiotherapy
Femke Vaassen, Colien Hazelaar, Ana Vaniqui, Mark Gooding, Brent van der Heyden, Richard Canters, and Wouter van Elmpt · 2020
Later among the works it cites.
Non-parametric calibration for classification
Jonathan Wenger, Hedvig Kjellström, and Rudolph Triebel · 2020
Later among the works it cites.
Carbontracker: Tracking and predicting the carbon footprint of training deep learning models
Lasse F Wolff Anthony, Benjamin Kanding, and Raghavendra Selvan · 2020
Later among the works it cites.
Deepphagy: a deep learning framework for quantitatively measuring autophagy activity in saccharomyces cerevisiae
Ying Zhang, Yubin Xie, Wenzhong Liu, Wankun Deng, Di Peng, Chenwei Wang, Haodong Xu, Chen Ruan, Yongjie Deng, Yaping Guo, et al · 2020
Later among the works it cites.
On the performance of matthews correlation coefficient (mcc) for imbalanced dataset
Qiuming Zhu · 2020
Later among the works it cites.
Deep semantic segmentation of natural and medical images: a review
Saeid Asgari Taghanaki, Kumar Abhishek, Joseph Paul Cohen, Julien Cohen-Adad, and Ghassan Hamarneh · 2021
Later among the works it cites.
The values encoded in machine learning research
Abeba Birhane, Pratyusha Kalluri, Dallas Card, William Agnew, Ravit Dotan, and Michelle Bao · 2021
Later among the works it cites.
Boundary iou: Improving object-centric image segmentation evaluation
Bowen Cheng, Ross Girshick, Piotr Dollár, Alexander C Berg, and Alexander Kirillov · 2021
Later among the works it cites.
Protocol for development of a reporting guideline (tripod-ai) and risk of bias tool (probast-ai) for diagnostic and prognostic prediction model studies based on artificial intelligence
Gary S Collins, Paula Dhiman, Constanza L Andaur Navarro, Jie Ma, Lotty Hooft, Johannes B Reitsma, Patricia Logullo, Andrew L Beam, Lily Peng, Ben Van Calster, et al · 2021
Later among the works it cites.
Qualitative criteria for feasible cranial implant designs
David G Ellis, Carlos M Alvarez, and Michele R Aizenberg · 2021
Later among the works it cites.
Health data poverty: an assailable barrier to equitable digital health care
Hussein Ibrahim, Xiaoxuan Liu, Nevine Zariffa, Andrew D Morris, and Alastair K Denniston · 2021
Later among the works it cites.
Towards responsible research in digital technology for health care
Pierre Jannin · 2021
Later among the works it cites.
Florian Kofler, Ivan Ezhov, Fabian Isensee, Christoph Berger, Maximilian Korner, Johannes Paetzold, Hongwei Li, Suprosanna Shit, Richard McKinley, Spyridon Bakas, et al · 2021
Later among the works it cites.
Green algorithms: quantifying the carbon footprint of computation
Loïc Lannelongue, Jason Grealey, and Michael Inouye · 2021
Later among the works it cites.
Baseline photos and confident annotation improve automated detection of cutaneous graft-versus-host disease
Xiaoqi Liu, Kelsey Parks, Inga Saknite, Tahsin Reasat, Austin D Cronin, Lee E Wheless, Benoit M Dawant, and Eric R Tkaczyk · 2021
Later among the works it cites.
Heidelberg colorectal data set for surgical data science in the sensor operating room
Lena Maier-Hein, Martin Wagner, Tobias Ross, Annika Reinke, Sebastian Bodenstedt, Peter M Full, Hellena Hempe, Diana Mindroc-Filimon, Patrick Scholz, Thuy Nuong Tran, et al · 2021
Later among the works it cites.
Comparison of metrics for the evaluation of medical segmentations using prostate mri dataset
Ying-Hwey Nai, Bernice W Teo, Nadya L Tan, Sophie O’Doherty, Mary C Stephenson, Yee Liang Thian, Edmund Chiong, and Anthonin Reilhac · 2021
Later among the works it cites.
Delphi methodology in healthcare research: how to decide its appropriateness
Prashant Nasa, Ravi Jain, and Deven Juneja · 2021
Later among the works it cites.
Clinically applicable segmentation of head and neck anatomy for radiotherapy: deep learning algorithm development and validation study
Stanislav Nikolov, Sam Blackwell, Alexei Zverovitch, Ruheena Mendes, Michelle Livne, Jeffrey De Fauw, Yojan Patel, Clemens Meyer, Harry Askham, Bernadino Romera-Paredes, et al · 2021
Later among the works it cites.
Carbon emissions and large neural network training
David Patterson, Joseph Gonzalez, Quoc Le, Chen Liang, Lluis-Miquel Munguia, Daniel Rothchild, David So, Maud Texier, and Jeff Dean · 2021
Later among the works it cites.
Common limitations of image processing metrics: A picture story
Annika Reinke, Matthias Eisenmann, Minu D Tizabi, Carole H Sudre, Tim Rädsch, Michela Antonelli, Tal Arbel, Spyridon Bakas, M Jorge Cardoso, Veronika Cheplygina, et al · 2021
Later among the works it cites.
Tobias Roß, Pierangela Bruno, Annika Reinke, Manuel Wiesenfarth, Lisa Koeppel, Peter M Full, Bünyamin Pekdemir, Patrick Godau, Darya Trofimova, Fabian Isensee, et al · 2021
Later among the works it cites.
cldice-a novel topology-preserving loss function for tubular structure segmentation
Suprosanna Shit, Johannes C Paetzold, Anjany Sekuboyina, Ivan Ezhov, Alexander Unger, Andrey Zhylka, Josien PW Pluim, Ulrich Bauer, and Bjoern H Menze · 2021
Later among the works it cites.
Nondeterminism and instability in neural network optimization
Cecilia Summers and Michael J Dinneen · 2021
Later among the works it cites.
Semantic segmentation of human oocyte images using deep neural networks
Anna Targosz, Piotr Przystałka, Ryszard Wiaderkiewicz, and Grzegorz Mrugacz · 2021
Later among the works it cites.
Comparing methods of detecting and segmenting unruptured intracranial aneurysms on tof-mras: The adam challenge
Kimberley M Timmins, Irene C van der Schaaf, Edwin Bennink, Ynte M Ruigrok, Xingle An, Michael Baumgartner, Pascal Bourdon, Riccardo De Feo, Tommaso Di Noto, Florian Dubost, et al · 2021
Later among the works it cites.
Dermoscopedia, 2021
Richard Usatine and Rachel Manci · 2021
Later among the works it cites.
Methods and open-source toolkit for analyzing and visualizing challenge results
Manuel Wiesenfarth, Annika Reinke, Bennett A Landman, Matthias Eisenmann, Laura Aguilera Saiz, M Jorge Cardoso, Lena Maier-Hein, and Annette Kopp-Schneider · 2021
Later among the works it cites.
Multi-output gaussian processes for uncertainty-aware recommender systems
Yinchong Yang and Florian Buettner · 2021
Later among the works it cites.
The medical segmentation decathlon
Michela Antonelli, Annika Reinke, Spyridon Bakas, Keyvan Farahani, Annette Kopp-Schneider, Bennett A Landman, Geert Litjens, Bjoern Menze, Olaf Ronneberger, Ronald M Summers, et al · 2022
Closest in time.
Mitosis domain generalization in histopathology images–the midog challenge
Marc Aubreville, Nikolas Stathonikos, Christof A Bertram, Robert Klopleisch, Natalie ter Hoeve, Francesco Ciompi, Frauke Wilm, Christian Marzahl, Taryn A Donovan, Andreas Maier, et al · 2022
Closest in time.
Uncertainty-informed deep learning models enable high-confidence predictions for digital histopathology
James M Dolezal, Andrew Srisuwananukorn, Dmitry Karpeyev, Siddhi Ramesh, Sara Kochanny, Brittany Cody, Aaron S Mansfield, Sagar Rakshit, Radhika Bansal, Melanie C Bois, et al · 2022
Closest in time.
Analysis and comparison of classification metrics
Luciana Ferrer · 2022
Closest in time.
Better uncertainty calibration via proper scores for classification and beyond
Sebastian Gregor Gruber and Florian Buettner · 2022
Closest in time.
-reliable uncertainty estimation for medical image segmentation
Thierry Judge, Olivier Bernard, Mihaela Porumb, Agisilaos Chartsias, Arian Beqiri, and Pierre-Marc Jodoin · 2022
Closest in time.
blob loss: instance imbalance aware loss functions for semantic segmentation
Florian Kofler, Suprosanna Shit, Ivan Ezhov, Lucas Fidon, Rami Al-Maskari, Hongwei Li, Harsharan Bhatia, Timo Loehr, Marie Piraud, Ali Erturk, et al · 2022
Closest in time.
Technology readiness levels for machine learning systems
Alexander Lavin, Ciarán M Gilligan-Lee, Alessya Visnjic, Siddha Ganju, Dava Newman, Sujoy Ganguly, Danny Lange, Atílím Güneş Baydin, Amit Sharma, Adam Gibson, et al · 2022
Closest in time.
A unifying force for the realization of medical ai
Jochen K Lennerz, Ursula Green, Drew FK Williamson, and Faisal Mahmood · 2022
Closest in time.
Estimating model performance under domain shifts with class-specific confidence scores
Zeju Li, Konstantinos Kamnitsas, Mobarakol Islam, Chen Chen, and Ben Glocker · 2022
Closest in time.
Metrics reloaded: Pitfalls and recommendations for image analysis validation
Lena Maier-Hein, Annika Reinke, Evangelia Christodoulou, Ben Glocker, Patrick Godau, Fabian Isensee, Jens Kleesiek, Michal Kozubek, Mauricio Reyes, Michael A Riegler, et al · 2022
Closest in time.
A research ethics framework for the clinical translation of healthcare machine learning
Melissa D McCradden, James A Anderson, Elizabeth A Stephenson, Erik Drysdale, Lauren Erdman, Anna Goldenberg, and Randi Zlotnik Shaul · 2022
Closest in time.
A searchable image resource of drosophila gal4-driver expression patterns with single neuron resolution
G. Meissner, A. Nern, Z. Dorman, DePasquale G.M., K. Forster, T. Gibney, Hausenfluck J.H., Y. He, N. Iyer, J. Jeter, et al · 2022
Closest in time.
A consistent and differentiable lp canonical calibration error estimator
Teodora Popordanoska, Raphael Sayer, and Matthew B Blaschko · 2022
Closest in time.
The institute for ethical AI & machine learning
The Institute for Ethical Ai and Machine Learning · 2022
Closest in time.
Sources of performance variability in deep learning-based polyp detection
Thuy N Tran, Tim Adler, Amine Yamlahi, Evangelia Christodoulou, Patrick Godau, Annika Reinke, Minu D Tizabi, Peter Sauer, Tillmann Persicke, Jörg G. Albert, and Lena Maier-Hein · 2022
Closest in time.
A call to reflect on evaluation practices for failure detection in image classification
Paul F Jaeger, Carsten T Lüth, Lukas Klein, and Till J Bungert · 2023
Closest in time.
Methods for Clinical Evaluation of Artificial Intelligence Algorithms for Medical Diagnosis
Seong Ho Park, Kyunghwa Han, Hye Young Jang, Ji Eun Park, June-Goo Lee, Dong Wook Kim, and Jaesoon Choi · 2023
Closest in time.
Beyond calibration: estimating the grouping loss of modern neural networks
Alexandre Perez-Lebel, Marine Le Morvan, and Gaël Varoquaux · 2023
Closest in time.
Understanding metric-related pitfalls in image analysis validation
Annika Reinke, Minu D Tizabi, Michael Baumgartner, Matthias Eisenmann, Doreen Heckmann-Nötzel, A Emre Kavur, Tim Rädsch, Carole H Sudre, Laura Acion, Michela Antonelli, et al · 2024
Closest in time.