Fetching the paper…
Reading the bibliography…
Machine learning models that achieve high overall accuracy often make systematic errors on important subsets (or slices) of data.
On Completeness-aware Concept-Based Explanations in Deep Neural Networks
Chih-Kuan Yeh, Been Kim, Sercan O Arik, Chun-Liang Li, Tomas Pfister, and Pradeep Ravikumar · 1910
Earlier work this paper cites.
WordNet: An Electronic Lexical Database
Christiane Fellbaum · 1998
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Contrastive learning of medical visual representations from paired images and text
Yuhao Zhang, Hang Jiang, Yasuhide Miura, Christopher D. Manning, and Curtis P. Langlotz · 2010
Earlier work this paper cites.
Novel dataset for fine-grained image categorization
Aditya Khosla, Nityananda Jayadevaprakash, Bangpeng Yao, and Li Fei-Fei · 2011
Earlier work this paper cites.
Svm based gender classification using iris images
Atul Bansal, Ravinder Agarwal, and R K Sharma · 2012
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Deep learning face attributes in the wild
Ziwei Liu, Ping Luo, Xiaogang Wang, and Xiaoou Tang · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
”why should i trust you?”: Explaining the predictions of any classifier
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin · 2016
Earlier work this paper cites.
Gender classification from the same iris code used for recognition
Juan E Tapia, Claudio A Perez, and Kevin W Bowyer · 2016
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization
Diederik P Kingma and Jimmy Ba · 2017
Earlier work this paper cites.
Gender-from-iris or gender-from-mascara?
Andrey Kuehlkamp, Benedict Becker, and Kevin Bowyer · 2017
Earlier work this paper cites.
Grad-cam: Visual explanations from deep networks via gradient-based localization
Ramprasaath R. Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Batra · 2017
Earlier work this paper cites.
Sanity Checks for Saliency Maps
Julius Adebayo, Justin Gilmer, Michael Muelly, Ian Goodfellow, Moritz Hardt, and Been Kim · 2018
Earlier work this paper cites.
Gender shades: Intersectional accuracy disparities in commercial gender classification
Joy Buolamwini and Timnit Gebru · 2018
Earlier work this paper cites.
BERT: pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Multiaccuracy: Black-Box Post-Processing for Fairness in Classification
Michael P Kim, Amirata Ghorbani, and James Zou · 2018
Earlier work this paper cites.
Variable generalization performance of a deep learning model to detect pneumonia in chest radiographs: a cross-sectional study
John R Zech, Marcus A Badgeley, Manway Liu, Anthony B Costa, Joseph J Titano, and Eric Karl Oermann · 2018
Cited alongside, same era.
Mitigating Unwanted Biases with Adversarial Learning
Brian Hu Zhang, Blake Lemoine, and Margaret Mitchell · 2018
Cited alongside, same era.
Deep learning predicts hip fracture using confounding patient and healthcare variables
Marcus A Badgeley, John R Zech, Luke Oakden-Rayner, Benjamin S Glicksberg, Manway Liu, William Gale, Michael V McConnell, Bethany Percha, Thomas M Snyder, and Joel T Dudley · 2019
Cited alongside, same era.
(de) constructing bias on skin lesion datasets
Alceu Bissoto, Michel Fornaciali, Eduardo Valle, and Sandra Avila · 2019
Cited alongside, same era.
Slice finder: Automated data slicing for model validation
Yeounoh Chung, Tim Kraska, Neoklis Polyzotis, Ki Hyun Tae, and Steven Euijong Whang · 2019
Cited alongside, same era.
Weak supervision as an efficient approach for automated seizure detection in electroencephalography
Khaled Saab, Jared Dunnmon, Christopher Ré, Daniel Rubin, and Christopher Lee-Messer · 2020
Later among the works it cites.
Distributionally Robust Neural Networks for Group Shifts: On the Importance of Regularization for Worst-Case Generalization
Shiori Sagawa, Pang Wei Koh, Tatsunori B Hashimoto, and Percy Liang · 2020
Later among the works it cites.
Combining automatic labelers and expert annotations for accurate radiology report labeling using BERT
Akshay Smit, Saahil Jain, Pranav Rajpurkar, Anuj Pareek, Andrew Ng, and Matthew Lungren · 2020
Later among the works it cites.
No subclass left behind: Fine-grained robustness in coarse-grained classification problems
Nimit Sohoni, Jared Dunnmon, Geoffrey Angus, Albert Gu, and Christopher Ré · 2020
Later among the works it cites.
On completeness-aware concept-based explanations in deep neural networks
Chih-Kuan Yeh, Been Kim, Sercan Arik, Chun-Liang Li, Tomas Pfister, and Pradeep Ravikumar · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Does object recognition work for everyone?
Terrance de Vries, Ishan Misra, Changhan Wang, and Laurens van der Maaten · 2019
Cited alongside, same era.
Weakly supervised classification of aortic valve malformations using unlabeled cardiac MRI sequences
Jason A Fries, Paroma Varma, Vincent S Chen, Ke Xiao, Heliodoro Tejeda, Priyanka Saha, Jared Dunnmon, Henry Chubb, Shiraz Maskatia, Madalina Fiterau, Scott Delp, Euan Ashley, Christopher Ré, and James R Priest · 2019
Cited alongside, same era.
Benchmarking neural network robustness to common corruptions and perturbations
Dan Hendrycks and Thomas Dietterich · 2019
Cited alongside, same era.
Mimic-cxr, a de-identified publicly available database of chest radiographs with free-text reports
Alistair E W Johnson, Tom J Pollard, Seth J Berkowitz, Nathaniel R Greenbaum, Matthew P Lungren, Chih-ying Deng, Roger G Mark, and Steven Horng · 2019
Cited alongside, same era.
Big transfer (bit): General visual representation learning
Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai, Joan Puigcerver, Jessica Yung, Sylvain Gelly, and Neil Houlsby · 2019
Cited alongside, same era.
Hidden stratification causes clinically meaningful failures in machine learning for medical imaging
Luke Oakden-Rayner, Jared Dunnmon, Gustavo Carneiro, and Christopher Ré · 2019
Cited alongside, same era.
Chrononet: a deep recurrent neural network for abnormal eeg identification
Subhrajit Roy, Isabell Kiral-Kornek, and Stefan Harrer · 2019
Cited alongside, same era.
Later among the works it cites.
Mandoline: Model evaluation under distribution shift
Mayee Chen, Karan Goel, Nimit S Sohoni, Fait Poms, Kayvon Fatahalian, and Christopher Re · 2021
Later among the works it cites.
AI for radiographic COVID-19 detection selects shortcuts over signal
Alex J DeGrave, Jose Janizek, and Su-In Lee · 2021
Later among the works it cites.
The spotlight: A general method for discovering systematic errors in deep learning models, 2021
Greg d’Eon, Jason d’Eon, James R. Wright, and Kevin Leyton-Brown · 2021
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale, 2021
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Later among the works it cites.
Robustness Gym: Unifying the NLP Evaluation Landscape
Karan Goel, Nazneen Rajani, Jesse Vig, Samson Tan, Jason Wu, Stephan Zheng, Caiming Xiong, Mohit Bansal, and Christopher Ré · 2021
Later among the works it cites.
Towards non-i.i.d. image classification: A dataset and baselines
Yue He, Zheyan Shen, and Peng Cui · 2021
Later among the works it cites.
WILDS: A Benchmark of in-the-Wild Distribution Shifts
Pang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie, Marvin Zhang, Akshay Balsubramani, Weihua Hu, Michihiro Yasunaga, Richard Lanas Phillips, Irena Gao, Tony Lee, Etienne David, Ian Stavness, Wei Guo, Berton A Earnshaw, Imran S Haque, Sara Beery, Jure Leskovec, Anshul Kundaje, Emma Pierson, Sergey Levine, Chelsea Finn, and Percy Liang · 2021
Later among the works it cites.
Metadataset: A dataset of datasets for evaluating distribution shifts and training conflicts
Weixin Liang and James Zou · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever · 2021
Later among the works it cites.
Sliceline: Fast, linear-algebra-based slice finding for ml model debugging
Svetlana Sagadeeva and Matthias Boehm · 2021
Later among the works it cites.
Understanding failures of deep networks via robust feature extraction
Sahil Singla, Besmira Nushi, Shital Shah, Ece Kamar, and Eric Horvitz · 2021
Later among the works it cites.
Cross-domain data integration for named entity disambiguation in biomedical text
Maya Varma, Laurel Orr, Sen Wu, Megan Leszczynski, Xiao Ling, and Christopher Ré · 2021
Later among the works it cites.
Vilmedic: a framework for research at the intersection of vision and language in medical ai
Jean-Benoit Delbrouck, Khaled Saab, Maya Varma, Sabri Eyuboglu, Jared A. Dunnmon, Pierre Chambon, Juan Manuel Zambrano, Akshay Chaudhari, and Curtis P. Langlotz · 2022
Closest in time.