Fetching the paper…
Reading the bibliography…
Machine learning (ML) models that achieve high average accuracy can still underperform on semantically coherent subsets ("slices") of data.
Fast, Cheap, and Creative: Evaluating Translation Quality Using Amazon’s Mechanical Turk
Callison-Burch, C. 2009 · 2009
Earlier work this paper cites.
Microsoft coco: Common objects in context
Lin, T.-Y.; Maire, M.; Belongie, S.; Hays, J.; Perona, P.; Ramanan, D.; Dollár, P.; and Zitnick, C. L. 2014 · 2014
Earlier work this paper cites.
Program Synthesis using Natural Language
Desai, A.; Gulwani, S.; Hingorani, V.; Jain, N.; Karkare, A.; Marron, M.; R, S.; and Roy, S. 2015 · 2015
Earlier work this paper cites.
MicroTalk: Using Argumentation to Improve Crowdsourcing Accuracy
Drapeau, R.; Chilton, L.; Bragg, J.; and Weld, D. 2016 · 2016
Earlier work this paper cites.
Revolt: Collaborative Crowdsourcing for Labeling Machine Learning Datasets
Chang, J. C.; Amershi, S.; and Kamar, E. 2017 · 2017
Earlier work this paper cites.
The Influence of Personality Traits and Cognitive Load on the Use of Adaptive User Interfaces
Gajos, K. Z.; and Chauncey, K. 2017 · 2017
Earlier work this paper cites.
Gender shades: Intersectional accuracy disparities in commercial gender classification
Buolamwini, J.; and Gebru, T. 2018 · 2018
Earlier work this paper cites.
Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (tcav)
Kim, B.; Wattenberg, M.; Gilmer, J.; Cai, C.; Wexler, J.; Viegas, F.; et al. 2018 · 2018
Earlier work this paper cites.
Improving Fairness in Machine Learning Systems
Holstein, K.; Vaughan, J. W.; Daumé, H.; Dudik, M.; and Wallach, H. 2019 · 2019
Earlier work this paper cites.
Slice Finder: Automated Data Slicing for Model Validation
Polyzotis, N.; Whang, S.; Kraska, T. K.; and Chung, Y. 2019 · 2019
Earlier work this paper cites.
Errudite: Scalable, Reproducible, and Testable Error Analysis
Wu, T.; Ribeiro, M. T.; Heer, J.; and Weld, D. 2019 · 2019
Earlier work this paper cites.
Evaluating Machine Accuracy on ImageNet
Shankar, V.; Roelofs, R.; Mania, H.; Fang, A.; Recht, B.; and Schmidt, L. 2020 · 2020
Earlier work this paper cites.
No subclass left behind: Fine-grained robustness in coarse-grained classification problems
Sohoni, N.; Dunnmon, J.; Angus, G.; Gu, A.; and Ré, C. 2020 · 2020
Earlier work this paper cites.
Discovering and Validating AI Errors With Crowdsourced Failure Reports
Cabrera, Á. A.; Druck, A. J.; Hong, J. I.; and Perer, A. 2021 · 2021
Earlier work this paper cites.
The spotlight: A general method for discovering systematic errors in deep learning models
d’Eon, G.; d’Eon, J.; Wright, J. R.; and Leyton-Brown, K. 2021 · 2021
Cited alongside, same era.
Just train twice: Improving group robustness without training group information
Liu, E. Z.; Haghgoo, B.; Chen, A. S.; Raghunathan, A.; Koh, P. W.; Sagawa, S.; Liang, P.; and Finn, C. 2021 · 2021
Cited alongside, same era.
Learning Transferable Visual Models From Natural Language Supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; Krueger, G.; and Sutskever, I. 2021 · 2021
Cited alongside, same era.
Understanding Failures of Deep Networks via Robust Feature Extraction
Singla, S.; Nushi, B.; Shah, S.; Kamar, E.; and Horvitz, E. 2021 · 2021
Cited alongside, same era.
When and why vision-language models behave like bags-of-words, and what to do about it?
SEAL : Interactive Tool for Systematic Error Analysis and Labeling
Rajani, N.; Liang, W.; Chen, L.; Mitchell, M.; and Zou, J. 2022 · 2022
Later among the works it cites.
Actionable Auditing Revisited: Investigating the Impact of Publicly Naming Biased Performance Results of Commercial AI Products
Raji, I. D.; and Buolamwini, J. 2022 · 2022
Later among the works it cites.
Annotators with Attitudes: How Annotator Beliefs And Identities Bias Toxic Language Detection
Sap, M.; Swayamdipta, S.; Vianna, L.; Zhou, X.; Choi, Y.; and Smith, N. A. 2022 · 2022
Later among the works it cites.
When does dough become a bagel? Analyzing the remaining mistakes on ImageNet
Vasudevan, V.; Caine, B.; Gontijo-Lopes, R.; Fridovich-Keil, S.; and Roelofs, R. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yuksekgonul, M.; Bianchi, F.; Kalluri, P.; Jurafsky, D.; and Zou, J. 2023 · 2021
Cited alongside, same era.
Post hoc Explanations may be Ineffective for Detecting Unknown Spurious Correlation
Adebayo, J.; Muelly, M.; Abelson, H.; and Kim, B. 2022 · 2022
Cited alongside, same era.
DendroMap: Visual Exploration of Large-Scale Image Datasets for Machine Learning with Treemaps
Bertucci, D.; Hamid, M. M.; Anand, Y.; Ruangrotsakun, A.; Tabatabai, D.; Perez, M.; and Kahng, M. 2022 · 2022
Cited alongside, same era.
The unseen Black faces of AI algorithms
Birhane, A. 2022 · 2022
Cited alongside, same era.
What Did My AI Learn? How Data Scientists Make Sense of Model Behavior
Cabrera, A. A.; Tulio Ribeiro, M.; Lee, B.; Deline, R.; Perer, A.; and Drucker, S. M. 2022 · 2022
Cited alongside, same era.
Domino: Discovering Systematic Errors with Cross-Modal Embeddings
Eyuboglu, S.; Varma, M.; Saab, K. K.; Delbrouck, J.-B.; Lee-Messer, C.; Dunnmon, J.; Zou, J.; and Re, C. 2022 · 2022
Cited alongside, same era.
Adaptive Testing of Computer Vision Models
Gao, I.; Ilharco, G.; Lundberg, S.; and Ribeiro, M. T. 2022 · 2022
Cited alongside, same era.
Simple data balancing achieves competitive worst-group-accuracy
Idrissi, B. Y.; Arjovsky, M.; Pezeshki, M.; and Lopez-Paz, D. 2022 · 2022
Cited alongside, same era.
Balayn, A.; Rikalo, N.; Yang, J.; and Bozzon, A. 2023 · 2023
Closest in time.
Zeno: An Interactive Framework for Behavioral Evaluation of Machine Learning
Cabrera, A. A.; Fu, E.; Bertucci, D.; Holstein, K.; Talwalkar, A.; Hong, J. I.; and Perer, A. 2023 · 2023
Closest in time.
Non-task expert physicians benefit from correct explainable AI advice when reviewing X-rays
Gaube, S.; Suresh, H.; Raue, M.; Lermer, E.; Koch, T.; Hudecek, M.; Ackery, A.; Grover, S.; Coughlin, J.; Frey, D.; Kitamura, C.; Ghassemi, M.; and Colak, E. 2023 · 2023
Closest in time.
FAIlureNotes: Supporting Designers in Understanding the Limits of AI Models for Computer Vision Tasks
Moore, S.; Liao, Q. V.; and Subramonyam, H. 2023 · 2023
Closest in time.
Towards a More Rigorous Science of Blindspot Discovery in Image Models
Plumb, G.; Johnson, N.; Ángel Alexander Cabrera; and Talwalkar, A. 2023 · 2023
Closest in time.
Kaleidoscope: Semantically-grounded, context-specific ML model evaluation
Suresh, H.; Shanmugam, D.; Bryan, A.; Chen, T.; D’Amour, A.; Guttag, J. V.; and Satyanarayan., A. 2023 · 2023
Closest in time.
Dataset Interfaces: Diagnosing Model Failures Using Controllable Counterfactual Generation
Vendrow, J.; Jain, S.; Engstrom, L.; and Madry, A. 2023 · 2023
Closest in time.
Error Discovery by Clustering Influence Embeddings
Wang, F.; Adebayo, J.; Tan, S.; Garcia-Olano, D.; and Kokhlikyan, N. 2023 · 2023
Closest in time.
Discovering Bugs in Vision Models using Off-the-shelf Image Generation and Captioning
Wiles, O.; Albuquerque, I.; and Gowal, S. 2023 · 2023
Closest in time.