Fetching the paper…
Reading the bibliography…
Understanding and explaining the mistakes made by trained models is critical to many machine learning objectives, such as improving robustness, addressing concept drift, and mitigating biases.
The Validity and Practicality of Sun-Reactive Skin Types I Through VI
Fitzpatrick, T. B · 1988
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Deep inside convolutional networks: Visualising image classification models and saliency maps
Simonyan, K., Vedaldi, A., and Zisserman, A · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Iterative orthogonal feature projection for diagnosing bias in black-box models
Adebayo, J. and Kagal, L · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
A baseline for detecting misclassified and out-of-distribution examples in neural networks
Hendrycks, D. and Gimpel, K · 2016
Earlier work this paper cites.
Squeezenet: Alexnet-level accuracy with 50x fewer parameters and¡ 0.5 mb model size
Iandola, F. N., Han, S., Moskewicz, M. W., Ashraf, K., Dally, W. J., and Keutzer, K · 2016
Earlier work this paper cites.
Learning deep features for discriminative localization
Zhou, B., Khosla, A., Lapedriza, A., Oliva, A., and Torralba, A · 2016
Earlier work this paper cites.
Densely connected convolutional networks
Huang, G., Liu, Z., Van Der Maaten, L., and Weinberger, K. Q · 2017
Earlier work this paper cites.
Understanding black-box predictions via influence functions
Koh, P. W. and Liang, P · 2017
Earlier work this paper cites.
Counterfactual explanations without opening the black box: Automated decisions and the gdpr
Wachter, S., Mittelstadt, B., and Russell, C · 2017
Earlier work this paper cites.
Hospital-scale chest x-ray database and benchmarks on weakly-supervised classification and localization of common thorax diseases
Wang, X., Peng, Y., Lu, L., Lu, Z., Bagheri, M., and Summers, R · 2017
Earlier work this paper cites.
Sanity checks for saliency maps
Adebayo, J., Gilmer, J., Muelly, M., Goodfellow, I. J., Hardt, M., and Kim, B · 2018
Earlier work this paper cites.
Grad-cam++: Generalized gradient-based visual explanations for deep convolutional networks
Chattopadhay, A., Sarkar, A., Howlader, P., and Balasubramanian, V. N · 2018
Earlier work this paper cites.
Explanations based on the missing: Towards contrastive explanations with pertinent negatives
Dhurandhar, A., Chen, P.-Y., Luss, R., Tu, C.-C., Ting, P., Shanmugam, K., and Das, P · 2018
Earlier work this paper cites.
Net2vec: Quantifying and explaining how concepts are encoded by filters in deep neural networks
Fong, R. and Vedaldi, A · 2018
Cited alongside, same era.
Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (tcav)
Kim, B., Wattenberg, M., Gilmer, J., Cai, C., Wexler, J., Viegas, F., et al · 2018
Cited alongside, same era.
Detecting and correcting for label shift with black box predictors
Lipton, Z., Wang, Y.-X., and Smola, A · 2018
Cited alongside, same era.
Learning under concept drift: A review
Lu, J., Liu, A., Dong, F., Gu, F., Gama, J., and Zhang, G · 2018
Cited alongside, same era.
A survey on automated melanoma detection
Okur, E. and Turkan, M · 2018
Cited alongside, same era.
Semidefinite relaxations for certifying robustness to adversarial examples
The what-if tool: Interactive probing of machine learning models
Wexler, J., Pushkarna, M., Bolukbasi, T., Wattenberg, M., Viégas, F., and Wilson, J · 2019
Later among the works it cites.
Cocox: Generating conceptual and counterfactual explanations via fault-lines
Akula, A., Wang, S., and Zhu, S.-C · 2020
Later among the works it cites.
Evaluating saliency map explanations for convolutional neural networks: a user study
Alqaraawi, A., Schuessler, M., Weiß, P., Costanza, E., and Bianchi-Berthouze, N · 2020
Later among the works it cites.
Wilds: A benchmark of in-the-wild distribution shifts
Koh, P. W., Sagawa, S., Marklund, H., Xie, S. M., Zhang, M., Balsubramani, A., Hu, W., Yasunaga, M., Phillips, R. L., Gao, I., et al · 2020
Later among the works it cites.
Gender imbalance in medical imaging datasets produces biased classifiers for computer-aided diagnosis
Larrazabal, A. J., Nieto, N., Peterson, V., Milone, D. H., and Ferrante, E · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Raghunathan, A., Steinhardt, J., and Liang, P · 2018
Cited alongside, same era.
Interpretable basis decomposition for visual explanation
Zhou, B., Sun, Y., Bau, D., and Torralba, A · 2018
Cited alongside, same era.
Gradio: Hassle-free sharing and testing of ml models in the wild
Abid, A., Abdalla, A., Abid, A., Khan, D., Alfozan, A., and Zou, J · 2019
Cited alongside, same era.
Towards automatic concept-based explanations
Ghorbani, A., Wexler, J., Zou, J., and Kim, B · 2019
Cited alongside, same era.
Counterfactual visual explanations
Goyal, Y., Wu, Z., Ernst, J., Batra, D., Parikh, D., and Lee, S · 2019
Cited alongside, same era.
Chexpert: A large chest radiograph dataset with uncertainty labels and expert comparison
Irvin, J., Rajpurkar, P., Ko, M., Yu, Y., Ciurea-Ilcus, S., Chute, C., Marklund, H., Haghgoo, B., Ball, R., Shpanskaya, K., et al · 2019
Cited alongside, same era.
Multiaccuracy: Black-box post-processing for fairness in classification
Kim, M. P., Ghorbani, A., and Zou, J · 2019
Cited alongside, same era.
Chexclusion: Fairness gaps in deep chest x-ray classifiers
Seyyed-Kalantari, L., Liu, G., McDermott, M., Chen, I. Y., and Ghassemi, M · 2020
Later among the works it cites.
Don’t judge an object by its context: Learning to overcome contextual bias
Singh, K. K., Mahajan, D. K., Grauman, K., Lee, Y. J., Feiszli, M., and Ghadiyaram, D · 2020
Later among the works it cites.
Explanation by progressive exaggeration
Singla, S., Pollack, B., Chen, J., and Batmanghelich, K · 2020
Later among the works it cites.
Counterfactual explanations for machine learning: A review
Verma, S., Dickerson, J., and Hines, K · 2020
Later among the works it cites.
The spotlight: A general method for discovering systematic errors in deep learning models
d’Eon, G., d’Eon, J., Wright, J. R., and Leyton-Brown, K · 2021
Closest in time.
Groh, M., Harris, C., Soenksen, L., Lau, F., Han, R., Kim, A., Koochek, A., and Badri, O · 2021
Closest in time.
Dynabench: Rethinking benchmarking in nlp
Kiela, D., Bartolo, M., Nie, Y., Kaushik, D., Geiger, A., Wu, Z., Vidgen, B., Prasad, G., Singh, A., Ringshia, P., et al · 2021
Closest in time.
Metadataset: A dataset of datasets for evaluating distribution shifts and training conflicts
Liang, W. and Zou, J · 2021
Closest in time.
A patient-centric dataset of images and metadata for identifying melanomas using clinical context
Rotemberg, V., Kurtansky, N., Betz-Stablein, B., Caffery, L., Chousakos, E., Codella, N., Combalia, M., Dusza, S., Guitera, P., Gutman, D., et al · 2021
Closest in time.
How medical ai devices are evaluated: limitations and recommendations from an analysis of fda approvals
Wu, E., Wu, K., Daneshjou, R., Ouyang, D., Ho, D. E., and Zou, J · 2021
Closest in time.