Fetching the paper…
Reading the bibliography…
When a deep learning model is deployed in the wild, it can encounter test data drawn from distributions different from the training data distribution and suffer drop in performance.
A database for handwritten text recognition research
Jonathan J. Hull · 1994
Earlier work this paper cites.
The mnist database of handwritten digits
Yann LeCun · 1998
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
Domain adaptation for sentiment classification
John Blitzer, Mark Dredze, and Fernando Pereira · 2007
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Unsupervised supervised learning i: Estimating classification and regression errors without labels
Pinar Donmez, Guy Lebanon, and Krishnakumar Balasubramanian · 2010
Earlier work this paper cites.
Adapting visual category models to new domains
Kate Saenko, Brian Kulis, Mario Fritz, and Trevor Darrell · 2010
Earlier work this paper cites.
Reading digits in natural images with unsupervised feature learning
Yuval Netzer, Tao Wang, Adam Coates, Alessandro Bissacco, Bo Wu, and Andrew Y Ng · 2011
Earlier work this paper cites.
Estimating accuracy from unlabeled data
Emmanouil Antonios Platanios, Avrim Blum, and Tom M. Mitchell · 2014
Earlier work this paper cites.
Estimating the accuracies of multiple classifiers without labeled data
Ariel Jaffe, Boaz Nadler, and Yuval Kluger · 2015
Earlier work this paper cites.
Dropout as a bayesian approximation: Representing model uncertainty in deep learning
Yarin Gal and Zoubin Ghahramani · 2016
Earlier work this paper cites.
Domain-adversarial training of neural networks
Yaroslav Ganin, Evgeniya Ustinova, Hana Ajakan, Pascal Germain, Hugo Larochelle, François Laviolette, Mario Marchand, and Victor Lempitsky · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
A baseline for detecting misclassified and out-of-distribution examples in neural networks
Dan Hendrycks and Kevin Gimpel · 2016
Cited alongside, same era.
Unsupervised risk estimation using only conditional independence structure
Jacob Steinhardt and Percy S Liang · 2016
Cited alongside, same era.
On calibration of modern neural networks
Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q Weinberger · 2017
Cited alongside, same era.
Densely connected convolutional networks
Gao Huang, Zhuang Liu, Laurens Van Der Maaten, and Kilian Q Weinberger · 2017
Cited alongside, same era.
Can you trust your model’s uncertainty? evaluating predictive uncertainty under dataset shift
Yaniv Ovadia, Emily Fertig, Jie Ren, Zachary Nado, David Sculley, Sebastian Nowozin, Joshua Dillon, Balaji Lakshminarayanan, and Jasper Snoek · 2019
Later among the works it cites.
The iwildcam 2020 competition dataset
Sara Beery, Elijah Cole, and Arvi Gjoka · 2020
Later among the works it cites.
Estimating generalization under distribution shifts via domain-invariant representations
Ching-Yao Chuang, Antonio Torralba, and Stefanie Jegelka · 2020
Later among the works it cites.
Wilds: A benchmark of in-the-wild distribution shifts
Pang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie, Marvin Zhang, Akshay Balsubramani, Weihua Hu, Michihiro Yasunaga, Richard Lanas Phillips, Sara Beery, et al · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Simple and scalable predictive uncertainty estimation using deep ensembles
Balaji Lakshminarayanan, Alexander Pritzel, and Charles Blundell · 2017
Cited alongside, same era.
Estimating accuracy from unlabeled data: A probabilistic logic approach
Emmanouil A. Platanios, Hoifung Poon, Tom M. Mitchell, and Eric Horvitz · 2017
Cited alongside, same era.
To trust or not to trust a classifier
Heinrich Jiang, Been Kim, Melody Guan, and Maya Gupta · 2018
Cited alongside, same era.
Addressing failure prediction by learning model confidence
Charles Corbière, Nicolas Thome, Avner Bar-Hen, Matthieu Cord, and Patrick Pérez · 2019
Cited alongside, same era.
To annotate or not? predicting performance drop under domain shift
Hady Elsahar and Matthias Gallé · 2019
Cited alongside, same era.
Deep ensembles: A loss landscape perspective
Stanislav Fort, Huiyi Hu, and Balaji Lakshminarayanan · 2019
Cited alongside, same era.
Benchmarking neural network robustness to common corruptions and perturbations
Dan Hendrycks and Thomas Dietterich · 2019
Cited alongside, same era.
Samuele Lo Piano · 2020
Later among the works it cites.
Learning to validate the predictions of black box classifiers on unseen data
Sebastian Schelter, Tammo Rukat, and Felix Biessmann · 2020
Later among the works it cites.
Mandoline: Model evaluation under distribution shift
Mayee F. Chen, Karan Goel, Nimit Sharad Sohoni, Fait Poms, Kayvon Fatahalian, and Christopher Ré · 2021
Closest in time.
What does rotation prediction tell us about classifier accuracy under varying testing environments?
Weijian Deng, Stephen Gould, and Liang Zheng · 2021
Closest in time.
Are labels always necessary for classifier accuracy evaluation?
Weijian Deng and Liang Zheng · 2021
Closest in time.
Predicting with confidence on unseen distributions
Devin Guillory, Vaishaal Shankar, Sayna Ebrahimi, Trevor Darrell, and Ludwig Schmidt · 2021
Closest in time.
Assessing generalization of SGD via disagreement
Yiding Jiang, Vaishnavh Nagarajan, Christina Baek, and J. Zico Kolter · 2021
Closest in time.