Fetching the paper…
Reading the bibliography…
For an AI system to be reliable, the confidence it expresses in its decisions must match its accuracy.
Verification of forecasts expressed in terms of probability
Brier, G. W. (1950) · 1950
Earlier work this paper cites.
A bias-corrected decomposition of the brier score
Ferro, C. A. T. and Fricker, T. E. (2012) · 1960
Earlier work this paper cites.
Signal detection theory and psychophysics
Green, D. M., Swets, J. A., et al. (1966) · 1966
Earlier work this paper cites.
Gradient-based learning applied to document recognition
LeCun, Y., Bottou, L., Bengio, Y., and Haffner, P. (1998) · 1998
Earlier work this paper cites.
Probabilistic outputs for support vector machines and comparisons to regularized likelihood methods
Platt, J. et al. (1999) · 1999
Earlier work this paper cites.
Obtaining calibrated probability estimates from decision trees and naive bayesian classifiers
Zadrozny, B. and Elkan, C. (2001) · 2001
Earlier work this paper cites.
Calibrating deep neural networks using focal loss
Mukhoti, J., Kulharia, V., Sanyal, A., Golodetz, S., Torr, P. H., and Dokania, P. K. (2020) · 2002
Earlier work this paper cites.
Transforming classifier scores into accurate multiclass probability estimates
Zadrozny, B. and Elkan, C. (2002) · 2002
Earlier work this paper cites.
Mix-n-match: Ensemble and compositional methods for uncertainty calibration in deep learning
Zhang, J., Kailkhura, B., and Han, T. (2020) · 2003
Earlier work this paper cites.
Bayesian Data Analysis
Gelman, A., Carlin, J. B., Stern, H. S., and Rubin, D. B. (2004) · 2004
Earlier work this paper cites.
Calibration of neural networks using splines
Gupta, K., Rahimi, A., Ajanthan, T., Mensink, T., Sminchisescu, C., and Hartley, R. (2020) · 2006
Earlier work this paper cites.
ImageNet: A Large-Scale Hierarchical Image Database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L. (2009) · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A. and Hinton, G. (2009) · 2009
Earlier work this paper cites.
On the convexity of roc curves estimated from radiological test results
Pesce, L. L., Metz, C. E., and Berbaum, K. S. (2010) · 2010
Earlier work this paper cites.
Estimating reliability and resolution of probability forecasts through decomposition of the empirical score
Brocker, J. (2012) · 2012
Earlier work this paper cites.
Improving model calibration with accuracy versus uncertainty optimization
Krishnan, R. and Tickoo, O. (2020) · 2012
Cited alongside, same era.
I-spline smoothing for calibrating predictive models
Wu, Y., Jiang, X., Kim, J., and Ohno-Machado, L. (2012) · 2012
Cited alongside, same era.
Vision meets robotics: The kitti dataset
Geiger, A., Lenz, P., Stiller, C., and Urtasun, R. (2013) · 2013
Cited alongside, same era.
Intriguing properties of neural networks
Szegedy, C., Zaremba, W., Sutskever, I., Bruna, J., Erhan, D., Goodfellow, I. J., and Fergus, R. (2013) · 2013
Cited alongside, same era.
Obtaining well calibrated probabilities using bayesian binning
Naeini, M. P., Cooper, G. F., and Hauskrecht, M. (2015) · 2015
Cited alongside, same era.
Calibration of medical diagnostic classifier scores to the probability of disease
Chen, W., Sahiner, B., Samuelson, F., Pezeshk, A., and Petrick, N. (2018) · 2018
Later among the works it cites.
Receiver operating characteristic (roc) curves
Gneiting, T. and Vogel, P. (2018) · 2018
Later among the works it cites.
Trainable calibration measures for neural networks from kernel mean embeddings
Kumar, A., Sarawagi, S., and Jain, U. (2018) · 2018
Later among the works it cites.
Focal loss for dense object detection
Lin, T.-Y., Goyal, P., Girshick, R., He, K., and Dollár, P. (2018) · 2018
Later among the works it cites.
A guide to deep learning in healthcare
Esteva, A., Robicquet, A., Ramsundar, B., Kuleshov, V., DePristo, M., Chou, K., Cui, C., Corrado, G., Thrun, S., and Dean, J. (2019) · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Development and validation of a deep learning algorithm for detection of diabetic retinopathy in retinal fundus photographs
Gulshan, V., Peng, L., Coram, M., Stumpe, M. C., Wu, D., Narayanaswamy, A., Venugopalan, S., Widner, K., Madams, T., Cuadros, J., et al. (2016) · 2016
Cited alongside, same era.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J. (2016) · 2016
Cited alongside, same era.
Deep networks with stochastic depth
Huang, G., Sun, Y., Liu, Z., Sedra, D., and Weinberger, K. Q. (2016) · 2016
Cited alongside, same era.
Zagoruyko, S. and Komodakis, N. (2016) · 2016
Cited alongside, same era.
Dermatologist-level classification of skin cancer with deep neural networks
Esteva, A., Kuprel, B., Novoa, R. A., Ko, J., Swetter, S. M., Blau, H. M., and Thrun, S. (2017) · 2017
Cited alongside, same era.
On calibration of modern neural networks
Guo, C., Pleiss, G., Sun, Y., and Weinberger, K. Q. (2017) · 2017
Cited alongside, same era.
Densely connected convolutional networks
Huang, G., Liu, Z., Van Der Maaten, L., and Weinberger, K. Q. (2017) · 2017
Cited alongside, same era.
Hendrycks, D. and Dietterich, T. (2019) · 2019
Later among the works it cites.
Verified uncertainty calibration
Kumar, A., Liang, P. S., and Ma, T. (2019) · 2019
Later among the works it cites.
Nn_calibration
Kängsepp, M. (2019) · 2019
Later among the works it cites.
Measuring calibration in deep learning
Nixon, J., Dusenberry, M. W., Zhang, L., Jerfel, G., and Tran, D. (2019) · 2019
Later among the works it cites.
Do imagenet classifiers generalize to imagenet?
Recht, B., Roelofs, R., Schmidt, L., and Shankar, V. (2019) · 2019
Later among the works it cites.
Calibration tests in multi-class classification: A unifying framework
Widmann, D., Lindsten, F., and Zachariah, D. (2019) · 2019
Later among the works it cites.
nuscenes: A multimodal dataset for autonomous driving
Caesar, H., Bankiti, V., Lang, A. H., Vora, S., Liong, V. E., Xu, Q., Krishnan, A., Pan, Y., Baldan, G., and Beijbom, O. (2020) · 2020
Closest in time.
Scalability in perception for autonomous driving: Waymo open dataset
Sun, P., Kretzschmar, H., Dotiwalla, X., Chouard, A., Patnaik, V., Tsui, P., Guo, J., Zhou, Y., Chai, Y., Caine, B., et al. (2020) · 2020
Closest in time.
Soft calibration objectives for neural networks
Karandikar, A., Cain, N., Tran, D., Lakshminarayanan, B., Shlens, J., Mozer, M. C., and Roelofs, B. (2021) · 2021
Closest in time.