Fetching the paper…
Reading the bibliography…
Identifying how much a model ${\widehat{p}}_{\theta}(Y|X)$ knows about the stochastic real-world process $p(Y|X)$ it was trained on is important to ensure it avoids producing incorrect or "hallucinated" answers or taking unsafe actions.
Sui confini della probabilita
Cantelli, F. P · 1929
Earlier work this paper cites.
The well-calibrated Bayesian
Dawid, A. P · 1982
Earlier work this paper cites.
Present position and potential developments: Some personal views statistical theory the prequential approach
Dawid, A. P · 1984
Earlier work this paper cites.
Calibration-based empirical probability
Dawid, A. P · 1985
Earlier work this paper cites.
Self-calibrating priors do not exist
Oakes, D · 1985
Earlier work this paper cites.
Signature verification using a” siamese” time delay neural network
Bromley, J., Guyon, I., LeCun, Y., Säckinger, E., and Shah, R · 1993
Earlier work this paper cites.
Probability inequalities for sums of bounded random variables
Hoeffding, W · 1994
Earlier work this paper cites.
Regression and classification using Gaussian process priors
Bernardo, J., Berger, J., Dawid, A., Smith, A., et al · 1998
Earlier work this paper cites.
Asymptotic calibration
Foster, D. P. and Vohra, R. V · 1998
Earlier work this paper cites.
Planning and acting in partially observable stochastic domains
Kaelbling, L. P., Littman, M. L., and Cassandra, A. R · 1998
Earlier work this paper cites.
Transforming classifier scores into accurate multiclass probability estimates
Zadrozny, B. and Elkan, C · 2002
Earlier work this paper cites.
Completely positive matrices
Berman, A. and Shaked-Monderer, N · 2003
Earlier work this paper cites.
Calibration with many checking rules
Sandroni, A., Smorodinsky, R., and Vohra, R. V · 2003
Earlier work this paper cites.
Phi‐Coefficient
Chedzoy, O. B · 2005
Earlier work this paper cites.
Classification
Rasmussen, C. E. and Williams, C. K. I · 2005
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A · 2009
Earlier work this paper cites.
Noise-contrastive estimation: A new estimation principle for unnormalized statistical models
Gutmann, M. and Hyvärinen, A · 2010
Earlier work this paper cites.
TensorFlow: Large-scale machine learning on heterogeneous systems, 2015
Abadi, M., Agarwal, A., Barham, P., Brevdo, E., Chen, Z., Citro, C., Corrado, G. S., Davis, A., Dean, J., Devin, M., Ghemawat, S., Goodfellow, I., Harp, A., Irving, G., Isard, M., Jia, Y., Jozefowicz, R., Kaiser, L., Kudlur, M., Levenberg, J., Mané, D., Monga, R., Moore, S., Murray, D., Olah, C., Schuster, M., Shlens, J., Steiner, B., Sutskever, I., Talwar, K., Tucker, P., Vanhoucke, V., Vasudevan, V., Viégas, F., Vinyals, O., Warden, P., Wattenberg, M., Wicke, M., Yu, Y., and Zheng, X · 2015
Earlier work this paper cites.
Dropout as a Bayesian approximation: Representing model uncertainty in deep learning
Gal, Y. and Ghahramani, Z · 2015
Earlier work this paper cites.
Scalable variational Gaussian process classification
Hensman, J., Matthews, A., and Ghahramani, Z · 2015
Earlier work this paper cites.
Novel decompositions of proper scoring rules for classification: Score adjustment as precursor to calibration
Kull, M. and Flach, P · 2015
Earlier work this paper cites.
Second order calibration: A simple way to get approximate posteriors
Muralidharan, O. and Najmi, A · 2015
Earlier work this paper cites.
Ba, J. L., Kiros, J. R., and Hinton, G. E · 2016
Earlier work this paper cites.
Garrabrant, S., Benson-Tilsen, T., Critch, A., Soares, N., and Taylor, J · 2016
Earlier work this paper cites.
Simple and scalable predictive uncertainty estimation using deep ensembles
Lakshminarayanan, B., Pritzel, A., and Blundell, C · 2016
Earlier work this paper cites.
Zagoruyko, S. and Komodakis, N · 2016
Earlier work this paper cites.
On calibration of modern neural networks
Guo, C., Pleiss, G., Sun, Y., and Weinberger, K. Q · 2017
Earlier work this paper cites.
What uncertainties do we need in Bayesian deep learning for computer vision?
Kendall, A. and Gal, Y · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2017
Earlier work this paper cites.
Equivalence between policy gradients and soft Q-learning
Schulman, J., Chen, X., and Abbeel, P · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Cited alongside, same era.
Nonparametric predictive distributions based on conformal prediction
Vovk, V., Shen, J., Manokhin, V., and Xie, M.-g · 2017
Cited alongside, same era.
JAX: composable transformations of Python+NumPy programs, 2018
Bradbury, J., Frostig, R., Hawkins, P., Johnson, M. J., Leary, C., Maclaurin, D., Necula, G., Paszke, A., VanderPlas, J., Wanderman-Milne, S., and Zhang, Q · 2018
Cited alongside, same era.
Multicalibration: Calibration for the (computationally-identifiable) masses
Hébert-Johnson, Ú., Kim, M. P., Reingold, O., and Rothblum, G. N · 2018
Cited alongside, same era.
Distribution-free calibration guarantees for histogram binning without sample splitting
Gupta, C. and Ramdas, A · 2021
Later among the works it cites.
DEUP: Direct epistemic uncertainty prediction
Lahlou, S., Jain, M., Nekoei, H., Butoi, V. I., Bertin, P., Rector-Brooks, J., Korablyov, M., and Bengio, Y · 2021
Later among the works it cites.
Uncertainty baselines: Benchmarks for uncertainty & robustness in deep learning
Nado, Z., Band, N., Collier, M., Djolonga, J., Dusenberry, M. W., Farquhar, S., Feng, Q., Filos, A., Havasi, M., Jenatton, R., et al · 2021
Later among the works it cites.
Shaking the foundations: delusions in sequence models for interaction and control
Ortega, P. A., Kunesch, M., Delétang, G., Genewein, T., Grau-Moya, J., Veness, J., Buchli, J., Degrave, J., Piot, B., Perolat, J., et al · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Pacgan: The power of two samples in generative adversarial networks
Lin, Z., Khetan, A., Fanti, G., and Oh, S · 2018
Cited alongside, same era.
Predictive uncertainty estimation via prior networks
Malinin, A. and Gales, M · 2018
Cited alongside, same era.
Evidential deep learning to quantify classification uncertainty
Sensoy, M., Kandemir, M., and Kaplan, L. M · 2018
Cited alongside, same era.
Augmix: A simple data processing method to improve robustness and uncertainty
Hendrycks, D., Mu, N., Cubuk, E. D., Zoph, B., Gilmer, J., and Lakshminarayanan, B · 2019
Cited alongside, same era.
Verified uncertainty calibration
Kumar, A., Liang, P. S., and Ma, T · 2019
Cited alongside, same era.
A simple baseline for Bayesian uncertainty in deep learning
Maddox, W. J., Garipov, T., Izmailov, P., Vetrov, D. P., and Wilson, A. G · 2019
Cited alongside, same era.
Human uncertainty makes classification more robust
Peterson, J. C., Battleday, R. M., Griffiths, T. L., and Russakovsky, O · 2019
Cited alongside, same era.
Osband, I., Wen, Z., Asghari, M., Ibrahimi, M., Lu, X., and Roy, B. V · 2021
Later among the works it cites.
Learning with noisy labels revisited: A study using real-world human annotations
Wei, J., Zhu, Z., Cheng, H., Liu, T., Niu, G., and Liu, Y · 2021
Later among the works it cites.
From predictions to decisions: The importance of joint predictive distributions
Wen, Z., Osband, I., Qin, C., Lu, X., Ibrahimi, M., Dwaracherla, V., Asghari, M., and Van Roy, B · 2021
Later among the works it cites.
Decoding methods in neural language generation: a survey
Zarrieß, S., Voigt, H., and Schüz, S · 2021
Later among the works it cites.
Pitfalls of epistemic uncertainty quantification through loss minimisation
Bengs, V., Hüllermeier, E., and Waegeman, W · 2022
Later among the works it cites.
The calibration generalization gap
Carrell, A., Mallinar, N. R., Lucas, J., and Nakkiran, P · 2022
Later among the works it cites.
Transfer and marginalize: Explaining away label noise with privileged information
Collier, M., Jenatton, R., Kokiopoulou, E., and Berent, J · 2022
Later among the works it cites.
Zigzag: Universal sampling-free uncertainty estimation through two-step inference
Durasov, N., Dorndorf, N., Le, H., and Fua, P · 2022
Later among the works it cites.
Survey of hallucination in natural language generation
Ji, Z., Lee, N., Frieske, R., Yu, T., Su, D., Xu, Y., Ishii, E., Bang, Y., Dai, W., Madotto, A., and Fung, P · 2022
Later among the works it cites.
Language models (mostly) know what they know
Kadavath, S., Conerly, T., Askell, A., Henighan, T. J., Drain, D., Perez, E., Schiefer, N., Dodds, Z., DasSarma, N., Tran-Johnson, E., Johnston, S., El-Showk, S., Jones, A., Elhage, N., Hume, T., Chen, A., Bai, Y., Bowman, S., Fort, S., Ganguli, D., Hernandez, D., Jacobson, J., Kernion, J., Kravec, S., Lovitt, L., Ndousse, K., Olsson, C., Ringer, S., Amodei, D., Brown, T. B., Clark, J., Joseph, N., Mann, B., McCandlish, S., Olah, C., and Kaplan, J · 2022
Later among the works it cites.
Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation
Kuhn, L., Gal, Y., and Farquhar, S · 2022
Later among the works it cites.
Beyond calibration: estimating the grouping loss of modern neural networks
Perez-Lebel, A., Morvan, M. L., and Varoquaux, G · 2022
Later among the works it cites.
Is one annotation enough?-a data-centric image classification benchmark for noisy and ambiguous label estimation
Schmarje, L., Grossmann, V., Zelenka, C., Dippel, S., Kiko, R., Oszust, M., Pastell, M., Stracke, J., Valros, A., Volkmann, N., et al · 2022
Later among the works it cites.
Self-consistency improves chain of thought reasoning in language models
Wang, X., Wei, J., Schuurmans, D., Le, Q., hsin Chi, E. H., and Zhou, D · 2022
Later among the works it cites.
On second-order scoring rules for epistemic uncertainty quantification
Bengs, V., Hüllermeier, E., and Waegeman, W · 2023
Later among the works it cites.
It’s mbr all the way down: Modern generation techniques through the lens of minimum bayes risk
Bertsch, A., Xie, A., Neubig, G., and Gormley, M. R · 2023
Later among the works it cites.
When does optimizing a proper loss yield calibration?
Błasiok, J., Gopalan, P., Hu, L., and Nakkiran, P · 2023
Later among the works it cites.
Universal self-consistency for large language model generation
Chen, X., Aksitov, R., Alon, U., Ren, J., Xiao, K., Yin, P., Prakash, S., Sutton, C., Wang, X., and Zhou, D · 2023
Later among the works it cites.
A first course in causal inference
Ding, P · 2023
Later among the works it cites.
Calibrated language models must hallucinate
Kalai, A. T. and Vempala, S. S · 2023
Later among the works it cites.
Collision probability matching loss for disentangling epistemic uncertainty from aleatoric uncertainty
Narimatsu, H., Ozawa, M., and Kumano, S · 2023
Later among the works it cites.
LEVER: learning to verify language-to-code generation with execution
Ni, A., Iyer, S., Radev, D. R., Stoyanov, V., tau Yih, W., Wang, S. I., and Lin, X. V · 2023
Later among the works it cites.
GPT-4 technical report, 2023
OpenAI · 2023
Later among the works it cites.
Second-order uncertainty quantification: A distance-based approach
Sale, Y., Bengs, V., Caprio, M., and Hüllermeier, E · 2023
Later among the works it cites.