Fetching the paper…
Reading the bibliography…
We consider the issue of calibration in large language models (LLM).
Verification of forecasts expressed in terms of probability
Brier, G. W · 1950
Earlier work this paper cites.
The well-calibrated Bayesian
Dawid, A. P · 1982
Earlier work this paper cites.
The Helmholtz machine
Dayan, P., Hinton, G. E., Neal, R. M., and Zemel, R. S · 1995
Earlier work this paper cites.
Probabilistic outputs for support vector machines and comparisons to regularized likelihood methods
Platt, J. et al · 1999
Earlier work this paper cites.
Obtaining calibrated probability estimates from decision trees and naive Bayesian classifiers
Zadrozny, B. and Elkan, C · 2001
Earlier work this paper cites.
Transforming classifier scores into accurate multiclass probability estimates
Zadrozny, B. and Elkan, C · 2002
Earlier work this paper cites.
Strictly proper scoring rules, prediction, and estimation
Gneiting, T. and Raftery, A. E · 2007
Earlier work this paper cites.
Probabilistic forecasts, calibration and sharpness
Gneiting, T., Balabdaoui, F., and Raftery, A. E · 2007
Earlier work this paper cites.
Obtaining well calibrated probabilities using Bayesian binning
Naeini, M. P., Cooper, G., and Hauskrecht, M · 2015
Earlier work this paper cites.
Dropout as a Bayesian approximation: Representing model uncertainty in deep learning
Gal, Y. and Ghahramani, Z · 2016
Earlier work this paper cites.
Rethinking the inception architecture for computer vision
Szegedy, C., Vanhoucke, V., Ioffe, S., Shlens, J., and Wojna, Z · 2016
Earlier work this paper cites.
On calibration of modern neural networks
Guo, C., Pleiss, G., Sun, Y., and Weinberger, K. Q · 2017
Earlier work this paper cites.
Simple and scalable predictive uncertainty estimation using deep ensembles
Lakshminarayanan, B., Pritzel, A., and Blundell, C · 2017
Earlier work this paper cites.
Regularizing neural networks by penalizing confident output distributions
Pereyra, G., Tucker, G., Chorowski, J., Kaiser, Ł., and Hinton, G · 2017
Earlier work this paper cites.
mixup: Beyond empirical risk minimization
Zhang, H., Cisse, M., Dauphin, Y. N., and Lopez-Paz, D · 2017
Earlier work this paper cites.
Multi-level variational autoencoder: Learning disentangled representations from grouped observations
Bouchacourt, D., Tomioka, R., and Nowozin, S · 2018
Earlier work this paper cites.
Mrqa 2019 shared task: Evaluating generalization in reading comprehension
Fisch, A., Talmor, A., Jia, R., Seo, M., Choi, E., and Chen, D · 2019
Earlier work this paper cites.
Beyond temperature scaling: Obtaining well-calibrated multi-class probabilities with dirichlet calibration
Kull, M., Perello Nieto, M., Kängsepp, M., Silva Filho, T., Song, H., and Flach, P · 2019
Cited alongside, same era.
EDA: Easy data augmentation techniques for boosting performance on text classification tasks
Wei, J. and Zou, K · 2019
Cited alongside, same era.
Improving question answering by commonsense-based pre-training
Zhong, W., Tang, D., Duan, N., Zhou, M., Wang, J., and Yin, J · 2019
Cited alongside, same era.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., and Askell, A · 2020
Cited alongside, same era.
Calibration of pre-trained transformers
Desai, S. and Durrett, G · 2020
Cited alongside, same era.
Teaching models to express their uncertainty in words
Lin, S., Hilton, J., and Evans, O · 2022
Later among the works it cites.
Reducing conversational agents’ overconfidence through linguistic calibration
Mielke, S. J., Szlam, A., Dinan, E., and Boureau, Y.-L · 2022
Later among the works it cites.
Park, S. Y. and Caragea, C · 2022
Later among the works it cites.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Srivastava, A., Rastogi, A., Rao, A., Shoeb, A. A. M., Abid, A., Fisch, A., Brown, A. R., Santoro, A., Gupta, A., Garriga-Alonso, A., et al · 2022
Later among the works it cites.
Uncertainty quantification with pre-trained language models: A large-scale empirical analysis
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2020
Cited alongside, same era.
Calibrating deep neural networks using focal loss
Mukhoti, J., Kulharia, V., Sanyal, A., Golodetz, S., Torr, P., and Dokania, P · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Cited alongside, same era.
Incorporating BERT into neural machine translation
Zhu, J., Xia, Y., Wu, L., He, D., Qin, T., Zhou, W., Li, H., and Liu, T.-Y · 2020
Cited alongside, same era.
Top-label calibration and multiclass-to-binary reductions
Gupta, C. and Ramdas, A · 2021
Cited alongside, same era.
What are Bayesian neural network posteriors really like?
Izmailov, P., Vikram, S., Hoffman, M. D., and Wilson, A. G. G · 2021
Cited alongside, same era.
How can we know when language models know? on the calibration of language models for question answering
Jiang, Z., Araki, J., Ding, H., and Neubig, G · 2021
Cited alongside, same era.
Xiao, Y., Liang, P. P., Bhatt, U., Neiswanger, W., Salakhutdinov, R., and Morency, L.-P · 2022
Later among the works it cites.
Robust calibration with multi-domain temperature scaling
Yu, Y., Bates, S., Ma, Y., and Jordan, M · 2022
Later among the works it cites.
Prototypical calibration for few-shot learning of language models
Han, Z., Hao, Y., Dong, L., Sun, Y., and Wei, F · 2023
Later among the works it cites.
Generative calibration for in-context learning
Jiang, Z., Zhang, Y., Liu, C., Zhao, J., and Liu, K · 2023
Later among the works it cites.
Sample-dependent adaptive temperature scaling for improved calibration
Joy, T., Pinto, F., Lim, S.-N., Torr, P. H., and Dokania, P. K · 2023
Later among the works it cites.
GPT-4 technical report, 2023
OpenAI · 2023
Later among the works it cites.
Just ask for calibration: Strategies for eliciting calibrated confidence scores from language models fine-tuned with human feedback
Tian, K., Mitchell, E., Zhou, A., Sharma, A., Rafailov, R., Yao, H., Finn, C., and Manning, C · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Later among the works it cites.
Can llms express their uncertainty? an empirical evaluation of confidence elicitation in llms
Xiong, M., Hu, Z., Lu, X., Li, Y., Fu, J., He, J., and Hooi, B · 2023
Later among the works it cites.
On the calibration of large language models and alignment
Zhu, C., Xu, B., Wang, Q., Zhang, Y., and Mao, Z · 2023
Later among the works it cites.
Enhancing in-context learning via linear probe calibration
Abbas, M., Zhou, Y., Ram, P., Baracaldo, N., Samulowitz, H., Salonidis, T., and Chen, T · 2024
Closest in time.
Are you using test log-likelihood correctly?
Deshpande, S., Ghosh, S., Nguyen, T. D., and Broderick, T · 2024
Closest in time.
Batch calibration: Rethinking calibration for in-context learning and prompt engineering
Zhou, H., Wan, X., Proleev, L., Mincu, D., Chen, J., Heller, K. A., and Roy, S · 2024
Closest in time.