Fetching the paper…
Reading the bibliography…
Applications of large language models often involve the generation of free-form responses, in which case uncertainty quantification becomes challenging.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. (2020) · 1901
Earlier work this paper cites.
The Foundations of Statistics
Savage, L. J. (1954) · 1954
Earlier work this paper cites.
Uncertainty, information, and sequential experiments
DeGroot, M. H. (1962) · 1962
Earlier work this paper cites.
Reliability of subjective probability forecasts of precipitation and temperature
Murphy, A. H. and Winkler, R. L. (1977) · 1977
Earlier work this paper cites.
Mathematical statistics: basic ideas and selected topics
Bickel, P. J. and Doksum, K. A. (1997) · 1997
Earlier work this paper cites.
Game theory, maximum entropy, minimum discrepancy and robust Bayesian decision theory
Grünwald, P. D. and Dawid, A. P. (2004) · 2004
Earlier work this paper cites.
Aleatory or Epistemic? Does it Matter?
Der Kiureghian, A. and Ditlevsen, O. (2009) · 2009
Earlier work this paper cites.
Statistical machine translation
Koehn, P. (2009) · 2009
Earlier work this paper cites.
Accuracy-rejection curves (arcs) for comparing classification methods with a reject option
Nadeem, M. S. A., Zucker, J.-D., and Hanczar, B. (2009) · 2009
Earlier work this paper cites.
Model uncertainty
Marinacci, M. (2015) · 2015
Earlier work this paper cites.
Obtaining well calibrated probabilities using bayesian binning
Naeini, M. P., Cooper, G., and Hauskrecht, M. (2015) · 2015
Earlier work this paper cites.
chrf: character n-gram f-score for automatic mt evaluation
Popović, M. (2015) · 2015
Earlier work this paper cites.
Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension
Joshi, M., Choi, E., Weld, D. S., and Zettlemoyer, L. (2017) · 2017
Earlier work this paper cites.
What Uncertainties do we Need in Bayesian deep learning for computer vision?
Kendall, A. and Gal, Y. (2017) · 2017
Earlier work this paper cites.
Crowdsourcing multiple choice science questions
Welbl, J., Liu, N. F., and Gardner, M. (2017) · 2017
Earlier work this paper cites.
Predict responsibly: improving fairness and accuracy by learning to defer
Madras, D., Pitassi, T., and Zemel, R. (2018) · 2018
Earlier work this paper cites.
A call for clarity in reporting bleu scores
Post, M. (2018) · 2018
Earlier work this paper cites.
Natural questions: A benchmark for question answering research
Kwiatkowski, T., Palomaki, J., Redfield, O., Collins, M., Parikh, A., Alberti, C., Epstein, D., Polosukhin, I., Devlin, J., Lee, K., Toutanova, K., Jones, L., Kelcey, M., Chang, M.-W., Dai, A. M., Uszkoreit, J., Le, Q., and Petrov, S. (2019) · 2019
Earlier work this paper cites.
Coqa: A conversational question answering challenge
Reddy, S., Chen, D., and Manning, C. D. (2019) · 2019
Earlier work this paper cites.
A theory of usable information under computational constraints
Xu, Y., Zhao, S., Song, J., Stewart, R., and Ermon, S. (2019) · 2019
Earlier work this paper cites.
Selective question answering under domain shift
Kamath, A., Jia, R., and Liang, P. (2020) · 2020
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W.-t., Rocktäschel, T., et al. (2020) · 2020
Earlier work this paper cites.
A class of models for Bayesian predictive inference
Berti, P., Dreassi, E., Pratelli, L., and Rigo, P. (2021) · 2021
Earlier work this paper cites.
How can we know when language models know? on the calibration of language models for question answering
Jiang, Z., Araki, J., Ding, H., and Neubig, G. (2021) · 2021
Earlier work this paper cites.
Truthfulqa: Measuring how models mimic human falsehoods
Lin, S., Hilton, J., and Evans, O. (2021) · 2021
Cited alongside, same era.
On hallucination and predictive uncertainty in conditional language generation
Xiao, Y. and Wang, W. Y. (2021) · 2021
Cited alongside, same era.
An explanation of in-context learning as implicit Bayesian inference
Xie, S. M., Raghunathan, A., Liang, P., and Ma, T. (2021) · 2021
Cited alongside, same era.
What learning algorithm is in-context learning? investigations with linear models
Akyürek, E., Schuurmans, D., Andreas, J., Ma, T., and Zhou, D. (2022) · 2022
Cited alongside, same era.
Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback
Tian, K., Mitchell, E., Zhou, A., Sharma, A., Rafailov, R., Yao, H., Finn, C., and Manning, C. (2023) · 2023
Later among the works it cites.
Is chatgpt a good nlg evaluator? a preliminary study
Wang, J., Liang, Y., Meng, F., Sun, Z., Shi, H., Li, Z., Xu, J., Qu, J., and Zhou, J. (2023) · 2023
Later among the works it cites.
Evaluating progress in automatic chest x-ray radiology report generation
Yu, F., Endo, M., Krishnan, R., Pan, I., Tsai, A., Reis, E. P., Fonseca, E. K. U. N., Lee, H. M. H., Abad, Z. S. H., Ng, A. Y., et al. (2023) · 2023
Later among the works it cites.
Zhang, J., Li, Z., Das, K., Malin, B. A., and Kumar, S. (2023) · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dohan, D., Xu, W., Lewkowycz, A., Austin, J., Bieber, D., Lopes, R. G., Wu, Y., Michalewski, H., Saurous, R. A., Sohl-Dickstein, J., et al. (2022) · 2022
Cited alongside, same era.
Survey of low-resource machine translation
Haddow, B., Bawden, R., Barone, A. V. M., Helcl, J., and Birch, A. (2022) · 2022
Cited alongside, same era.
Deep learning for text style transfer: A survey
Jin, D., Jin, Z., Hu, Z., Vechtomova, O., and Mihalcea, R. (2022) · 2022
Cited alongside, same era.
Language Models (Mostly) Know What They Know
Kadavath, S., Conerly, T., Askell, A., Henighan, T., Drain, D., Perez, E., Schiefer, N., Hatfield-Dodds, Z., DasSarma, N., Tran-Johnson, E., Johnston, S., El-Showk, S., Jones, A., Elhage, N., Hume, T., Chen, A., Bai, Y., Bowman, S., Fort, S., Ganguli, D., Hernandez, D., Jacobson, J., Kernion, J., Kravec, S., Lovitt, L., Ndousse, K., Olsson, C., Ringer, S., Amodei, D., Brown, T., Clark, J., Joseph, N., Mann, B., McCandlish, S., Olah, C., and Kaplan, J. (2022) · 2022
Cited alongside, same era.
Reducing conversational agents’ overconfidence through linguistic calibration
Mielke, S. J., Szlam, A., Dinan, E., and Boureau, Y.-L. (2022) · 2022
Cited alongside, same era.
No language left behind: Scaling human-centered machine translation
NLLB Team, Costa-jussà, M. R., Cross, J., Çelebi, O., Elbayad, M., Heafield, K., Heffernan, K., Kalbassi, E., Lam, J., Licht, D., Maillard, J., Sun, A., Wang, S., Wenzek, G., Youngblood, A., Akula, B., Barrault, L., Mejia-Gonzalez, G., Hansanti, P., Hoffman, J., Jarrett, S., Sadagopan, K. R., Rowe, D., Spruit, S., Tran, C., Andrews, P., Ayan, N. F., Bhosale, S., Edunov, S., Fan, A., Gao, C., Goswami, V., Guzmán, F., Koehn, P., Mourachko, A., Ropers, C., Saleem, S., Schwenk, H., and Wang, J. (2022) · 2022
Cited alongside, same era.
Self-consistency improves chain of thought reasoning in language models
Wang, X., Wei, J., Schuurmans, D., Le, Q., Chi, E., Narang, S., Chowdhery, A., and Zhou, D. (2022) · 2022
Cited alongside, same era.
From Predictions to Decisions: The Importance of Joint Predictive Distributions
Wen, Z., Osband, I., Qin, C., Lu, X., Ibrahimi, M., Dwaracherla, V., Asghari, M., and Van Roy, B. (2022) · 2022
Cited alongside, same era.
Agarwal, R., Singh, A., Zhang, L. M., Bohnet, B., Chan, S., Anand, A., Abbas, Z., Nova, A., Co-Reyes, J. D., Chu, E., Behbahani, F., Faust, A., and Larochelle, H. (2024) · 2024
Closest in time.
Distinguishing the Knowable from the Unknowable with Language Models
Ahdritz, G., Qin, T., Vyas, N., Barak, B., and Edelman, B. L. (2024) · 2024
Closest in time.
How many Opinions does your LLM have? Improving Uncertainty Estimation in NLG
Aichberger, L., Schweighofer, K., Ielanskyi, M., and Hochreiter, S. (2024) · 2024
Closest in time.
Claude 3.5 sonnet model card addendum
Anthropic (2024) · 2024
Closest in time.
Linguistic Calibration of Long-Form Generations
Band, N., Li, X., Ma, T., and Hashimoto, T. (2024) · 2024
Closest in time.
Maira-2: Grounded radiology report generation
Bannur, S., Bouzid, K., Castro, D. C., Schwaighofer, A., Bond-Taylor, S., Ilse, M., P’erez-Garc’ia, F., Salvatelli, V., Sharma, H., Meissen, F., Ranjit, M. P., Srivastav, S., Gong, J., Falck, F., Oktay, O., Thieme, A., Lungren, M. P., Wetscherek, M. T. A., Alvarez-Valle, J., and Hyland, S. L. (2024) · 2024
Closest in time.
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al. (2024) · 2024
Closest in time.
Are large language models bayesian? a martingale perspective on in-context learning
Falck, F., Wang, Z., and Holmes, C. C. (2024) · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Gemini Team (2024) · 2024
Closest in time.
Gemma 2: Improving open language models at a practical size
Gemma Team (2024) · 2024
Closest in time.
Language model cascades: Token-level uncertainty and beyond
Gupta, N., Narasimhan, H., Jitkrittum, W., Rawat, A. S., Menon, A. K., and Kumar, S. (2024) · 2024
Closest in time.
Experts Don’t Cheat: Learning What You Don’t Know By Predicting Pairs
Johnson, D. D., Tarlow, D., Duvenaud, D., and Maddison, C. J. (2024) · 2024
Closest in time.
Predictive Uncertainty Quantification via Risk Decompositions for Strictly Proper Scoring Rules
Kotelevskii, N. and Panov, M. (2024) · 2024
Closest in time.
Generating with confidence: Uncertainty quantification for black-box large language models
Lin, Z., Trivedi, S., and Sun, J. (2024) · 2024
Closest in time.
SILO language models: Isolating legal risk in a nonparametric datastore
Min, S., Gururangan, S., Wallace, E., Shi, W., Hajishirzi, H., Smith, N. A., and Zettlemoyer, L. (2024) · 2024
Closest in time.
Language Models with Conformal Factuality Guarantees
Mohri, C. and Hashimoto, T. (2024) · 2024
Closest in time.
Kernel language entropy: Fine-grained uncertainty quantification for llms from semantic similarities
Nikitin, A., Kossen, J., Gal, Y., and Marttinen, P. (2024) · 2024
Closest in time.
Gpt-4o system card
OpenAI (2024) · 2024
Closest in time.
Modern bayesian experimental design
Rainforth, T., Foster, A., Ivanova, D. R., and Bickford Smith, F. (2024) · 2024
Closest in time.
Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs
Xiong, M., Hu, Z., Lu, X., Li, Y., Fu, J., He, J., and Hooi, B. (2024) · 2024
Closest in time.