Fetching the paper…
Reading the bibliography…
We present an automatic large language model (LLM) conversion approach that produces uncertainty-aware LLMs capable of estimating uncertainty with every prediction.
Quizbowl: The Case for Incremental Question Answering
Rodriguez, P.; Feng, S.; Iyyer, M.; He, H.; and Boyd-Graber, J. 2021 · 1904
Earlier work this paper cites.
Nemo: a toolkit for building ai applications using neural modules
Kuchaiev, O.; Li, J.; Nguyen, H.; Hrinchuk, O.; Leary, R.; Ginsburg, B.; Kriman, S.; Beliaev, S.; Lavrukhin, V.; Cook, J.; et al. 2019 · 1909
Earlier work this paper cites.
Posing fair generalization tasks for natural language inference
Geiger, A.; Cases, I.; Karttunen, L.; and Potts, C. 2019 · 1911
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Srivastava, N.; Hinton, G.; Krizhevsky, A.; Sutskever, I.; and Salakhutdinov, R. 2014 · 1958
Earlier work this paper cites.
Estimating the mean and variance of the target probability distribution
Nix, D. A.; and Weigend, A. S. 1994 · 1994
Earlier work this paper cites.
BLEURT: Learning robust metrics for text generation
Sellam, T.; Das, D.; and Parikh, A. P. 2020 · 2004
Earlier work this paper cites.
Selective question answering under domain shift
Kamath, A.; Jia, R.; and Liang, P. 2020 · 2006
Earlier work this paper cites.
A survey of answer extraction techniques in factoid question answering
Wang, M.; et al. 2006 · 2006
Earlier work this paper cites.
Overview of ResPubliQA 2009: Question answering evaluation over European legislation
Peñas, A.; Forner, P.; Sutcliffe, R.; Rodrigo, Á.; Forăscu, C.; Alegria, I.; Giampiccolo, D.; Moreau, N.; and Osenova, P. 2010 · 2009
Earlier work this paper cites.
On the Foundations of Noise-free Selective Classification
El-Yaniv, R.; et al. 2010 · 2010
Earlier work this paper cites.
A framework for merging and ranking of answers in DeepQA
Gondek, D. C.; Lally, A.; Kalyanpur, A.; Murdock, J. W.; Duboué, P. A.; Zhang, L.; Pan, Y.; Qiu, Z. M.; and Welty, C. 2012 · 2012
Earlier work this paper cites.
Deep gaussian processes
Damianou, A.; and Lawrence, N. D. 2013 · 2013
Earlier work this paper cites.
Weight uncertainty in neural network
Blundell, C.; Cornebise, J.; Kavukcuoglu, K.; and Wierstra, D. 2015 · 2015
Earlier work this paper cites.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Zhu, Y.; Kiros, R.; Zemel, R.; Salakhutdinov, R.; Urtasun, R.; Torralba, A.; and Fidler, S. 2015 · 2015
Earlier work this paper cites.
Dropout as a bayesian approximation: Representing model uncertainty in deep learning
Gal, Y.; and Ghahramani, Z. 2016 · 2016
Cited alongside, same era.
SQuAD: 100,000+ Questions for Machine Comprehension of Text
Rajpurkar, P.; Zhang, J.; Lopyrev, K.; and Liang, P. 2016 · 2016
Cited alongside, same era.
Google’s Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
Wu, Y.; Schuster, M.; Chen, Z.; Le, Q. V.; Norouzi, M.; Macherey, W.; Krikun, M.; Cao, Y.; Gao, Q.; Macherey, K.; Klingner, J.; Shah, A.; Johnson, M.; Liu, X.; Łukasz Kaiser; Gouws, S.; Kato, Y.; Kudo, T.; Kazawa, H.; Stevens, K.; Kurian, G.; Patil, N.; Wang, W.; Young, C.; Smith, J.; Riesa, J.; Rudnick, A.; Vinyals, O.; Corrado, G.; Hughes, M.; and Dean, J. 2016 · 2016
Cited alongside, same era.
Selective classification for deep neural networks
Geifman, Y.; and El-Yaniv, R. 2017 · 2017
Cited alongside, same era.
On calibration of modern neural networks
Guo, C.; Pleiss, G.; Sun, Y.; and Weinberger, K. Q. 2017 · 2017
Capsa: A Unified Framework for Quantifying Risk in Deep Neural Networks
Lolla, S.; Elistratov, I.; Perez, A.; Ahmadi, E.; Rus, D.; and Amini, A. 2022 · 2022
Later among the works it cites.
Post-hoc Uncertainty Learning using a Dirichlet Meta-Model
Shen, M.; Bu, Y.; Sattigeri, P.; Ghosh, S.; Das, S.; and Wornell, G. 2022 · 2022
Later among the works it cites.
Large language models are reasoners with self-verification
Weng, Y.; Zhu, M.; He, S.; Liu, K.; and Zhao, J. 2022 · 2022
Later among the works it cites.
Capsa Software Library
Amini, A.; et al. 2023 · 2023
Closest in time.
Quantifying Uncertainty in Answers from any Language Model and Enhancing their Trustworthiness
Chen, J.; and Mueller, J. 2023 · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Simple and scalable predictive uncertainty estimation using deep ensembles
Lakshminarayanan, B.; Pritzel, A.; and Blundell, C. 2017 · 2017
Cited alongside, same era.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2018 · 2018
Cited alongside, same era.
Confidence modeling for neural semantic parsing
Dong, L.; Quirk, C.; and Lapata, M. 2018 · 2018
Cited alongside, same era.
Know What You Don’t Know: Unanswerable Questions for SQuAD
Rajpurkar, P.; Jia, R.; and Liang, P. 2018 · 2018
Cited alongside, same era.
How can we know when language models know? on the calibration of language models for question answering
Jiang, Z.; Araki, J.; Ding, H.; and Neubig, G. 2021 · 2021
Cited alongside, same era.
Truthfulqa: Measuring how models mimic human falsehoods
Lin, S.; Hilton, J.; and Evans, O. 2021 · 2021
Cited alongside, same era.
Dropconnect is effective in modeling uncertainty of bayesian deep networks
Mobiny, A.; Yuan, P.; Moulik, S. K.; Garg, N.; Wu, C. C.; and Van Nguyen, H. 2021 · 2021
Cited alongside, same era.
DoLa: Decoding by Contrasting Layers Improves Factuality in Large Language Models
Chuang, Y.-S.; Xie, Y.; Luo, H.; Kim, Y.; Glass, J.; and He, P. 2023 · 2023
Closest in time.
Human Uncertainty in Concept-Based AI Systems
Collins, K. M.; Barker, M.; Zarlenga, M. E.; Raman, N.; Bhatt, U.; Jamnik, M.; Sucholutsky, I.; Weller, A.; and Dvijotham, K. 2023 · 2023
Closest in time.
Generating with Confidence: Uncertainty Quantification for Black-box Large Language Models
Lin, Z.; Trivedi, S.; and Sun, J. 2023 · 2023
Closest in time.
Capsa: A Unified Framework for Quantifying Risk in Deep Neural Networks
Lolla, S.; Elistratov, I.; Perez, A.; Ahmadi, E.; Rus, D.; and Amini, A. 2023 · 2023
Closest in time.
SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models
Manakul, P.; Liusie, A.; and Gales, M. J. F. 2023 · 2023
Closest in time.
SelfCheck: Using LLMs to Zero-Shot Check Their Own Step-by-Step Reasoning
Miao, N.; Teh, Y. W.; and Rainforth, T. 2023 · 2023
Closest in time.
Quach, V.; Fisch, A.; Schuster, T.; Yala, A.; Sohn, J. H.; Jaakkola, T. S.; and Barzilay, R. 2023 · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H.; Martin, L.; Stone, K.; Albert, P.; Almahairi, A.; Babaei, Y.; Bashlykov, N.; Batra, S.; Bhargava, P.; Bhosale, S.; et al. 2023 · 2023
Closest in time.