Fetching the paper…
Reading the bibliography…
Uncertainty quantification in Large Language Models (LLMs) is crucial for applications where safety and reliability are important.
Prediction and entropy of printed english
C. E. Shannon · 1951
Earlier work this paper cites.
Spectral graph theory , volume 92
F. R. Chung · 1997
Earlier work this paper cites.
Diffusion kernels on graphs and other discrete structures
R. I. Kondor and J. Lafferty · 2002
Earlier work this paper cites.
Information theory, inference and learning algorithms
D. J. MacKay · 2003
Earlier work this paper cites.
Evaluating wordnet-based measures of lexical semantic relatedness
A. Budanitsky and G. Hirst · 2006
Earlier work this paper cites.
The effective rank: A measure of effective dimensionality
O. Roy and M. Vetterli · 2007
Earlier work this paper cites.
A tutorial on spectral clustering
U. Von Luxburg · 2007
Earlier work this paper cites.
Accuracy-rejection curves (ARCs) for comparing classification methods with a reject option
M. S. A. Nadeem, J.-D. Zucker, and B. Hanczar · 2009
Earlier work this paper cites.
Matrix analysis
R. A. Horn and C. R. Johnson · 2012
Earlier work this paper cites.
Dropout as a Bayesian approximation: Representing model uncertainty in deep learning
Y. Gal and Z. Ghahramani · 2016
Earlier work this paper cites.
Geometry of quantum states: an introduction to quantum entanglement
I. Bengtsson and K. Życzkowski · 2017
Earlier work this paper cites.
TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension
M. Joshi, E. Choi, D. S. Weld, and L. Zettlemoyer · 2017
Earlier work this paper cites.
Simple and scalable predictive uncertainty estimation using deep ensembles
B. Lakshminarayanan, A. Pritzel, and C. Blundell · 2017
Earlier work this paper cites.
The Cambridge encyclopedia of the English language
D. Crystal · 2018
Earlier work this paper cites.
Know what you don’t know: Unanswerable questions for SQuAD
P. Rajpurkar, R. Jia, and P. Liang · 2018
Earlier work this paper cites.
Mathematical foundations of quantum mechanics: New edition , volume 53
J. Von Neumann · 2018
Earlier work this paper cites.
Calibration of encoder decoder models for neural machine translation
A. Kumar and S. Sarawagi · 2019
Earlier work this paper cites.
Natural questions: a benchmark for question answering research
T. Kwiatkowski, J. Palomaki, O. Redfield, M. Collins, A. Parikh, C. Alberti, D. Epstein, I. Polosukhin, J. Devlin, K. Lee, et al · 2019
Earlier work this paper cites.
Can you trust your model’s uncertainty? evaluating predictive uncertainty under dataset shift
Y. Ovadia, E. Fertig, J. Ren, Z. Nado, D. Sculley, S. Nowozin, J. Dillon, B. Lakshminarayanan, and J. Snoek · 2019
Earlier work this paper cites.
Quantifying uncertainties in natural language processing tasks
Y. Xiao and W. Y. Wang · 2019
Earlier work this paper cites.
Calibration of pre-trained transformers
S. Desai and G. Durrett · 2020
Earlier work this paper cites.
Controlled hallucinations: Learning to generate faithfully from noisy data
K. Filippova · 2020
Earlier work this paper cites.
Deberta: Decoding-enhanced bert with disentangled attention
P. He, X. Liu, J. Gao, and W. Chen · 2020
Earlier work this paper cites.
Simple and principled uncertainty estimation with deterministic deep learning via distance awareness
J. Liu, Z. Lin, S. Padhy, D. Tran, T. Bedrax Weiss, and B. Lakshminarayanan · 2020
Earlier work this paper cites.
Uncertainty estimation in autoregressive structured prediction
A. Malinin and M. Gales · 2020
Earlier work this paper cites.
On faithfulness and factuality in abstractive summarization
J. Maynez, S. Narayan, B. Bohnet, and R. McDonald · 2020
Cited alongside, same era.
S. J. Mielke, A. Szlam, Y.-L. Boureau, and E. Dinan · 2020
Cited alongside, same era.
Matérn Gaussian processes on graphs
V. Borovitskiy, I. Azangulov, A. Terenin, P. Mostowsky, M. Deisenroth, and N. Durrande · 2021
Cited alongside, same era.
Unsolved problems in ML safety
D. Hendrycks, N. Carlini, J. Schulman, and J. Steinhardt · 2021
Cited alongside, same era.
How can we know when language models know? on the calibration of language models for question answering
Z. Jiang, J. Araki, H. Ding, and G. Neubig · 2021
A. Q. Jiang, A. Sablayrolles, A. Mensch, C. Bamford, D. S. Chaplot, D. d. l. Casas, F. Bressand, G. Lengyel, G. Lample, L. Saulnier, et al · 2023
Later among the works it cites.
Chatgpt for good? on opportunities and challenges of large language models for education
E. Kasneci, K. Seßler, S. Küchemann, M. Bannert, D. Dementieva, F. Fischer, U. Gasser, G. Groh, S. Günnemann, E. Hüllermeier, et al · 2023
Later among the works it cites.
Vne: An effective method for improving deep representation by manipulating eigenvalue distribution
J. Kim, S. Kang, D. Hwang, J. Shin, and W. Rhee · 2023
Later among the works it cites.
BioASQ-QA: A manually curated corpus for biomedical question answering
A. Krithara, A. Nentidis, K. Bougiatiotis, and G. Paliouras · 2023
Later among the works it cites.
L. Kuhn, Y. Gal, and S. Farquhar · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Are NLP models really able to solve simple math word problems?
A. Patel, S. Bhattamishra, and N. Goyal · 2021
Cited alongside, same era.
On hallucination and predictive uncertainty in conditional language generation
Y. Xiao and W. Y. Wang · 2021
Cited alongside, same era.
Introduction to quantum information science II lecture notes, 2022
S. Aaronson · 2022
Cited alongside, same era.
Information theory with kernel methods
F. Bach · 2022
Cited alongside, same era.
Discovering latent knowledge in language models without supervision, 2022
C. Burns, H. Ye, D. Klein, and J. Steinhardt · 2022
Cited alongside, same era.
Language models (mostly) know what they know
S. Kadavath, T. Conerly, A. Askell, T. Henighan, D. Drain, E. Perez, N. Schiefer, Z. Hatfield-Dodds, N. DasSarma, E. Tran-Johnson, et al · 2022
Cited alongside, same era.
Coderl: Mastering code generation through pretrained models and deep reinforcement learning
H. Le, Y. Wang, A. D. Gotmare, S. Savarese, and S. C. H. Hoi · 2022
Cited alongside, same era.
Later among the works it cites.
Chain-of-knowledge: Grounding large language models via dynamic knowledge adapting over heterogeneous sources
X. Li, R. Zhao, Y. K. Chia, B. Ding, S. Joty, S. Poria, and L. Bing · 2023
Later among the works it cites.
Generating with confidence: Uncertainty quantification for black-box large language models
Z. Lin, S. Trivedi, and J. Sun · 2023
Later among the works it cites.
S. Liu, L. Xing, and J. Zou · 2023
Later among the works it cites.
SelfCheckGPT: Zero-resource black-box hallucination detection for generative large language models
P. Manakul, A. Liusie, and M. J. Gales · 2023
Later among the works it cites.
Deep deterministic uncertainty: A new simple baseline
J. Mukhoti, A. Kirsch, J. van Amersfoort, P. H. Torr, and Y. Gal · 2023
Later among the works it cites.
GPT-4 technical report
OpenAI · 2023
Later among the works it cites.
V. Quach, A. Fisch, T. Schuster, A. Yala, J. H. Sohn, T. S. Jaakkola, and R. Barzilay · 2023
Later among the works it cites.
Self-evaluation improves selective generation in large language models
J. Ren, Y. Zhao, T. Vu, P. J. Liu, and B. Lakshminarayanan · 2023
Later among the works it cites.
Gemini: a family of highly capable multimodal models
G. Team, R. Anil, S. Borgeaud, Y. Wu, J.-B. Alayrac, J. Yu, R. Soricut, J. Schalkwyk, A. M. Dai, A. Hauth, et al · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, et al · 2023
Later among the works it cites.
N. Varshney, W. Yao, H. Zhang, J. Chen, and D. Yu · 2023
Later among the works it cites.
Bayesian low-rank adaptation for large language models
A. X. Yang, M. Robeyns, X. Wang, and L. Aitchison · 2023
Later among the works it cites.
Representation engineering: A top-down approach to ai transparency
A. Zou, L. Phan, S. Chen, J. Campbell, P. Guo, R. Ren, A. Pan, X. Yin, M. Mazeika, A.-K. Dombrowski, et al · 2023
Later among the works it cites.
How many opinions does your llm have? improving uncertainty estimation in nlg
L. Aichberger, K. Schweighofer, M. Ielanskyi, and S. Hochreiter · 2024
Closest in time.
Personal Communication, 2024
S. Farquhar, J. Kossen, L. Kuhn, and Y. Gal · 2024
Closest in time.
Unfamiliar finetuning examples control how language models hallucinate
K. Kang, E. Wallace, C. Tomlin, A. Kumar, and S. Levine · 2024
Closest in time.
Inference-time intervention: Eliciting truthful answers from a language model
K. Li, O. Patel, F. Viégas, H. Pfister, and M. Wattenberg · 2024
Closest in time.
Simple probes can catch sleeper agents, 2024
M. MacDiarmid, T. Maxwell, N. Schiefer, J. Mu, J. Kaplan, D. Duvenaud, S. Bowman, A. Tamkin, E. Perez, M. Sharma, C. Denison, and E. Hubinger · 2024
Closest in time.
Luq: Long-text uncertainty quantification for llms
C. Zhang, F. Liu, M. Basaldella, and N. Collier · 2024
Closest in time.