Fetching the paper…
Reading the bibliography…
In recent years, Large Language Models (LLMs) have become fundamental to a broad spectrum of artificial intelligence applications.
Knowing with certainty: The appropriateness of extreme confidence
Baruch Fischhoff, Paul Slovic, and Sarah Lichtenstein · 1977
Earlier work this paper cites.
Reasons for confidence
Asher Koriat, Sarah Lichtenstein, and Baruch Fischhoff · 1980
Earlier work this paper cites.
Pitfalls of in-domain uncertainty estimation and ensembling in deep learning, 2021
Arsenii Ashukha, Alexander Lyzhov, Dmitry Molchanov, and Dmitry Vetrov · 2002
Earlier work this paper cites.
Jesse Dodge, Gabriel Ilharco, Roy Schwartz, Ali Farhadi, Hannaneh Hajishirzi, and Noah Smith · 2002
Earlier work this paper cites.
A noisy-channel model of human sentence comprehension under uncertain input
Roger Levy · 2008
Earlier work this paper cites.
Aleatory or epistemic? does it matter?
Armen Der Kiureghian and Ove Ditlevsen · 2009
Earlier work this paper cites.
Towards maximizing the representation gap between in-domain & out-of-distribution examples, 2021
Jay Nandy, Wynne Hsu, and Mong Li Lee · 2010
Earlier work this paper cites.
Risk, unexpected uncertainty, and estimation uncertainty: Bayesian learning in unstable settings
Elise Payzan-LeNestour and Peter Bossaerts · 2011
Earlier work this paper cites.
The communicative function of ambiguity in language
Steven T. Piantadosi, Harry J. Tily, and Edward Gibson · 2011
Earlier work this paper cites.
Hyter: Meaning-equivalent semantics for translation evaluation
Markus Dreyer and Daniel Marcu · 2012
Earlier work this paper cites.
On calibration of modern neural networks
Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q Weinberger · 2017
Earlier work this paper cites.
What uncertainties do we need in bayesian deep learning for computer vision?
Alex Kendall and Yarin Gal · 2017
Earlier work this paper cites.
Simple and scalable predictive uncertainty estimation using deep ensembles
Balaji Lakshminarayanan, Alexander Pritzel, and Charles Blundell · 2017
Earlier work this paper cites.
Relational inductive biases, deep learning, and graph networks, 2018
Peter W. Battaglia, Jessica B. Hamrick, Victor Bapst, Alvaro Sanchez-Gonzalez, Vinicius Zambaldi, Mateusz Malinowski, Andrea Tacchetti, David Raposo, Adam Santoro, Ryan Faulkner, Caglar Gulcehre, Francis Song, Andrew Ballard, Justin Gilmer, George Dahl, Ashish Vaswani, Kelsey Allen, Charles Nash, Victoria Langston, Chris Dyer, Nicolas Heess, Daan Wierstra, Pushmeet Kohli, Matt Botvinick, Oriol Vinyals, Yujia Li, and Razvan Pascanu · 2018
Earlier work this paper cites.
Data statements for natural language processing: Toward mitigating system bias and enabling better science
Emily M. Bender and Batya Friedman · 2018
Earlier work this paper cites.
Hierarchical neural story generation
Angela Fan, Mike Lewis, and Yann Dauphin · 2018
Earlier work this paper cites.
Diachronic word embeddings and semantic shifts: a survey
Andrey Kutuzov, Lilja Øvrelid, Terrence Szymanski, and Erik Velldal · 2018
Earlier work this paper cites.
Neural relation extraction via inner-sentence noise reduction and transfer learning
Tianyi Liu, Xinsong Zhang, Wanhao Zhou, and Weijia Jia · 2018
Earlier work this paper cites.
Predictive uncertainty estimation via prior networks
Andrey Malinin and Mark Gales · 2018
Earlier work this paper cites.
Analyzing uncertainty in neural machine translation, 2018
Myle Ott, Michael Auli, David Grangier, and Marc’Aurelio Ranzato · 2018
Earlier work this paper cites.
Quantifying uncertainties in natural language processing tasks, 2018
Yijun Xiao and William Yang Wang · 2018
Earlier work this paper cites.
A variational dirichlet framework for out-of-distribution detection, 2019
Wenhu Chen, Yilin Shen, Hongxia Jin, and William Wang · 2019
Earlier work this paper cites.
Characterizing sources of uncertainty to proxy calibration and disambiguate annotator and data bias
Asma Ghandeharioun, Brian Eoff, Brendan Jou, and Rosalind Picard · 2019
Earlier work this paper cites.
Matthias Hein, Maksym Andriushchenko, and Julian Bitterwolf · 2019
Earlier work this paper cites.
Reverse kl-divergence training of prior networks: Improved uncertainty and adversarial robustness
Andrey Malinin and Mark Gales · 2019
Earlier work this paper cites.
Measuring calibration in deep learning
Jeremy Nixon, Michael W. Dusenberry, Linchuan Zhang, Ghassen Jerfel, and Dustin Tran · 2019
Earlier work this paper cites.
On NMT search errors and model errors: Cat got your tongue?
Felix Stahlberg and Bill Byrne · 2019
Earlier work this paper cites.
Aleatoric uncertainty estimation with test-time augmentation for medical image segmentation with convolutional neural networks
Guotai Wang, Wenqi Li, Michael Aertsen, Jan Deprest, Sébastien Ourselin, and Tom Vercauteren · 2019
Earlier work this paper cites.
Language (technology) is power: A critical survey of “bias” in NLP
Su Lin Blodgett, Solon Barocas, Hal Daumé III, and Hanna Wallach · 2020
Earlier work this paper cites.
Is MAP decoding all you need? the inadequacy of the mode in neural machine translation
Bryan Eikema and Wilker Aziz · 2020
Earlier work this paper cites.
The curious case of neural text degeneration
Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi · 2020
Earlier work this paper cites.
Calibrated language model fine-tuning for in- and out-of-distribution data
Lingkai Kong, Haoming Jiang, Yuchen Zhuang, Jie Lyu, Tuo Zhao, and Chao Zhang · 2020
Earlier work this paper cites.
If beam search is the answer, what was the question?
Clara Meister, Ryan Cotterell, and Tim Vieira · 2020
Earlier work this paper cites.
Uncertainty-based rejection wrappers for black-box classifiers
José Mena, Oriol Pujol, and Jordi Vitrià · 2020
Earlier work this paper cites.
Calibrating deep neural networks using focal loss
Jishnu Mukhoti, Viveka Kulharia, Amartya Sanyal, Stuart Golodetz, Philip H. S. Torr, and Puneet K. Dokania · 2020
Earlier work this paper cites.
Uncertainty quantification and deep ensembles
Rahul Rahaman and Alexandre H Thiery · 2020
Cited alongside, same era.
A review of uncertainty quantification in deep learning: Techniques, applications and challenges
Moloud Abdar, Farhad Pourpanah, Sadiq Hussain, Dana Rezazadegan, Li Liu, Mohammad Ghavamzadeh, Paul Fieguth, Xiaochun Cao, Abbas Khosravi, U. Rajendra Acharya, Vladimir Makarenkov, and Saeid Nahavandi · 2021
Cited alongside, same era.
A general language assistant as a laboratory for alignment
Amanda Askell, Yuntao Bai, Anna Chen, Dawn Drain, Deep Ganguli, Tom Henighan, Andy Jones, Nicholas Joseph, Ben Mann, Nova DasSarma, et al · 2021
Cited alongside, same era.
A survey of uncertainty in deep neural networks
Jakob Gawlikowski, Cedrique Rovile Njieutcheu Tassi, Mohsin Ali, Jongseok Lee, Matthias Humt, Jianxiang Feng, Anna Kruspe, Rudolph Triebel, Peter Jung, Ribana Roscher, et al · 2021
Cited alongside, same era.
Aleatoric and epistemic uncertainty in machine learning: an introduction to concepts and methods
Pseudo outlier exposure for out-of-distribution detection using pretrained transformers, 2023
Jaeyoung Kim, Kyuheon Jung, Dongbin Na, Sion Jang, Eunbin Park, and Sungchul Choi · 2023
Later among the works it cites.
Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation, 2023
Lorenz Kuhn, Yarin Gal, and Sebastian Farquhar · 2023
Later among the works it cites.
Deup: Direct epistemic uncertainty prediction, 2023
Salem Lahlou, Moksh Jain, Hadi Nekoei, Victor Ion Butoi, Paul Bertin, Jarrid Rector-Brooks, Maksym Korablyov, and Yoshua Bengio · 2023
Later among the works it cites.
Generating with confidence: Uncertainty quantification for black-box large language models
Zhen Lin, Shubhendu Trivedi, and Jimeng Sun · 2023
Later among the works it cites.
Selfcheckgpt: Zero-resource black-box hallucination detection for generative large language models, 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Eyke Hüllermeier and Willem Waegeman · 2021
Cited alongside, same era.
How can we know when language models know? on the calibration of language models for question answering
Zhengbao Jiang, Jun Araki, Haibo Ding, and Graham Neubig · 2021
Cited alongside, same era.
Hannah Kirk, Yennie Jun, Haider Iqbal, Elias Benussi, Filippo Volpin, Frederic A. Dreyer, Aleksandar Shtedritski, and Yuki M. Asano · 2021
Cited alongside, same era.
A survey on uncertainty estimation in deep learning classification systems from a bayesian perspective
José Mena, Oriol Pujol, and Jordi Vitrià · 2021
Cited alongside, same era.
Improving paragraph-level question generation with extended answer network and uncertainty-aware beam search
Hongwei Zeng, Zhuo Zhi, Jun Liu, and Bifan Wei · 2021
Cited alongside, same era.
Uncertainty analysis in ontology-based knowledge representation
Sanjay Kumar Anand and Suresh Kumar · 2022
Cited alongside, same era.
Exploring length generalization in large language models, 2022
Cem Anil, Yuhuai Wu, Anders Andreassen, Aitor Lewkowycz, Vedant Misra, Vinay Ramasesh, Ambrose Slone, Guy Gur-Ari, Ethan Dyer, and Behnam Neyshabur · 2022
Cited alongside, same era.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al · 2022
Cited alongside, same era.
Potsawee Manakul, Adian Liusie, and Mark J. F. Gales · 2023
Later among the works it cites.
Strength in numbers: Estimating confidence of large language models by prompt agreement
Gwenyth Portillo Wightman, Alexandra Delucia, and Mark Dredze · 2023
Later among the works it cites.
Distributional preference learning: Understanding and accounting for hidden context in rlhf
Anand Siththaranjan, Cassidy Laidlaw, and Dylan Hadfield-Menell · 2023
Later among the works it cites.
Katherine Tian, Eric Mitchell, Allan Zhou, Archit Sharma, Rafael Rafailov, Huaxiu Yao, Chelsea Finn, and Christopher D Manning · 2023
Later among the works it cites.
Efficient out-of-domain detection for sequence to sequence models
Artem Vazhentsev, Akim Tsvigun, Roman Vashurin, Sergey Petrakov, Daniil Vasilev, Maxim Panov, Alexander Panchenko, and Artem Shelmanov · 2023
Later among the works it cites.
Calibration in deep learning: A survey of the state-of-the-art
Cheng Wang · 2023
Later among the works it cites.
Self-consistency improves chain of thought reasoning in language models, 2023
Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou · 2023
Later among the works it cites.
Self-evaluation guided beam search for reasoning, 2023
Yuxi Xie, Kenji Kawaguchi, Yiran Zhao, Xu Zhao, Min-Yen Kan, Junxian He, and Qizhe Xie · 2023
Later among the works it cites.
Can llms express their uncertainty? an empirical evaluation of confidence elicitation in llms
Miao Xiong, Zhiyuan Hu, Xinyang Lu, Yifei Li, Jie Fu, Junxian He, and Bryan Hooi · 2023
Later among the works it cites.
Do large language models know what they don’t know?
Zhangyue Yin, Qiushi Sun, Qipeng Guo, Jiawen Wu, Xipeng Qiu, and Xuanjing Huang · 2023
Later among the works it cites.
LLMaAA: Making large language models as active annotators
Ruoyu Zhang, Yanzeng Li, Yongliang Ma, Ming Zhou, and Lei Zou · 2023
Later among the works it cites.
A survey of large language models
Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, et al · 2023
Later among the works it cites.
Navigating the grey area: Expressions of overconfidence and uncertainty in language models, 2023
Kaitlyn Zhou, Dan Jurafsky, and Tatsunori Hashimoto · 2023
Later among the works it cites.
Internalinspector i 2 i^{2} : Robust confidence estimation in llms through internal states, 2024
Mohammad Beigi, Ying Shen, Runing Yang, Zihao Lin, Qifan Wang, Ankith Mohan, Jianfeng He, Ming Jin, Chang-Tien Lu, and Lifu Huang · 2024
Closest in time.
Direct preference optimization with unobserved preference heterogeneity
Keertana Chidambaram, Karthik Vinay Seetharaman, and Vasilis Syrgkanis · 2024
Closest in time.
Sycophancy to subterfuge: Investigating reward-tampering in large language models, 2024
Carson Denison, Monte MacDiarmid, Fazl Barez, David Duvenaud, Shauna Kravec, Samuel Marks, Nicholas Schiefer, Ryan Soklaski, Alex Tamkin, Jared Kaplan, Buck Shlegeris, Samuel R. Bowman, Ethan Perez, and Evan Hubinger · 2024
Closest in time.
Data augmentation using llms: Data perspectives, learning paradigms and challenges, 2024
Bosheng Ding, Chengwei Qin, Ruochen Zhao, Tianze Luo, Xinze Li, Guizhen Chen, Wenhan Xia, Junjie Hu, Anh Tuan Luu, and Shafiq Joty · 2024
Closest in time.
How does beam search improve span-level confidence estimation in generative sequence labeling?
Kazuma Hashimoto, Iftekhar Naim, and Karthik Raman · 2024
Closest in time.
A survey on uncertainty quantification methods for deep learning, 2024
Wenchong He, Zhe Jiang, Tingsong Xiao, Zelin Xu, and Yukun Li · 2024
Closest in time.
Investigating data contamination for pre-training language models, 2024
Minhao Jiang, Ken Ziyu Liu, Ming Zhong, Rylan Schaeffer, Siru Ouyang, Jiawei Han, and Sanmi Koyejo · 2024
Closest in time.
How good are llms at out-of-distribution detection?, 2024
Bo Liu, Liming Zhan, Zexin Lu, Yujie Feng, Lei Xue, and Xiao-Ming Wu · 2024
Closest in time.
Principled rlhf from heterogeneous feedback via personalization and preference aggregation
Chanwoo Park, Mingyang Liu, Kaiqing Zhang, and Asuman Ozdaglar · 2024
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn · 2024
Closest in time.
The effect of sampling temperature on problem solving in large language models, 2024
Matthew Renze and Erhan Guven · 2024
Closest in time.
Junkang Wu, Yuexiang Xie, Zhengyi Yang, Jiancan Wu, Jiawei Chen, Jinyang Gao, Bolin Ding, Xiang Wang, and Xiangnan He · 2024
Closest in time.
Can llms express their uncertainty? an empirical evaluation of confidence elicitation in llms, 2024
Miao Xiong, Zhiyuan Hu, Xinyang Lu, Yifei Li, Jie Fu, Junxian He, and Bryan Hooi · 2024
Closest in time.
Large language models for anomaly and out-of-distribution detection: A survey, 2024
Ruiyao Xu and Kaize Ding · 2024
Closest in time.
Relying on the unreliable: The impact of language models’ reluctance to express uncertainty, 2024
Kaitlyn Zhou, Jena D. Hwang, Xiang Ren, and Maarten Sap · 2024
Closest in time.
Data augmentation for low-resource neural machine translation
Marzieh Fadaee, Arianna Bisazza, and Christof Monz · 2090
Closest in time.
A survey on large language model based autonomous agents
Lei Wang, Chen Ma, Xueyang Feng, Zeyu Zhang, Hao Yang, Jingsen Zhang, Zhiyuan Chen, Jiakai Tang, Xu Chen, Yankai Lin, Wayne Xin Zhao, Zhewei Wei, and Jirong Wen · 2095
Closest in time.