Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) require robust confidence estimation, particularly in critical domains like healthcare and law where unreliable outputs can lead to significant consequences.
Are humans good intuitive statisticians after all? rethinking some conclusions from the literature on judgment under uncertainty
Leda Cosmides and John Tooby. 1996 · 1996
Earlier work this paper cites.
Learning and making decisions when costs and probabilities are both unknown
Bianca Zadrozny and Charles Elkan. 2001 · 2001
Earlier work this paper cites.
Conformal prediction with neural networks
Harris Papadopoulos, Volodya Vovk, and Alex Gammerman. 2007 · 2007
Earlier work this paper cites.
Accuracy-rejection curves (arcs) for comparing classification methods with a reject option
Malik Sajjad Ahmed Nadeem, Jean-Daniel Zucker, and Blaise Hanczar. 2009 · 2009
Earlier work this paper cites.
On the foundations of noise-free selective classification
Ran El-Yaniv et al. 2010 · 2010
Earlier work this paper cites.
Selective classification for deep neural networks
Yonatan Geifman and Ran El-Yaniv. 2017 · 2017
Earlier work this paper cites.
On calibration of modern neural networks
Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q. Weinberger. 2017 · 2017
Earlier work this paper cites.
Race: Large-scale reading comprehension dataset from examinations
Guokun Lai, Qizhe Xie, Hanxiao Liu, Yiming Yang, and Eduard Hovy. 2017 · 2017
Earlier work this paper cites.
Why we need new evaluation metrics for nlg
Jekaterina Novikova, Ondřej Dušek, Amanda Cercas Curry, and Verena Rieser. 2017 · 2017
Earlier work this paper cites.
Survey of the state of the art in natural language generation: Core tasks, applications and evaluation
Albert Gatt and Emiel Krahmer. 2018 · 2018
Earlier work this paper cites.
Beyond temperature scaling: Obtaining well-calibrated multi-class probabilities with dirichlet calibration
Meelis Kull, Miquel Perello Nieto, Markus Kängsepp, Telmo Silva Filho, Hao Song, and Peter Flach. 2019 · 2019
Earlier work this paper cites.
CoQA: A conversational question answering challenge
Siva Reddy, Danqi Chen, and Christopher D. Manning. 2019 · 2019
Earlier work this paper cites.
CommonsenseQA: A question answering challenge targeting commonsense knowledge
Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. 2019 · 2019
Earlier work this paper cites.
Qasc: A dataset for question answering via sentence composition
Tushar Khot, Peter Clark, Michal Guerquin, Peter Jansen, and Ashish Sabharwal. 2020 · 2020
Earlier work this paper cites.
Mix-n-match : Ensemble and compositional methods for uncertainty calibration in deep learning
Jize Zhang, Bhavya Kailkhura, and T. Yong-Jin Han. 2020 · 2020
Earlier work this paper cites.
Bertscore: Evaluating text generation with bert
Tianyi Zhang*, Varsha Kishore*, Felix Wu*, Kilian Q. Weinberger, and Yoav Artzi. 2020 · 2020
Earlier work this paper cites.
Deberta: Decoding-enhanced bert with disentangled attention
Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen. 2021 · 2021
Earlier work this paper cites.
What disease does this patient have? a large-scale open domain question answering dataset from medical exams
Di Jin, Eileen Pan, Nassim Oufattole, Wei-Hung Weng, Hanyi Fang, and Peter Szolovits. 2021 · 2021
Cited alongside, same era.
Meta-cal: Well-controlled post-hoc calibration by ranking
Xingchen Ma and Matthew B. Blaschko. 2021 · 2021
Cited alongside, same era.
Towards better selective classification
Leo Feng, Mohamed Osama Ahmed, Hossein Hajimirsadeghi, and Amir Abdi. 2022 · 2022
Cited alongside, same era.
Truthfulqa: Measuring how models mimic human falsehoods
Stephanie Lin, Jacob Hilton, and Owain Evans. 2022 · 2022
Cited alongside, same era.
The internal state of an LLM knows when it‘s lying
Amos Azaria and Tom Mitchell. 2023 · 2023
Cited alongside, same era.
Selectively answering ambiguous questions
Large language model validity via enhanced conformal prediction methods
John Cherian, Isaac Gibbs, and Emmanuel Candes. 2024 · 2024
Later among the works it cites.
Longchao Da, Tiejin Chen, Lu Cheng, and Hua Wei. 2024 · 2024
Later among the works it cites.
Shifting attention to relevance: Towards the predictive uncertainty quantification of free-form large language models
Jinhao Duan, Hao Cheng, Shiqi Wang, Alex Zavalny, Chenan Wang, Renjing Xu, Bhavya Kailkhura, and Kaidi Xu. 2024 · 2024
Later among the works it cites.
Conformal alignment: Knowing when to trust foundation models with guarantees
Yu Gui, Ying Jin, and Zhimei Ren. 2024 · 2024
Later among the works it cites.
Unveiling llm evaluation focused on metrics: Challenges and solutions
Taojun Hu and Xiao-Hua Zhou. 2024 · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jeremy Cole, Michael Zhang, Daniel Gillick, Julian Eisenschlos, Bhuwan Dhingra, and Jacob Eisenstein. 2023 · 2023
Cited alongside, same era.
Optimal strategies for reject option classifiers
Vojtech Franc, Daniel Prusa, and Vaclav Voracek. 2023 · 2023
Cited alongside, same era.
A survey of language model confidence estimation and calibration
Jiahui Geng, Fengyu Cai, Yuxia Wang, Heinz Koeppl, Preslav Nakov, and Iryna Gurevych. 2023 · 2023
Cited alongside, same era.
Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation
Lorenz Kuhn, Yarin Gal, and Sebastian Farquhar. 2023 · 2023
Cited alongside, same era.
Efficient memory management for large language model serving with pagedattention
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, and Ion Stoica. 2023 · 2023
Cited alongside, same era.
Inference-time intervention: Eliciting truthful answers from a language model
Kenneth Li, Oam Patel, Fernanda Viégas, Hanspeter Pfister, and Martin Wattenberg. 2023 · 2023
Cited alongside, same era.
SelfCheckGPT: Zero-resource black-box hallucination detection for generative large language models
Potsawee Manakul, Adian Liusie, and Mark Gales. 2023 · 2023
Cited alongside, same era.
Later among the works it cites.
Uncertainty in language models: Assessment through rank-calibration
Xinmeng Huang, Shuo Li, Mengxin Yu, Matteo Sesia, Hamed Hassani, Insup Lee, Osbert Bastani, and Edgar Dobriban. 2024 · 2024
Later among the works it cites.
Selective generation for controllable language models
Minjae Lee, Kyungmin Kim, Taesoo Kim, and Sangdon Park. 2024 · 2024
Later among the works it cites.
From generation to judgment: Opportunities and challenges of llm-as-a-judge
Dawei Li, Bohan Jiang, Liangjie Huang, Alimohammad Beigi, Chengshuai Zhao, Zhen Tan, Amrita Bhattacharjee, Yuxuan Jiang, Canyu Chen, Tianhao Wu, et al. 2024 · 2024
Later among the works it cites.
Contextualized sequence likelihood: Enhanced confidence scores for natural language generation
Zhen Lin, Shubhendu Trivedi, and Jimeng Sun. 2024a · 2024
Later among the works it cites.
Language models with conformal factuality guarantees
Christopher Mohri and Tatsunori Hashimoto. 2024 · 2024
Later among the works it cites.
Llm evaluators recognize and favor their own generations
Arjun Panickssery, Samuel R Bowman, and Shi Feng. 2024 · 2024
Later among the works it cites.
Combining confidence elicitation and sample-based methods for uncertainty quantification in misinformation mitigation
Mauricio Rivera, Jean-François Godbout, Reihaneh Rabbany, and Kellin Pelrine. 2024 · 2024
Later among the works it cites.
Judgebench: A benchmark for evaluating llm-based judges
Sijun Tan, Siyuan Zhuang, Kyle Montgomery, William Y Tang, Alejandro Cuadron, Chenguang Wang, Raluca Ada Popa, and Ion Stoica. 2024 · 2024
Later among the works it cites.
Benchmarking uncertainty quantification methods for large language models with lm-polygraph
Roman Vashurin, Ekaterina Fadeeva, Artem Vazhentsev, Lyudmila Rvanova, Akim Tsvigun, Daniil Vasilev, Rui Xing, Abdelrahman Boda Sadallah, Kirill Grishchenkov, Sergey Petrakov, et al. 2024 · 2024
Later among the works it cites.
Can LLMs express their uncertainty? an empirical evaluation of confidence elicitation in LLMs
Miao Xiong, Zhiyuan Hu, Xinyang Lu, YIFEI LI, Jie Fu, Junxian He, and Bryan Hooi. 2024 · 2024
Later among the works it cites.
Mitigating llm hallucinations via conformal abstention
Yasin Abbasi Yadkori, Ilja Kuzborskij, David Stutz, András György, Adam Fisch, Arnaud Doucet, Iuliya Beloshapka, Wei-Hung Weng, Yao-Yuan Yang, Csaba Szepesvári, et al. 2024 · 2024
Later among the works it cites.
An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, et al. 2024 · 2024
Later among the works it cites.