Fetching the paper…
Reading the bibliography…
While past works have shown how uncertainty quantification can be applied to large language model (LLM) outputs, the question of whether resulting uncertainty guarantees still hold within sub-groupings of data remains open.
Probabilistic outputs for support vector machines and comparisons to regularized likelihood methods
John C Platt · 1999
Earlier work this paper cites.
Obtaining calibrated probability estimates from decision trees and naive bayesian classifiers
Bianca Zadrozny and Charles Elkan · 2001
Earlier work this paper cites.
A tutorial on conformal prediction
Glenn Shafer and Vladimir Vovk · 2008
Earlier work this paper cites.
Reliability, sufficiency, and the decomposition of proper scores
Jochen Bröcker · 2009
Earlier work this paper cites.
On calibration of modern neural networks
Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q Weinberger · 2017
Earlier work this paper cites.
Multicalibration: Calibration for the (computationally-identifiable) masses
Ursula Hébert-Johnson, Michael Kim, Omer Reingold, and Guy Rothblum · 2018
Earlier work this paper cites.
Bias in bios: A case study of semantic representation bias in a high-stakes setting
Maria De-Arteaga, Alexey Romanov, Hanna Wallach, Jennifer Chayes, Christian Borgs, Alexandra Chouldechova, Sahin Geyik, Krishnaram Kenthapadi, and Adam Tauman Kalai · 2019
Earlier work this paper cites.
Natural questions: a benchmark for question answering research
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, et al · 2019
Earlier work this paper cites.
Conformalized quantile regression
Yaniv Romano, Evan Patterson, and Emmanuel Candes · 2019
Earlier work this paper cites.
Low-degree multicalibration
Parikshit Gopalan, Michael P Kim, Mihir A Singhal, and Shengjia Zhao · 2022
Earlier work this paper cites.
Nested conformal prediction and quantile out-of-bag ensemble methods
Chirag Gupta, Arun K Kuchibhotla, and Aaditya Ramdas · 2022
Earlier work this paper cites.
Batch multivalid conformal prediction
Christopher Jung, Georgy Noarov, Ramya Ramalingam, and Aaron Roth · 2022
Earlier work this paper cites.
Language models (mostly) know what they know
Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield-Dodds, Nova DasSarma, Eli Tran-Johnson, et al · 2022
Earlier work this paper cites.
Uncertain: Modern topics in uncertainty estimation
Aaron Roth · 2022
Earlier work this paper cites.
Self-consistency improves chain of thought reasoning in language models
Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V Le, Ed H Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou · 2022
Earlier work this paper cites.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Cited alongside, same era.
Rohan Anil, Andrew M Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, et al · 2023
Cited alongside, same era.
Sparks of artificial general intelligence: Early experiments with gpt-4
Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, et al · 2023
Cited alongside, same era.
Lm-polygraph: Uncertainty estimation for language models
Ekaterina Fadeeva, Roman Vashurin, Akim Tsvigun, Artem Vazhentsev, Sergey Petrakov, Kirill Fedyanin, Daniil Vasilev, Elizaveta Goncharova, Alexander Panchenko, Maxim Panov, et al · 2023
Cited alongside, same era.
Yuqing Yang, Ethan Chern, Xipeng Qiu, Graham Neubig, and Pengfei Liu · 2023
Later among the works it cites.
Alignscore: Evaluating factual consistency with a unified alignment function
Yuheng Zha, Yichi Yang, Ruichen Li, and Zhiting Hu · 2023
Later among the works it cites.
Alleviating hallucinations of large language models through induced hallucinations
Yue Zhang, Leyang Cui, Wei Bi, and Shuming Shi · 2023
Later among the works it cites.
Factbench: A dynamic benchmark for in-the-wild language model factuality evaluation
Farima Fatahi Bayat, Lechen Zhang, Sheza Munir, and Lu Wang · 2024
Closest in time.
Large language model validity via enhanced conformal prediction methods
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Isaac Gibbs, John J Cherian, and Emmanuel J Candès · 2023
Cited alongside, same era.
Lei Huang, Weijiang Yu, Weitao Ma, Weihong Zhong, Zhangyin Feng, Haotian Wang, Qianglong Chen, Weihua Peng, Xiaocheng Feng, Bing Qin, et al · 2023
Cited alongside, same era.
Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al · 2023
Cited alongside, same era.
Conformal prediction with large language models for multi-choice question answering
Bhawesh Kumar, Charlie Lu, Gauri Gupta, Anil Palepu, David Bellamy, Ramesh Raskar, and Andrew Beam · 2023
Cited alongside, same era.
Evaluating verifiability in generative search engines
Nelson F Liu, Tianyi Zhang, and Percy Liang · 2023
Cited alongside, same era.
Selfcheckgpt: Zero-resource black-box hallucination detection for generative large language models
Potsawee Manakul, Adian Liusie, and Mark Gales · 2023
Cited alongside, same era.
Factscore: Fine-grained atomic evaluation of factual precision in long form text generation
Sewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis, Wen-tau Yih, Pang Koh, Mohit Iyyer, Luke Zettlemoyer, and Hannaneh Hajishirzi · 2023
Cited alongside, same era.
Conformal language modeling
Victor Quach, Adam Fisch, Tal Schuster, Adam Yala, Jae Ho Sohn, Tommi S Jaakkola, and Regina Barzilay · 2023
Cited alongside, same era.
John J Cherian, Isaac Gibbs, and Emmanuel J Candès · 2024
Closest in time.
Multicalibration for confidence scoring in llms
Gianluca Detommaso, Martin Bertran, Riccardo Fogliato, and Aaron Roth · 2024
Closest in time.
A survey of confidence estimation and calibration in large language models
Jiahui Geng, Fengyu Cai, Yuxia Wang, Heinz Koeppl, Preslav Nakov, and Iryna Gurevych · 2024
Closest in time.
Language models with conformal factuality guarantees
Christopher Mohri and Tatsunori Hashimoto · 2024
Closest in time.
Anthony Sicilia, Hyunwoo Kim, Khyathi Raghavi Chandu, Malihe Alikhani, and Jack Hessel · 2024
Closest in time.
Veriscore: Evaluating the factuality of verifiable claims in long-form text generation
Yixiao Song, Yekyung Kim, and Mohit Iyyer · 2024
Closest in time.
Benchmarking uncertainty quantification methods for large language models with lm-polygraph
Roman Vashurin, Ekaterina Fadeeva, Artem Vazhentsev, Akim Tsvigun, Daniil Vasilev, Rui Xing, Abdelrahman Boda Sadallah, Lyudmila Rvanova, Sergey Petrakov, Alexander Panchenko, et al · 2024
Closest in time.
Fine-grained self-endorsement improves factuality and reasoning
Ante Wang, Linfeng Song, Baolin Peng, Ye Tian, Lifeng Jin, Haitao Mi, Jinsong Su, and Dong Yu · 2024
Closest in time.
Long-form factuality in large language models
Jerry Wei, Chengrun Yang, Xinying Song, Yifeng Lu, Nathan Zixia Hu, Jie Huang, Dustin Tran, Daiyi Peng, Ruibo Liu, Da Huang, et al · 2024
Closest in time.
Calibrating llms for text-to-sql parsing by leveraging sub-clause frequencies
Terrance Liu, Shuyi Wang, Daniel Preotiuc-Pietro, Yash Chandarana, and Chirag Gupta · 2025
Closest in time.
Nonparametric masked language modeling
Sewon Min, Weijia Shi, Mike Lewis, Xilun Chen, Wen-tau Yih, Hannaneh Hajishirzi, and Luke Zettlemoyer · 2097
Closest in time.