Fetching the paper…
Reading the bibliography…
We develop new conformal inference methods for obtaining validity guarantees on the output of large language models (LLMs).
Optimality and degeneracy in linear programming
A. Charnes · 1952
Earlier work this paper cites.
Inductive confidence machines for regression
H. Papadopoulos, K. Proedrou, V. Vovk, and A. Gammerman · 2002
Earlier work this paper cites.
Linear programming: Theory and extensions , volume 2
G. B. Dantzig and M. N. Thapa · 2003
Earlier work this paper cites.
Algorithmic Learning in a Random World
V. Vovk, A. Gammerman, and G. Shafer · 2005
Earlier work this paper cites.
Conditional validity of inductive conformal predictors
V. Vovk · 2012
Earlier work this paper cites.
Overview of the medical question answering task at trec 2017 liveqa
A. B. Abacha, E. Agichtein, Y. Pinter, and D. Demner-Fushman · 2017
Earlier work this paper cites.
Multicalibration: Calibration for the (computationally-identifiable) masses
U. Hébert-Johnson, M. Kim, O. Reingold, and G. Rothblum · 2018
Earlier work this paper cites.
Bridging the gap between consumers’ medication questions and trusted answers
A. Ben Abacha, Y. Mrabet, M. Sharp, T. Goodwin, S. E. Shooshan, and D. Demner-Fushman · 2019
Earlier work this paper cites.
Learning optimal conformal classifiers
D. Stutz, K. D. Dvijotham, A. T. Cemgil, and A. Doucet · 2021
Earlier work this paper cites.
Learn then test: Calibrating predictive algorithms to achieve risk control
A. N. Angelopoulos, S. Bates, E. J. Candès, M. I. Jordan, and L. Lei · 2022
Earlier work this paper cites.
Low-degree multicalibration
P. Gopalan, M. P. Kim, M. A. Singhal, and S. Zhao · 2022
Earlier work this paper cites.
Language models (mostly) know what they know
S. Kadavath, T. Conerly, A. Askell, T. Henighan, D. Drain, E. Perez, N. Schiefer, Z. Hatfield-Dodds, N. DasSarma, E. Tran-Johnson, et al · 2022
Cited alongside, same era.
The internal state of an llm knows when it’s lying
A. Azaria and T. Mitchell · 2023
Cited alongside, same era.
Conformal prediction with conditional guarantees
I. Gibbs, J. J. Cherian, and E. J. Candès · 2023
Cited alongside, same era.
Conformal prediction with large language models for multi-choice question answering
B. Kumar, C. Lu, G. Gupta, A. Palepu, D. Bellamy, R. Raskar, and A. Beam · 2023
Cited alongside, same era.
Selfcheckgpt: Zero-resource black-box hallucination detection for generative large language models
Multicalibration for confidence scoring in llms
G. Detommaso, M. Bertran, R. Fogliato, and A. Roth · 2024
Closest in time.
Olaph: Improving factuality in biomedical long-form question answering
M. Jeong, H. Hwang, C. Yoon, T. Lee, and J. Kang · 2024
Closest in time.
K-qa: A real-world medical q&a benchmark
I. Manes, N. Ronn, D. Cohen, R. I. Ber, Z. Horowitz-Kugler, and G. Stanovsky · 2024
Closest in time.
Language models with conformal factuality guarantees
C. Mohri and T. Hashimoto · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
P. Manakul, A. Liusie, and M. J. Gales · 2023
Cited alongside, same era.
Factscore: Fine-grained atomic evaluation of factual precision in long form text generation
S. Min, K. Krishna, X. Lyu, M. Lewis, W.-t. Yih, P. W. Koh, M. Iyyer, L. Zettlemoyer, and H. Hajishirzi · 2023
Cited alongside, same era.
Large language models encode clinical knowledge
K. Singhal, S. Azizi, T. Tu, S. S. Mahdavi, J. Wei, H. W. Chung, N. Scales, A. Tanwani, H. Cole-Lewis, S. Pfohl, et al · 2023
Cited alongside, same era.
N. Varshney, W. Yao, H. Zhang, J. Chen, and D. Yu · 2023
Cited alongside, same era.
Posterior conformal prediction
Y. Zhang and E. J. Candès · 2023
Cited alongside, same era.
Conformal risk control
A. N. Angelopoulos, S. Bates, A. Fisch, L. Lei, and T. Schuster · 2024
Cited alongside, same era.
D. Nadeau, M. Kroutikov, K. McNeil, and S. Baribeau · 2024
Closest in time.
Chatgpt active user count reaches new milestone, November 2023
J. Porter · 2024
Closest in time.
Conformal language modeling
V. Quach, A. Fisch, T. Schuster, A. Yala, J. H. Sohn, T. S. Jaakkola, and R. Barzilay · 2024
Closest in time.
Avianca airline lawsuit uses chatgpt, May 2023
B. Weiser · 2024
Closest in time.
Benchmarking llms via uncertainty quantification, 2024
F. Ye, M. Yang, J. Pang, L. Wang, D. F. Wong, E. Yilmaz, S. Shi, and Z. Tu · 2024
Closest in time.
The limits of distribution-free conditional predictive inference
R. F. Barber, E. J. Candès, A. Ramdas, and R. J. Tibshirani · 2049
Closest in time.