Fetching the paper…
Reading the bibliography…
When using large language models (LLMs) in high-stakes applications, we need to know when we can trust their predictions.
Calibration of probabilities: The state of the art
Sarah Lichtenstein, Baruch Fischhoff, and Lawrence D Phillips · 1977
Earlier work this paper cites.
Calibration and probability judgements: Conceptual and methodological issues
Gideon Keren · 1991
Earlier work this paper cites.
Learning from demonstration
Stefan Schaal · 1996
Earlier work this paper cites.
Unskilled and unaware of it: how difficulties in recognizing one’s own incompetence lead to inflated self-assessments
Justin Kruger and David Dunning · 1999
Earlier work this paper cites.
Probabilistic outputs for support vector machines and comparisons to regularized likelihood methods
John Platt et al · 1999
Earlier work this paper cites.
Unskilled and unaware–but why? a reply to krueger and mueller (2002)
Justin Kruger and David Dunning · 2002
Earlier work this paper cites.
Learning question classifiers
Xin Li and Dan Roth · 2002
Earlier work this paper cites.
Information theory, inference, and learning algorithms
David John Cameron MacKay · 2004
Earlier work this paper cites.
Pattern recognition and machine learning
Christopher M Bishop · 2006
Earlier work this paper cites.
Strictly proper scoring rules, prediction, and estimation
Tilmann Gneiting and Adrian E Raftery · 2007
Earlier work this paper cites.
Updating methods improved the performance of a clinical prediction model in new patients
KJM Janssen, KGM Moons, CJ Kalkman, DE Grobbee, and Y Vergouwe · 2008
Earlier work this paper cites.
Semeval-2012 task 7: Choice of plausible alternatives: An evaluation of commonsense causal reasoning
Andrew S. Gordon, Zornitsa Kozareva, and Melissa Roemmele · 2011
Earlier work this paper cites.
Truth is a lie: Crowd truth and the seven myths of human annotation
Lora Aroyo and Chris Welty · 2015
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning · 2015
Earlier work this paper cites.
Obtaining well calibrated probabilities using bayesian binning
Mahdi Pakdaman Naeini, Gregory F. Cooper, and Milos Hauskrecht · 2015
Earlier work this paper cites.
On calibration of modern neural networks
Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q. Weinberger · 2017
Earlier work this paper cites.
Fixing weight decay regularization in adam
Ilya Loshchilov and Frank Hutter · 2017
Earlier work this paper cites.
Crowdsourcing multiple choice science questions
Johannes Welbl, Nelson F. Liu, and Matt Gardner · 2017
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord · 2018
Earlier work this paper cites.
Looking beyond the surface: A challenge set for reading comprehension over multiple sentences
Daniel Khashabi, Snigdha Chaturvedi, Michael Roth, Shyam Upadhyay, and Dan Roth · 2018
Earlier work this paper cites.
Can a suit of armor conduct electricity? a new dataset for open book question answering
Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal · 2018
Cited alongside, same era.
Prolific. ac—a subject pool for online experiments
Stefan Palan and Christian Schitter · 2018
Cited alongside, same era.
MathQA: Towards interpretable math word problem solving with operation-based formalisms
Aida Amini, Saadia Gabriel, Shanchuan Lin, Rik Koncel-Kedziorski, Yejin Choi, and Hannaneh Hajishirzi · 2019
Cited alongside, same era.
Piqa: Reasoning about physical commonsense in natural language
Yonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao, and Yejin Choi · 2019
Cited alongside, same era.
Boolq: Exploring the surprising difficulty of natural yes/no questions
Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova · 2019
Teaching models to express their uncertainty in words
Stephanie C. Lin, Jacob Hilton, and Owain Evans · 2022
Later among the works it cites.
Fine-tuning language models via epistemic neural networks
Ian Osband, Seyed Mohammad Asghari, Benjamin Van Roy, Nat McAleese, John Aslanides, and Geoffrey Irving · 2022
Later among the works it cites.
Uncalibrated models can improve human-ai collaboration
Kailas Vodrahalli, Tobias Gerstenberg, and James Y Zou · 2022
Later among the works it cites.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Later among the works it cites.
The internal state of an llm knows when its lying
Amos Azaria and Tom M. Mitchell · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
The commitmentbank: Investigating projection in naturally occurring discourse
Marie-Catherine De Marneffe, Mandy Simons, and Judith Tonhauser · 2019
Cited alongside, same era.
Cosmos qa: Machine reading comprehension with contextual commonsense reasoning
Lifu Huang, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi · 2019
Cited alongside, same era.
Adversarial nli: A new benchmark for natural language understanding
Yixin Nie, Adina Williams, Emily Dinan, Mohit Bansal, Jason Weston, and Douwe Kiela · 2019
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Köpf, Edward Yang, Zach DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala · 2019
Cited alongside, same era.
Sentence-bert: Sentence embeddings using siamese bert-networks
Nils Reimers and Iryna Gurevych · 2019
Cited alongside, same era.
Winogrande: An adversarial winograd schema challenge at scale
Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi · 2019
Cited alongside, same era.
Commonsenseqa: A question answering challenge targeting commonsense knowledge
Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant · 2019
Cited alongside, same era.
Later among the works it cites.
Learning personalized decision support policies
Umang Bhatt, Valerie Chen, Katherine M Collins, Parameswaran Kamalaruban, Emma Kallina, Adrian Weller, and Ameet Talwalkar · 2023
Later among the works it cites.
Can large language models be an alternative to human evaluations?
Cheng-Han Chiang and Hung yi Lee · 2023
Later among the works it cites.
Human uncertainty in concept-based ai systems
Katherine Maeve Collins, Matthew Barker, Mateo Espinosa Zarlenga, Naveen Raman, Umang Bhatt, Mateja Jamnik, Ilia Sucholutsky, Adrian Weller, and Krishnamurthy Dvijotham · 2023
Later among the works it cites.
Using logprobs, Dec 2023
James Hills and Shyamal Anadkat · 2023
Later among the works it cites.
Albert Qiaochu Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de Las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, L’elio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed · 2023
Later among the works it cites.
Lorenz Kuhn, Yarin Gal, and Sebastian Farquhar · 2023
Later among the works it cites.
Alphafold predictions are valuable hypotheses and accelerate but do not replace experimental structure determination
Thomas C Terwilliger, Dorothee Liebschner, Tristan I Croll, Christopher J Williams, Airlie J McCoy, Billy K Poon, Pavel V Afonine, Robert D Oeffner, Jane S Richardson, Randy J Read, et al · 2023
Later among the works it cites.
Can llms express their uncertainty? an empirical evaluation of confidence elicitation in llms
Miao Xiong, Zhiyuan Hu, Xinyang Lu, Yifei Li, Jie Fu, Junxian He, and Bryan Hooi · 2023
Later among the works it cites.
Do large language models know what they don’t know?
Zhangyue Yin, Qiushi Sun, Qipeng Guo, Jiawen Wu, Xipeng Qiu, and Xuanjing Huang · 2023
Later among the works it cites.
R-tuning: Teaching large language models to refuse unknown questions
Hanning Zhang, Shizhe Diao, Yong Lin, Yi R Fung, Qing Lian, Xingyao Wang, Yangyi Chen, Heng Ji, and Tong Zhang · 2023
Later among the works it cites.
Livecodebench: Holistic and contamination free evaluation of large language models for code
Naman Jain, King Han, Alex Gu, Wen-Ding Li, Fanjia Yan, Tianjun Zhang, Sida Wang, Armando Solar-Lezama, Koushik Sen, and Ion Stoica · 2024
Closest in time.
Softmax probabilities (mostly) predict large language model correctness on multiple-choice q&a
Benjamin Plaut, Khanh Nguyen, and Tu Trinh · 2024
Closest in time.
Gemini: A family of highly capable multimodal models, 2024
Gemini Team · 2024
Closest in time.
Calibrating large language models using their generations only
Dennis Ulmer, Martin Gubri, Hwaran Lee, Sangdoo Yun, and Seong Joon Oh · 2024
Closest in time.