Fetching the paper…
Reading the bibliography…
Language Models (LMs) have shown promising performance in natural language generation.
Verification of forecasts expressed in terms of probability
Glenn W Brier. 1950 · 1950
Earlier work this paper cites.
The evaluation of economic forecasts
Jacob A Mincer and Victor Zarnowitz. 1969 · 1969
Earlier work this paper cites.
Elicitation of personal probabilities and expectations
Leonard J Savage. 1971 · 1971
Earlier work this paper cites.
Calibration of probabilities: The state of the art
Sarah Lichtenstein, Baruch Fischhoff, and Lawrence D Phillips. 1977 · 1975
Earlier work this paper cites.
Perplexity—a measure of the difficulty of speech recognition tasks
Fred Jelinek, Robert L Mercer, Lalit R Bahl, and James K Baker. 1977 · 1977
Earlier work this paper cites.
The comparison and evaluation of forecasters
Morris H DeGroot and Stephen E Fienberg. 1983 · 1983
Earlier work this paper cites.
Scoring rules and the evaluation of probabilities
Robert L Winkler, Javier Munoz, José L Cervera, José M Bernardo, Gail Blattenberger, Joseph B Kadane, Dennis V Lindley, Allan H Murphy, Robert M Oliver, and David Ríos-Insua. 1996 · 1996
Earlier work this paper cites.
Probabilistic outputs for support vector machines and comparisons to regularized likelihood methods
John Platt. 1999 · 1999
Earlier work this paper cites.
Confidence estimation methods for neural networks: A practical comparison
Georgios Papadopoulos, Peter J Edwards, and Alan F Murray. 2001 · 2001
Earlier work this paper cites.
Obtaining calibrated probability estimates from decision trees and naive bayesian classifiers
Bianca Zadrozny and Charles Elkan. 2001 · 2001
Earlier work this paper cites.
Transforming classifier scores into accurate multiclass probability estimates
Bianca Zadrozny and Charles Elkan. 2002 · 2002
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
METEOR: An automatic metric for MT evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie. 2005 · 2005
Earlier work this paper cites.
Predicting good probabilities with supervised learning
Alexandru Niculescu-Mizil and Rich Caruana. 2005 · 2005
Earlier work this paper cites.
Statistical comparisons of classifiers over multiple data sets
Janez Demšar. 2006 · 2006
Earlier work this paper cites.
Strictly proper scoring rules, prediction, and estimation
Tilmann Gneiting and Adrian E Raftery. 2007 · 2007
Earlier work this paper cites.
Natural language processing with Python: analyzing text with the natural language toolkit
Steven Bird, Ewan Klein, and Edward Loper. 2009 · 2009
Earlier work this paper cites.
Introduction to Nonparametric Estimation
Alexandre B Tsybakov. 2009 · 2009
Earlier work this paper cites.
Regression modeling strategies with applications to linear models, logistic and ordinal regression, and survival analysis
Frank E Harrell. 2015 · 2015
Earlier work this paper cites.
Obtaining well calibrated probabilities using bayesian binning
Mahdi Pakdaman Naeini, Gregory Cooper, and Milos Hauskrecht. 2015 · 2015
Earlier work this paper cites.
Dropout as a bayesian approximation: Representing model uncertainty in deep learning
Yarin Gal and Zoubin Ghahramani. 2016 · 2016
Earlier work this paper cites.
Squad: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Earlier work this paper cites.
On calibration of modern neural networks
Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q Weinberger. 2017 · 2017
Cited alongside, same era.
spaCy 2: Natural language understanding with Bloom embeddings, convolutional neural networks and incremental parsing
Matthew Honnibal and Ines Montani. 2017 · 2017
Cited alongside, same era.
TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension
Mandar Joshi, Eunsol Choi, Daniel Weld, and Luke Zettlemoyer. 2017 · 2017
Cited alongside, same era.
Simple and scalable predictive uncertainty estimation using deep ensembles
Balaji Lakshminarayanan, Alexander Pritzel, and Charles Blundell. 2017 · 2017
Cited alongside, same era.
A discourse-aware attention model for abstractive summarization of long documents
Arman Cohan, Franck Dernoncourt, Doo Soon Kim, Trung Bui, Seokhwan Kim, Walter Chang, and Nazli Goharian. 2018 · 2018
Cited alongside, same era.
On hallucination and predictive uncertainty in conditional language generation
Yijun Xiao and William Yang Wang. 2021 · 2021
Later among the works it cites.
Calibrate before use: Improving few-shot performance of language models
Zihao Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh. 2021 · 2021
Later among the works it cites.
Language models (mostly) know what they know
Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield-Dodds, Nova DasSarma, Eli Tran-Johnson, et al. 2022 · 2022
Later among the works it cites.
Teaching models to express their uncertainty in words
Stephanie Lin, Jacob Hilton, and Owain Evans. 2022 · 2022
Later among the works it cites.
Reducing conversational agents’ overconfidence through linguistic calibration
Sabrina J Mielke, Arthur Szlam, Emily Dinan, and Y-Lan Boureau. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Enhancing the reliability of out-of-distribution image detection in neural networks
Shiyu Liang, Yixuan Li, and R. Srikant. 2018 · 2018
Cited alongside, same era.
Deep k-nearest neighbors: Towards confident, interpretable and robust deep learning
Nicolas Papernot and Patrick McDaniel. 2018 · 2018
Cited alongside, same era.
Deep bayesian bandits showdown: An empirical comparison of bayesian deep networks for thompson sampling
Carlos Riquelme, George Tucker, and Jasper Snoek. 2018 · 2018
Cited alongside, same era.
Beyond temperature scaling: Obtaining well-calibrated multi-class probabilities with dirichlet calibration
Meelis Kull, Miquel Perello Nieto, Markus Kängsepp, Telmo Silva Filho, Hao Song, and Peter Flach. 2019 · 2019
Cited alongside, same era.
Verified uncertainty calibration
Ananya Kumar, Percy S Liang, and Tengyu Ma. 2019 · 2019
Cited alongside, same era.
Natural questions: A benchmark for question answering research
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov. 2019 · 2019
Cited alongside, same era.
Measuring calibration in deep learning
Jeremy Nixon, Michael W Dusenberry, Linchuan Zhang, Ghassen Jerfel, and Dustin Tran. 2019 · 2019
Cited alongside, same era.
Training language models to follow instructions with human feedback
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe. 2022 · 2022
Later among the works it cites.
Re-examining calibration: The case of question answering
Chenglei Si, Chen Zhao, Sewon Min, and Jordan Boyd-Graber. 2022 · 2022
Later among the works it cites.
Jiuhai Chen and Jonas Mueller. 2023 · 2023
Later among the works it cites.
Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation
Lorenz Kuhn, Yarin Gal, and Sebastian Farquhar. 2023 · 2023
Later among the works it cites.
T-cal: An optimal test for the calibration of predictive models
Donghwan Lee, Xinmeng Huang, Hamed Hassani, and Edgar Dobriban. 2023 · 2023
Later among the works it cites.
Generating with confidence: Uncertainty quantification for black-box large language models
Zhen Lin, Shubhendu Trivedi, and Jimeng Sun. 2023 · 2023
Later among the works it cites.
Selfcheckgpt: Zero-resource black-box hallucination detection for generative large language models
Potsawee Manakul, Adian Liusie, and Mark JF Gales. 2023 · 2023
Later among the works it cites.
Llama access request form - meta ai
Meta. 2023 · 2023
Later among the works it cites.
Gpt-4 technical report
OpenAI. 2023 · 2023
Later among the works it cites.
Quantifying uncertainty in natural language explanations of large language models
Sree Harsha Tanneru, Chirag Agarwal, and Himabindu Lakkaraju. 2023 · 2023
Later among the works it cites.
Just ask for calibration: Strategies for eliciting calibrated confidence scores from language models fine-tuned with human feedback
Katherine Tian, Eric Mitchell, Allan Zhou, Archit Sharma, Rafael Rafailov, Huaxiu Yao, Chelsea Finn, and Christopher Manning. 2023 · 2023
Later among the works it cites.
A comprehensive survey on summarization techniques
Padma Jyothi Uppalapati, Madhavi Dabbiru, · K. Venkata Rao, Omer F. Rana, Rajiv Misra, Alexander Pfeiffer, Luigi Troiano, Nishtha Kesswani, and K. Venkata Rao. 2023 · 2023
Later among the works it cites.
INSIDE: LLMs’ internal states retain the power of hallucination detection
Chao Chen, Kai Liu, Ze Chen, Yi Gu, Yue Wu, Mingyuan Tao, Zhihang Fu, and Jieping Ye. 2024 · 2024
Closest in time.
One-shot safety alignment for large language models via optimal dualization
Xinmeng Huang, Shuo Li, Edgar Dobriban, Osbert Bastani, Hamed Hassani, and Dongsheng Ding. 2024 · 2024
Closest in time.
Can LLMs express their uncertainty? an empirical evaluation of confidence elicitation in LLMs
Miao Xiong, Zhiyuan Hu, Xinyang Lu, YIFEI LI, Jie Fu, Junxian He, and Bryan Hooi. 2024 · 2024
Closest in time.
Luq: Long-text uncertainty quantification for llms
Caiqi Zhang, Fangyu Liu, Marco Basaldella, and Nigel Collier. 2024 · 2024
Closest in time.