Fetching the paper…
Reading the bibliography…
As large language models (LLMs) are increasingly deployed in user-facing applications, building trust and maintaining safety by accurately quantifying a model's confidence in its prediction becomes even more important.
Verification of forecasts expressed in terms of probability
Glenn W Brier. 1950 · 1950
Earlier work this paper cites.
An introduction to the bootstrap
Bradley Efron and Robert J Tibshirani. 1994 · 1994
Earlier work this paper cites.
A density-based algorithm for discovering clusters in large spatial databases with noise
Martin Ester, Hans-Peter Kriegel, Jörg Sander, and Xiaowei Xu. 1996 · 1996
Earlier work this paper cites.
Probabilistic outputs for support vector machines and comparisons to regularized likelihood methods
John Platt et al. 1999 · 1999
Earlier work this paper cites.
The next decade in ai: four steps towards robust artificial intelligence
Gary Marcus. 2020 · 2002
Earlier work this paper cites.
Inductive confidence machines for regression
Harris Papadopoulos, Kostas Proedrou, Volodya Vovk, and Alex Gammerman. 2002 · 2002
Earlier work this paper cites.
Confidence estimation for machine translation
John Blatz, Erin Fitzgerald, George Foster, Simona Gandrabur, Cyril Goutte, Alex Kulesza, Alberto Sanchis, and Nicola Ueffing. 2004 · 2004
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Training a sentence-level machine translation confidence measure
Christopher Quirk. 2004 · 2004
Earlier work this paper cites.
Algorithmic learning in a random world , volume 29
Vladimir Vovk, Alexander Gammerman, and Glenn Shafer. 2005 · 2005
Earlier work this paper cites.
Wat zei je? detecting out-of-distribution translations with variational transformers
Tim Z Xiao, Aidan N Gomez, and Yarin Gal. 2020 · 2006
Earlier work this paper cites.
Practical bayesian optimization of machine learning algorithms
Jasper Snoek, Hugo Larochelle, and Ryan P Adams. 2012 · 2012
Earlier work this paper cites.
Density-based clustering based on hierarchical density estimates
Ricardo JGB Campello, Davoud Moulavi, and Jörg Sander. 2013 · 2013
Earlier work this paper cites.
Obtaining well calibrated probabilities using bayesian binning
Mahdi Pakdaman Naeini, Gregory Cooper, and Milos Hauskrecht. 2015 · 2015
Earlier work this paper cites.
Exploring prediction uncertainty in machine translation quality estimation
Daniel Beck, Lucia Specia, and Trevor Cohn. 2016 · 2016
Earlier work this paper cites.
On calibration of modern neural networks
Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q Weinberger. 2017 · 2017
Earlier work this paper cites.
Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension
Mandar Joshi, Eunsol Choi, Daniel S. Weld, and Luke Zettlemoyer. 2017 · 2017
Earlier work this paper cites.
An optimal transportation approach for assessing almost stochastic order
Eustasio Del Barrio, Juan A Cuesta-Albertos, and Carlos Matrán. 2018 · 2018
Earlier work this paper cites.
Fixing weight decay regularization in adam
Ilya Loshchilov and Frank Hutter. 2018 · 2018
Earlier work this paper cites.
Deep dominance - how to properly compare deep neural models
Rotem Dror, Segev Shlomov, and Roi Reichart. 2019 · 2019
Earlier work this paper cites.
The practical implementation of artificial intelligence technologies in medicine
Jianxing He, Sally L Baxter, Jie Xu, Jiming Xu, Xingtao Zhou, and Kang Zhang. 2019 · 2019
Earlier work this paper cites.
Quantifying the carbon emissions of machine learning
Alexandre Lacoste, Alexandra Luccioni, Victor Schmidt, and Thomas Dandres. 2019 · 2019
Earlier work this paper cites.
Energy usage reports: Environmental awareness as part of algorithmic accountability
Kadan Lottick, Silvia Susai, Sorelle A. Friedler, and Jonathan P. Wilson. 2019 · 2019
Earlier work this paper cites.
Can you trust your model’s uncertainty? evaluating predictive uncertainty under dataset shift
Yaniv Ovadia, Emily Fertig, Jie Ren, Zachary Nado, David Sculley, Sebastian Nowozin, Joshua Dillon, Balaji Lakshminarayanan, and Jasper Snoek. 2019 · 2019
Earlier work this paper cites.
Coqa: A conversational question answering challenge
Siva Reddy, Danqi Chen, and Christopher D Manning. 2019 · 2019
Earlier work this paper cites.
Sentence-bert: Sentence embeddings using siamese bert-networks
Nils Reimers and Iryna Gurevych. 2019 · 2019
Earlier work this paper cites.
Improving back-translation with uncertainty-based confidence estimation
Shuo Wang, Yang Liu, Chao Wang, Huanbo Luan, and Maosong Sun. 2019 · 2019
Earlier work this paper cites.
Experiment tracking with weights and biases
Lukas Biewald. 2020 · 2020
Earlier work this paper cites.
ELECTRA: pre-training text encoders as discriminators rather than generators
Kevin Clark, Minh-Thang Luong, Quoc V. Le, and Christopher D. Manning. 2020 · 2020
Earlier work this paper cites.
Calibration of pre-trained transformers
Shrey Desai and Greg Durrett. 2020 · 2020
Cited alongside, same era.
Comparison of deep learning models and various text pre-processing techniques for the toxic comments classification
Viera Maslej-Krešňáková, Martin Sarnovskỳ, Peter Butka, and Kristína Machová. 2020 · 2020
Cited alongside, same era.
Trust issues: Uncertainty estimation does not enable reliable ood detection on medical tabular data
Dennis Ulmer, Lotta Meijerink, and Giovanni Cinà. 2020 · 2020
Cited alongside, same era.
A gentle introduction to conformal prediction and distribution-free uncertainty quantification
Anastasios N Angelopoulos and Stephen Bates. 2021 · 2021
Cited alongside, same era.
Uncertainty as a form of transparency: Measuring, communicating, and using uncertainty
Umang Bhatt, Javier Antorán, Yunfeng Zhang, Q Vera Liao, Prasanna Sattigeri, Riccardo Fogliato, Gabrielle Melançon, Ranganath Krishnan, Jason Stanley, Omesh Tickoo, et al. 2021 · 2021
A diachronic perspective on user trust in ai under uncertainty
Shehzaad Dhuliawala, Vilém Zouhar, Mennatallah El-Assady, and Mrinmaya Sachan. 2023 · 2023
Later among the works it cites.
Shortcut learning of large language models in natural language understanding
Mengnan Du, Fengxiang He, Na Zou, Dacheng Tao, and Xia Hu. 2023 · 2023
Later among the works it cites.
Perspectives on the state and future of deep learning–2023
Micah Goldblum, Anima Anandkumar, Richard Baraniuk, Tom Goldstein, Kyunghyun Cho, Zachary C Lipton, Melanie Mitchell, Preetum Nakkiran, Max Welling, and Andrew Gordon Wilson. 2023 · 2023
Later among the works it cites.
Mehedi Hasan, Moloud Abdar, Abbas Khosravi, Uwe Aickelin, Pietro Lio, Ibrahim Hossain, Ashikur Rahman, and Saeid Nahavandi. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Uncertainty-aware machine translation evaluation
Taisiya Glushkova, Chrysoula Zerva, Ricardo Rei, and André F. T. Martins. 2021 · 2021
Cited alongside, same era.
Deberta: decoding-enhanced bert with disentangled attention
Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen. 2021 · 2021
Cited alongside, same era.
Formalizing trust in artificial intelligence: Prerequisites, causes and goals of human trust in ai
Alon Jacovi, Ana Marasović, Tim Miller, and Yoav Goldberg. 2021 · 2021
Cited alongside, same era.
How can we know When language models know? on the calibration of language models for question answering
Zhengbao Jiang, Jun Araki, Haibo Ding, and Graham Neubig. 2021 · 2021
Cited alongside, same era.
Question and answer test-train overlap in open-domain question answering datasets
Patrick S. H. Lewis, Pontus Stenetorp, and Sebastian Riedel. 2021 · 2021
Cited alongside, same era.
Uncertainty estimation in autoregressive structured prediction
Andrey Malinin and Mark J. F. Gales. 2021 · 2021
Cited alongside, same era.
CodeCarbon: Estimate and Track Carbon Emissions from Machine Learning Computing
Victor Schmidt, Kamal Goyal, Aditya Joshi, Boris Feld, Liam Conell, Nikolas Laskaris, Doug Blank, Jonathan Wilson, Sorelle Friedler, and Sasha Luccioni. 2021 · 2021
Cited alongside, same era.
Debertav3: Improving deberta using electra-style pre-training with gradient-disentangled embedding sharing
Pengcheng He, Jianfeng Gao, and Weizhu Chen. 2023 · 2023
Later among the works it cites.
On the richness of calibration
Benedikt Höltgen and Robert C Williamson. 2023 · 2023
Later among the works it cites.
Decomposing uncertainty for large language models through input clarification ensembling
Bairu Hou, Yujian Liu, Kaizhi Qian, Jacob Andreas, Shiyu Chang, and Yang Zhang. 2023 · 2023
Later among the works it cites.
Calibrating language models via augmented prompt ensembles
Mingjian Jiang, Yangjun Ruan, Sicong Huang, Saifei Liao, Silviu Pitis, Roger Baker Grosse, and Jimmy Ba. 2023 · 2023
Later among the works it cites.
Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation
Lorenz Kuhn, Yarin Gal, and Sebastian Farquhar. 2023 · 2023
Later among the works it cites.
DEUP: direct epistemic uncertainty prediction
Salem Lahlou, Moksh Jain, Hadi Nekoei, Victor Butoi, Paul Bertin, Jarrid Rector-Brooks, Maksym Korablyov, and Yoshua Bengio. 2023 · 2023
Later among the works it cites.
Generating with confidence: Uncertainty quantification for black-box large language models
Zhen Lin, Shubhendu Trivedi, and Jimeng Sun. 2023 · 2023
Later among the works it cites.
Deep deterministic uncertainty: A new simple baseline
Jishnu Mukhoti, Andreas Kirsch, Joost van Amersfoort, Philip HS Torr, and Yarin Gal. 2023 · 2023
Later among the works it cites.
How to catch an ai liar: Lie detection in black-box llms by asking unrelated questions
Lorenzo Pacchiardi, Alex J Chan, Sören Mindermann, Ilan Moscovitz, Alexa Y Pan, Yarin Gal, Owain Evans, and Jan Brauner. 2023 · 2023
Later among the works it cites.
Just ask for calibration: Strategies for eliciting calibrated confidence scores from language models fine-tuned with human feedback
Katherine Tian, Eric Mitchell, Allan Zhou, Archit Sharma, Rafael Rafailov, Huaxiu Yao, Chelsea Finn, and Christopher D. Manning. 2023 · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023 · 2023
Later among the works it cites.
Miles Turpin, Julian Michael, Ethan Perez, and Samuel R Bowman. 2023 · 2023
Later among the works it cites.
Intensive care unit physicians’ perspectives on artificial intelligence–based clinical decision support tools: Preimplementation survey study
Siri L van der Meijden, Anne AH de Hond, Patrick J Thoral, Ewout W Steyerberg, Ilse MJ Kant, Giovanni Cinà, and M Sesmu Arbous. 2023 · 2023
Later among the works it cites.
Hybrid uncertainty quantification for selective text classification in ambiguous tasks
Artem Vazhentsev, Gleb Kuzmin, Akim Tsvigun, Alexander Panchenko, Maxim Panov, Mikhail Burtsev, and Artem Shelmanov. 2023 · 2023
Later among the works it cites.
Fine-grained human feedback gives better rewards for language model training
Zeqiu Wu, Yushi Hu, Weijia Shi, Nouha Dziri, Alane Suhr, Prithviraj Ammanabrolu, Noah A. Smith, Mari Ostendorf, and Hannaneh Hajishirzi. 2023 · 2023
Later among the works it cites.
Can llms express their uncertainty? an empirical evaluation of confidence elicitation in llms
Miao Xiong, Zhiyuan Hu, Xinyang Lu, Yifei Li, Jie Fu, Junxian He, and Bryan Hooi. 2023 · 2023
Later among the works it cites.
Bayesian low-rank adaptation for large language models
Adam X Yang, Maxime Robeyns, Xi Wang, and Laurence Aitchison. 2023 · 2023
Later among the works it cites.
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al. 2023 · 2023
Later among the works it cites.
Navigating the grey area: Expressions of overconfidence and uncertainty in language models
Kaitlyn Zhou, Dan Jurafsky, and Tatsunori Hashimoto. 2023 · 2023
Later among the works it cites.
Can LLMs express their uncertainty? an empirical evaluation of confidence elicitation in LLMs
Anonymous. 2024 · 2024
Closest in time.
Mars: Meaning-aware response scoring for uncertainty estimation in generative llms
Yavuz Faruk Bakman, Duygu Nur Yaldiz, Baturalp Buyukates, Chenyang Tao, Dimitrios Dimitriadis, and Salman Avestimehr. 2024 · 2024
Closest in time.
Reconfidencing llms from the grouping loss perspective
Lihu Chen, Alexandre Perez-Lebel, Fabian M Suchanek, and Gaël Varoquaux. 2024 · 2024
Closest in time.
Large legal fictions: Profiling legal hallucinations in large language models
Matthew Dahl, Varun Magesh, Mirac Suzgun, and Daniel E Ho. 2024 · 2024
Closest in time.
Relying on the unreliable: The impact of language models’ reluctance to express uncertainty
Kaitlyn Zhou, Jena D Hwang, Xiang Ren, and Maarten Sap. 2024 · 2024
Closest in time.