Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) exhibit impressive performance across diverse domains but often suffer from overconfidence, limiting their reliability in critical applications.
Obtaining calibrated probability estimates from decision trees and naive bayesian classifiers
Bianca Zadrozny and Charles Elkan · 2001
Earlier work this paper cites.
Area under the precision-recall curve: point estimates and confidence intervals
Kendrick Boyd, Kevin H Eng, and C David Page · 2013
Earlier work this paper cites.
Bigbench: Towards an industry standard benchmark for big data analytics
Ahmad Ghazal, Tilmann Rabl, Minqing Hu, Francois Raab, Meikel Poess, Alain Crolotte, and Hans-Arno Jacobsen · 2013
Earlier work this paper cites.
Obtaining well calibrated probabilities using bayesian binning
Mahdi Pakdaman Naeini, Gregory Cooper, and Milos Hauskrecht · 2015
Earlier work this paper cites.
Dropout as a bayesian approximation: Representing model uncertainty in deep learning
Yarin Gal and Zoubin Ghahramani · 2016
Earlier work this paper cites.
On calibration of modern neural networks
Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q. Weinberger · 2017
Earlier work this paper cites.
Simple and scalable predictive uncertainty estimation using deep ensembles
Balaji Lakshminarayanan, Alexander Pritzel, and Charles Blundell · 2017
Earlier work this paper cites.
Ai-driven sales automation: Using chatbots to boost sales
Christian Hildebrand and Anouk Bergner · 2019
Earlier work this paper cites.
Learning to count objects with few exemplar annotations
Jianfeng Wang, Rong Xiao, Yandong Guo, and Lei Zhang · 2019
Earlier work this paper cites.
Mix-n-match: Ensemble and compositional methods for uncertainty calibration in deep learning
Jize Zhang, Bhavya Kailkhura, and T Yong-Jin Han · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman · 2021
Earlier work this paper cites.
Did aristotle use a laptop? A question answering benchmark with implicit reasoning strategies
Mor Geva, Daniel Khashabi, Elad Segal, Tushar Khot, Dan Roth, and Jonathan Berant · 2021
Earlier work this paper cites.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt · 2021
Earlier work this paper cites.
How can we know when language models know? on the calibration of language models for question answering
Zhengbao Jiang, Jun Araki, Haibo Ding, and Graham Neubig · 2021
Earlier work this paper cites.
Sports understanding in bigbench, 2021
Ethan Kim · 2021
Cited alongside, same era.
Revisiting the calibration of modern neural networks
Matthias Minderer, Josip Djolonga, Rob Romijnders, Frances Hubis, Xiaohua Zhai, Neil Houlsby, Dustin Tran, and Mario Lucic · 2021
Cited alongside, same era.
Data understanding in bigbench, 2021
Xinyi Wu and Zijian Wang · 2021
Cited alongside, same era.
Large-scale robust deep auc maximization: A new surrogate loss and empirical studies on medical image classification
Zhuoning Yuan, Yan Yan, Milan Sonka, and Tianbao Yang · 2021
Cited alongside, same era.
Language models (mostly) know what they know
Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield Dodds, Nova DasSarma, Eli Tran-Johnson, et al · 2022
Cited alongside, same era.
OpenAI · 2023
Later among the works it cites.
Large language models in medicine
Arun James Thirunavukarasu, Darren Shu Jeng Ting, Kabilan Elangovan, Laura Gutierrez, Ting Fang Tan, and Daniel Shu Wei Ting · 2023
Later among the works it cites.
Just ask for calibration: Strategies for eliciting calibrated confidence scores from language models fine-tuned with human feedback
Katherine Tian, Eric Mitchell, Allan Zhou, Archit Sharma, Rafael Rafailov, Huaxiu Yao, Chelsea Finn, and Christopher D Manning · 2023
Later among the works it cites.
Proximity-informed calibration for deep neural networks
Miao Xiong, Ailin Deng, Pang Wei W Koh, Jiaying Wu, Shen Li, Jianqing Xu, and Bryan Hooi · 2023
Later among the works it cites.
Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs
Miao Xiong, Zhiyuan Hu, Xinyang Lu, Yifei Li, Jie Fu, Junxian He, and Bryan Hooi · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Stephanie Lin, Jacob Hilton, and Owain Evans · 2022
Cited alongside, same era.
Reducing conversational agents’ overconfidence through linguistic calibration
Sabrina J Mielke, Arthur Szlam, Emily Dinan, and Y-Lan Boureau · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al · 2022
Cited alongside, same era.
Birds of a feather trust together: Knowing when to trust a classifier via adaptive neighborhood aggregation
Miao Xiong, Shen Li, Wenjie Feng, Ailin Deng, Jihai Zhang, and Bryan Hooi · 2022
Cited alongside, same era.
A close look into the calibration of pre-trained language models
Yangyi Chen, Lifan Yuan, Ganqu Cui, Zhiyuan Liu, and Heng Ji · 2023
Cited alongside, same era.
Great models think alike: Improving model reliability via inter-model latent agreement
Ailin Deng, Miao Xiong, and Bryan Hooi · 2023
Cited alongside, same era.
A survey of uncertainty in deep neural networks
Jakob Gawlikowski, Cedrique Rovile Njieutcheu Tassi, Mohsin Ali, Jongseok Lee, Matthias Humt, Jianxiang Feng, Anna M. Kruspe, Rudolph Triebel, Peter Jung, Ribana Roscher, Muhammad Shahzad, Wen Yang, Richard Bamler, and Xiaoxiang Zhu · 2023
Cited alongside, same era.
Navigating the grey area: How expressions of uncertainty and overconfidence affect language models
Kaitlyn Zhou, Dan Jurafsky, and Tatsunori B Hashimoto · 2023
Later among the works it cites.
Large language model influence on diagnostic reasoning: a randomized clinical trial
Ethan Goh, Robert Gallo, Jason Hom, Eric Strong, Yingjie Weng, Hannah Kerman, Joséphine A Cool, Zahir Kanjee, Andrew S Parsons, Neera Ahuja, et al · 2024
Later among the works it cites.
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, et al · 2024
Later among the works it cites.
Leveraging the potential of large language models in education through playful and game-based learning
Stefan E Huber, Kristian Kiili, Steve Nebel, Richard M Ryan, Michael Sailer, and Manuel Ninaus · 2024
Later among the works it cites.
A survey on large language models for code generation
Juyong Jiang, Fan Wang, Jiasi Shen, Sungju Kim, and Sunghun Kim · 2024
Later among the works it cites.
Large language models in law: A survey
Jinqi Lai, Wensheng Gan, Jiayang Wu, Zhenlian Qi, and Philip S Yu · 2024
Later among the works it cites.
The application of large language models in medicine: A scoping review
Xiangbin Meng, Xiangyu Yan, Kuo Zhang, Da Liu, Xiaojuan Cui, Yaodong Yang, Muhan Zhang, Chunxia Cao, Jingjia Wang, Xuliang Wang, et al · 2024
Later among the works it cites.
Using ai-driven chatbots to foster chinese efl students’ academic engagement: An intervention study
Yongliang Wang and Lina Xue · 2024
Later among the works it cites.
Large language models for automated q&a involving legal documents: a survey on algorithms, frameworks and applications
Xiaoxian Yang, Zhifeng Wang, Qi Wang, Ke Wei, Kaiqi Zhang, and Jiangang Shi · 2024
Later among the works it cites.