Fetching the paper…
Reading the bibliography…
Black-box large language models (LLMs) are increasingly deployed in various environments, making it essential for these models to effectively convey their confidence and uncertainty, especially in high-stakes settings.
Pubmedqa: A dataset for biomedical research question answering
Qiao Jin, Bhuwan Dhingra, Zhengping Liu, William W. Cohen, and Xinghua Lu. 2019 · 1909
Earlier work this paper cites.
Strictly proper scoring rules, prediction, and estimation
Tilmann Gneiting and Adrian E Raftery. 2007 · 2007
Earlier work this paper cites.
Diagnostic difficulty and error in primary care–a systematic review
O Kostopoulou, BC Delaney, and CW Munro. 2008 · 2008
Earlier work this paper cites.
Di Jin, Eileen Pan, Nassim Oufattole, Wei-Hung Weng, Hanyi Fang, and Peter Szolovits. 2020 · 2009
Earlier work this paper cites.
Obtaining well calibrated probabilities using bayesian binning
Mahdi Pakdaman Naeini, Gregory F Cooper, and Milos Hauskrecht. 2015 · 2015
Earlier work this paper cites.
A survey of outpatient internal medicine clinician perceptions of diagnostic error
JC Matulis, SN Kok, EC Dankbar, and AJ Majka. 2020 · 2020
Earlier work this paper cites.
Atypical Presentations of Illness
Michael Goldrich and Amit Shah. 2021 · 2021
Earlier work this paper cites.
When you hear hoof beats, look for the zebras: Atypical presentation of illness in the older adult
Cassandra Vonnes and Rosalie El-Rady. 2021 · 2021
Earlier work this paper cites.
Teaching models to express their uncertainty in words
Stephanie Lin, Jacob Hilton, and Owain Evans. 2022 · 2022
Earlier work this paper cites.
Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering
Ankit Pal, Logesh Kumar Umapathi, and Malaikannan Sankarasubbu. 2022 · 2022
Earlier work this paper cites.
The benefits of using atypical presentations and rare diseases in problem-based learning in undergraduate medical education
S Bai, L Zhang, Z Ye, D Yang, T Wang, and Y Zhang. 2023 · 2023
Cited alongside, same era.
Clinical ai tools must convey predictive uncertainty for each individual patient
C. R. S. Banerji, T. Chakraborti, C. Harbron, et al. 2023 · 2023
Cited alongside, same era.
Quantifying uncertainty in answers from any language model and enhancing their trustworthiness
Jiuhai Chen and Jonas Mueller. 2023 · 2023
Cited alongside, same era.
Investigating uncertainty calibration of aligned language models under the multiple-choice setting
Guande He, Peng Cui, Jianfei Chen, Wenbo Hu, and Jun Zhu. 2023 · 2023
Cited alongside, same era.
Claude 3
Anthropic. 2023 · 2024
Closest in time.
Gemini 1.0 pro
Google DeepMind. 2023 · 2024
Closest in time.
Definitions and measurements for atypical presentations at risk for diagnostic errors in internal medicine: Protocol for a scoping review
Y Harada, R Kawamura, M Yokose, T Shimizu, and H Singh. 2024 · 2024
Closest in time.
Sunnie S. Y. Kim, Q. Vera Liao, Mihaela Vorvoreanu, Stephanie Ballard, and Jennifer Wortman Vaughan. 2024 · 2024
Closest in time.
Gpt-3.5-turbo
OpenAI. 2023 · 2024
Closest in time.
OpenAI. 2024 · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2023 · 2023
Cited alongside, same era.
Lorenz Kuhn, Yarin Gal, and Sebastian Farquhar. 2023 · 2023
Cited alongside, same era.
Llamas know what gpts don’t show: Surrogate models for confidence estimation
Vaishnavi Shrivastava, Percy Liang, and Ananya Kumar. 2023 · 2023
Cited alongside, same era.
Quantifying uncertainty in natural language explanations of large language models
Sree Harsha Tanneru, Chirag Agarwal, and Himabindu Lakkaraju. 2023 · 2023
Cited alongside, same era.
Katherine Tian, Eric Mitchell, Allan Zhou, Archit Sharma, Rafael Rafailov, Huaxiu Yao, Chelsea Finn, and Christopher D. Manning. 2023 · 2023
Cited alongside, same era.
Beyond confidence: Reliable models should also consider atypicality
Mert Yuksekgonul, Linjun Zhang, James Zou, and Carlos Guestrin. 2023 · 2023
Cited alongside, same era.
Closest in time.
Mauricio Rivera, Jean-François Godbout, Reihaneh Rabbany, and Kellin Pelrine. 2024 · 2024
Closest in time.
Can llms express their uncertainty? an empirical evaluation of confidence elicitation in llms
Miao Xiong, Zhiyuan Hu, Xinyang Lu, Yifei Li, Jie Fu, Junxian He, and Bryan Hooi. 2024 · 2024
Closest in time.
Benchmarking llms via uncertainty quantification
Fanghua Ye, Mingming Yang, Jianhui Pang, Longyue Wang, Derek F. Wong, Emine Yilmaz, Shuming Shi, and Zhaopeng Tu. 2024 · 2024
Closest in time.