Fetching the paper…
Reading the bibliography…
As artificial intelligence (AI) systems, particularly large language models (LLMs), become increasingly integrated into decision-making processes, the ability to trust their outputs is crucial.
The effects of involvement on responses to argument quantity and quality: Central and peripheral routes to persuasion
Richard E Petty and John T Cacioppo · 1984
Earlier work this paper cites.
Words or numbers? the evaluation of probability expressions in general practice
Bernie J O’Brien · 1989
Earlier work this paper cites.
Misremembrance of options past: Source monitoring and choice
Mara Mather, Eldar Shafir, and Marcia K Johnson · 2000
Earlier work this paper cites.
Toward a universal translator of verbal probabilities
Tzur M. Karelitz, Mandeep K. Dhami, David V. Budescu, and Thomas S. Wallsten · 2002
Earlier work this paper cites.
Consequences of erudite vernacular utilized irrespective of necessity: Problems with using long words needlessly
Daniel M Oppenheimer · 2006
Earlier work this paper cites.
Exploring intelligence analysts’ selection and interpretation of probability terms
Thomas S. Wallsten, Yaron Shlomi, and Hisuchi Ting · 2008
Earlier work this paper cites.
Default bayes factors for anova designs
Jeffrey N Rouder, Richard D Morey, Paul L Speckman, and Jordan M Province · 2012
Earlier work this paper cites.
The interpretation of IPCC probabilistic statements around the world
David V Budescu, Han-Hui Por, and Michael Broomell, Stephen B andSmithson · 2014
Earlier work this paper cites.
How to measure metacognition
Stephen M Fleming and Hakwan C Lau · 2014
Earlier work this paper cites.
Improving the communication of uncertainty in climate science and intelligence analysis
Emily H Ho, David V Budescu, Mandeep K Dhami, and David R Mandel · 2015
Earlier work this paper cites.
Obtaining well calibrated probabilities using bayesian binning
Mahdi Pakdaman Naeini, Gregory Cooper, and Milos Hauskrecht · 2015
Earlier work this paper cites.
On calibration of modern neural networks
Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q. Weinberger · 2017
Earlier work this paper cites.
Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension
Mandar Joshi, Eunsol Choi, Daniel S Weld, and Luke Zettlemoyer · 2017
Earlier work this paper cites.
What can ai do for me? evaluating machine learning interpretations in cooperative play
Shi Feng and Jordan Boyd-Graber · 2019
Earlier work this paper cites.
Verified uncertainty calibration
Ananya Kumar, Percy S Liang, and Tengyu Ma · 2019
Earlier work this paper cites.
On mixup training: Improved calibration and predictive uncertainty for deep neural networks
Sunil Thulasidasan, Gopinath Chennupati, Jeff A Bilmes, Tanmoy Bhattacharya, and Sarah Michalak · 2019
Earlier work this paper cites.
No explainability without accountability: An empirical study of explanations and feedback in interactive ml
Alison Smith-Renner, Ron Fan, Melissa Birchfield, Tongshuang Wu, Jordan Boyd-Graber, Daniel S. Weld, and Leah Findlater · 2020
Earlier work this paper cites.
Does the whole exceed its parts? the effect of ai explanations on complementary team performance
Gagan Bansal, Tongshuang Wu, Joyce Zhou, Raymond Fok, Besmira Nushi, Ece Kamar, Marco Tulio Ribeiro, and Daniel Weld · 2021
Earlier work this paper cites.
To trust or to think: Cognitive forcing functions can reduce overreliance on ai in ai-assisted decision-making
Zana Buçinca, Maja Barbara Malaya, and Krzysztof Z Gajos · 2021
Cited alongside, same era.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt · 2021
Cited alongside, same era.
How can we know when language models know? on the calibration of language models for question answering
Zhengbao Jiang, Jun Araki, Haibo Ding, and Graham Neubig · 2021
Cited alongside, same era.
Scaling language models: Methods, analysis & insights from training gopher
Jack W Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, et al · 2021
Cited alongside, same era.
Better uncertainty calibration via proper scores for classification and beyond
Sebastian Gruber and Florian Buettner · 2022
The promise and peril of generative ai
A Jo · 2023
Later among the works it cites.
Gpt-3.5, Nov 2022a
OpenAI · 2023
Later among the works it cites.
Introducing chatgpt, Nov 2022b
OpenAI · 2023
Later among the works it cites.
Towards human-centered explainable ai: A survey of user studies for model explanations
Yao Rong, Tobias Leemann, Thai-Trang Nguyen, Lisa Fiedler, Peizhu Qian, Vaibhav Unhelkar, Tina Seidel, Gjergji Kasneci, and Enkelejda Kasneci · 2023
Later among the works it cites.
Verbosity bias in preference labeling by large language models
Keita Saito, Akifumi Wachi, Koki Wataoka, and Youhei Akimoto · 2023
Later among the works it cites.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R Brown, Adam Santoro, Aditya Gupta, Adrià Garriga-Alonso, et al · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Training compute-optimal large language models
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al · 2022
Cited alongside, same era.
Language models (mostly) know what they know
Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield-Dodds, Nova DasSarma, Eli Tran-Johnson, et al · 2022
Cited alongside, same era.
Teaching models to express their uncertainty in words
Stephanie Lin, Jacob Hilton, and Owain Evans · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Cited alongside, same era.
Effects of explanations in ai-assisted decision making: Principles and comparisons
Xinru Wang and Ming Yin · 2022
Cited alongside, same era.
Uncertainty quantification with pre-trained language models: A large-scale empirical analysis
Yuxin Xiao, Paul Pu Liang, Umang Bhatt, Willie Neiswanger, Ruslan Salakhutdinov, and Louis-Philippe Morency · 2022
Cited alongside, same era.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Cited alongside, same era.
Later among the works it cites.
Three challenges for ai-assisted decision-making
Mark Steyvers and Aakriti Kumar · 2023
Later among the works it cites.
Just ask for calibration: Strategies for eliciting calibrated confidence scores from language models fine-tuned with human feedback
Katherine Tian, Eric Mitchell, Allan Zhou, Archit Sharma, Rafael Rafailov, Huaxiu Yao, Chelsea Finn, and Christopher Manning · 2023
Later among the works it cites.
Chatgpt: Challenges, opportunities, and implications for teacher education
Jeromie Whalen, Chrystalla Mouza, et al · 2023
Later among the works it cites.
Do large language models know what they don’t know?
Zhangyue Yin, Qiushi Sun, Qipeng Guo, Jiawen Wu, Xipeng Qiu, and Xuanjing Huang · 2023
Later among the works it cites.
From ncoder to chatgpt: From automated coding to refining human coding
Andres Felipe Zambrano, Xiner Liu, Amanda Barany, Ryan S Baker, Juhan Kim, and Nidhi Nasiar · 2023
Later among the works it cites.
Navigating the grey area: How expressions of uncertainty and overconfidence affect language models
Kaitlyn Zhou, Dan Jurafsky, and Tatsunori Hashimoto · 2023
Later among the works it cites.
How experts and novices judge other people’s knowledgeability from language use
Alexander H Bower, Nicole Han, Ansh Soni, Miguel P Eckstein, and Mark Steyvers · 2024
Closest in time.
Detecting hallucinations in large language models using semantic entropy
Sebastian Farquhar, Jannik Kossen, Lorenz Kuhn, and Yarin Gal · 2024
Closest in time.
Quantifying uncertainty in natural language explanations of large language models
Sree Harsha Tanneru, Chirag Agarwal, and Himabindu Lakkaraju · 2024
Closest in time.
Can LLMs express their uncertainty? an empirical evaluation of confidence elicitation in LLMs
Miao Xiong, Zhiyuan Hu, Xinyang Lu, YIFEI LI, Jie Fu, Junxian He, and Bryan Hooi · 2024
Closest in time.
Relying on the unreliable: The impact of language models’ reluctance to express uncertainty
Kaitlyn Zhou, Jena Hwang, Xiang Ren, and Maarten Sap · 2024
Closest in time.