Fetching the paper…
Reading the bibliography…
Widely deployed large language models (LLMs) can produce convincing yet incorrect outputs, potentially misleading users who may rely on them as if they were correct.
Coefficient alpha and the internal structure of tests
Lee J. Cronbach. 1951 · 1951
Earlier work this paper cites.
Verbal Vs. Numerical Processing of Subjective Probabilities
Alf C. Zimmer. 1983 · 1983
Earlier work this paper cites.
Verbal uncertainty expressions: A critical review of two decades of research
Dominic A. Clark. 1990 · 1990
Earlier work this paper cites.
Preferences and reasons for communicating probabilistic information in verbal or numerical terms
Thomas S. Wallsten, David V. Budescu, Rami Zwick, and Steven M. Kemp. 1993 · 1993
Earlier work this paper cites.
An Integrative Model of Organizational Trust
Roger C. Mayer, James H. Davis, and F. David Schoorman. 1995 · 1995
Earlier work this paper cites.
Methodology matters: Doing research in the behavioral and social sciences
Joseph E McGrath. 1995 · 1995
Earlier work this paper cites.
Measuring psychological uncertainty: Verbal versus numeric methods
Paul D Windschitl and Gary L Wells. 1996 · 1996
Earlier work this paper cites.
Transforming qualitative information: Thematic analysis and code development
Richard E Boyatzis. 1998 · 1998
Earlier work this paper cites.
Machines and mindlessness: Social responses to computers
Clifford Nass and Youngme Moon. 2000 · 2000
Earlier work this paper cites.
Understanding Self-Report Bias in Organizational Behavior Research
Stewart I. Donaldson and Elisa J. Grant-Vallone. 2002 · 2002
Earlier work this paper cites.
The impact of initial consumer trust on intentions to transact with a web site: a trust building model
D. Harrison McKnight, Vivek Choudhury, and Charles Kacmar. 2002 · 2002
Earlier work this paper cites.
Probability Neglect: Emotions, Worst Cases, and Law
Cass R. Sunstein. 2002 · 2002
Earlier work this paper cites.
Researchers misunderstand confidence intervals and standard error bars
Sarah Belia, Fiona Fidler, Jennifer Williams, and Geoff Cumming. 2005 · 2005
Earlier work this paper cites.
Advice taking and decision-making: An integrative literature review, and implications for the organizational sciences
Silvia Bonaccio and Reeshad S Dalal. 2006 · 2006
Earlier work this paper cites.
Using thematic analysis in psychology
Virginia Braun and Victoria Clarke. 2006 · 2006
Earlier work this paper cites.
G*Power 3: A flexible statistical power analysis program for the social, behavioral, and biomedical sciences
Franz Faul, Edgar Erdfelder, Albert-Georg Lang, and Axel Buchner. 2007 · 2007
Earlier work this paper cites.
The Effects of Hedges in Persuasive Arguments: A Nuanced Analysis of Language
Amanda M. Durik, M. Anne Britt, Rebecca Reynolds, and Jennifer Storey. 2008 · 2008
Earlier work this paper cites.
Measurement Instruments for the Anthropomorphism, Animacy, Likeability, Perceived Intelligence, and Perceived Safety of Robots
Christoph Bartneck, Dana Kulić, Elizabeth Croft, and Susana Zoghbi. 2009 · 2009
Earlier work this paper cites.
Statistical power analyses using G*Power 3.1: Tests for correlation and regression analyses
Franz Faul, Edgar Erdfelder, Axel Buchner, and Albert-Georg Lang. 2009 · 2009
Earlier work this paper cites.
Verbal Expressions of Confidence and Doubt
Caroline J. Wesson and Briony D. Pulford. 2009 · 2009
Earlier work this paper cites.
Running experiments on Amazon Mechanical Turk
Gabriele Paolacci, Jesse Chandler, and Panagiotis G. Ipeirotis. 2010 · 2010
Earlier work this paper cites.
Who Are the Crowdworkers? Shifting Demographics in Mechanical Turk. In CHI ’10 Extended Abstracts on Human Factors in Computing Systems (Atlanta, Georgia, USA) (CHI EA ’10) . Association for Computing Machinery, New York, NY, USA, 2863–2872
Joel Ross, Lilly Irani, M. Six Silberman, Andrew Zaldivar, and Bill Tomlinson. 2010 · 2010
Earlier work this paper cites.
Amazon’s Mechanical Turk: A New Source of Inexpensive, Yet High-Quality, Data?
Michael Buhrmester, Tracy Kwang, and Samuel Gosling. 2011 · 2011
Earlier work this paper cites.
Evaluating Online Labor Markets for Experimental Research: Amazon.com’s Mechanical Turk
Adam Berinsky, Gregory Huber, Gabriel Lenz, and R. Alvarez. 2012 · 2012
Earlier work this paper cites.
Conducting behavioral research on Amazon’s Mechanical Turk
Winter Mason and Siddharth Suri. 2012 · 2012
Earlier work this paper cites.
Separate but equal? A comparison of participants and data gathered via Amazon’s MTurk, social media, and face-to-face behavioral testing
Krista Casler, Lydia Bickel, and Elizabeth Hackett. 2013 · 2013
Earlier work this paper cites.
Thinking, Fast and Slow
Daniel Kahneman. 2013 · 2013
Earlier work this paper cites.
How a robot should give advice. In 2013 8th ACM/IEEE International Conference on Human-Robot Interaction (HRI) . 275–282
Cristen Torrey, Susan R. Fussell, and Sara Kiesler. 2013 · 2013
Earlier work this paper cites.
Inside the Turk: Understanding Mechanical Turk as a Participant Pool
Gabriele Paolacci and Jesse Chandler. 2014 · 2014
Earlier work this paper cites.
“Who are these people?” Evaluating the demographic characteristics and political preferences of MTurk survey respondents
Connor Huff and Dustin Tingley. 2015 · 2015
Earlier work this paper cites.
Research in the Crowdsourcing Age, a Case Study
Paul Hitlin. 2016 · 2016
Earlier work this paper cites.
Persuasive effects of point of view, protagonist competence, and similarity in a health narrative about type 2 diabetes
Meng Chen, Robert A Bell, and Laramie D Taylor. 2017 · 2017
Earlier work this paper cites.
Theory of Machine: When Do People Rely on Algorithms? (2017)
Jennifer M. Logg. 2017 · 2017
Earlier work this paper cites.
Resilient Chatbots: Repair Strategy Preferences for Conversational Breakdowns. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems (Glasgow, Scotland Uk) (CHI ’19) . Association for Computing Machinery, New York, NY, USA, 1–12
Zahra Ashktorab, Mohit Jain, Q. Vera Liao, and Justin D. Weisz. 2019 · 2019
Earlier work this paper cites.
A Question-Entailment Approach to Question Answering
Asma Ben Abacha and Dina Demner-Fushman. 2019 · 2019
Earlier work this paper cites.
"Hello AI": Uncovering the Onboarding Needs of Medical Practitioners for Human-AI Collaborative Decision-Making
Carrie J. Cai, Samantha Winter, David Steiner, Lauren Wilcox, and Michael Terry. 2019 · 2019
Earlier work this paper cites.
Generalizing from Survey Experiments Conducted on Mechanical Turk: A Replication Approach
Alexander Coppock. 2019 · 2019
Earlier work this paper cites.
On Human Predictions with Explanations and Predictions of Machine Learning Models: A Case Study on Deception Detection. In Proceedings of the Conference on Fairness, Accountability, and Transparency (Atlanta, GA, USA) (FAT* ’19) . Association for Computing Machinery, New York, NY, USA, 29–38
Vivian Lai and Chenhao Tan. 2019 · 2019
Earlier work this paper cites.
Algorithm appreciation: People prefer algorithmic to human judgment
Jennifer M. Logg, Julia A. Minson, and Don A. Moore. 2019 · 2019
Earlier work this paper cites.
Model Cards for Model Reporting. In Proceedings of the Conference on Fairness, Accountability, and Transparency (Atlanta, GA, USA) (FAT* ’19) . Association for Computing Machinery, New York, NY, USA, 220–229
Margaret Mitchell, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasserman, Ben Hutchinson, Elena Spitzer, Inioluwa Deborah Raji, and Timnit Gebru. 2019 · 2019
Earlier work this paper cites.
Understanding the Effect of Accuracy on Trust in Machine Learning Models. In Proceedings of the 2019 ACM CHI Conference on Human Factors in Computing Systems
Ming Yin, Jennifer Wortman Vaughan, and Hanna Wallach. 2019 · 2019
Cited alongside, same era.
Language Models are Few-Shot Learners. In Advances in Neural Information Processing Systems , H. Larochelle, M. Ranzato, R. Hadsell, M.F. Balcan, and H. Lin (Eds.), Vol. 33. Curran Associates, Inc., 1877–1901
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 2020
Cited alongside, same era.
An MTurk Crisis? Shifts in Data Quality and the Impact on Study Results
Michael Chmielewski and Sarah C. Kucker. 2020 · 2020
Cited alongside, same era.
Conversational Repair in Chatbots for Customer Service: The Effect of Expressing Uncertainty and Suggesting Alternatives. In Chatbot Research and Design , Asbjørn Følstad, Theo Araujo, Symeon Papadopoulos, Effie Lai-Chong Law, Ole-Christoffer Granmo, Ewa Luger, and Petter Bae Brandtzaeg (Eds.). Springer International Publishing, Cham, 201–214
Mirages. On Anthropomorphism in Dialogue Systems. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , Houda Bouamor, Juan Pino, and Kalika Bali (Eds.). Association for Computational Linguistics, Singapore, 4776–4790
Gavin Abercrombie, Amanda Cercas Curry, Tanvi Dinkar, Verena Rieser, and Zeerak Talat. 2023 · 2023
Later among the works it cites.
Announcing our series A funding round and mobile app launch
Perplexity AI. 2023 · 2023
Later among the works it cites.
Knowledge of Knowledge: Exploring Known-Unknowns Uncertainty with Large Language Models
Alfonso Amayuelas, Liangming Pan, Wenhu Chen, and William Wang. 2023 · 2023
Later among the works it cites.
Uncertainty in Natural Language Generation: From Theory to Applications
Joris Baan, Nico Daheim, Evgenia Ilia, Dennis Ulmer, Haau-Sing Li, Raquel Fernández, Barbara Plank, Rico Sennrich, Chrysoula Zerva, and Wilker Aziz. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Asbjørn Følstad and Cameron Taylor. 2020 · 2020
Cited alongside, same era.
How Visualizing Inferential Uncertainty Can Mislead Readers About Treatment Effects in Scientific Results. In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems
Jake M. Hofman, Daniel G. Goldstein, and Jessica Hullman. 2020 · 2020
Cited alongside, same era.
The shape of and solutions to the MTurk quality crisis
Ryan Kennedy, Scott Clifford, Tyler Burleigh, Philip D. Waggoner, Ryan Jewell, and Nicholas J. G. Winter. 2020 · 2020
Cited alongside, same era.
The intuitive use of contextual information in decisions made with verbal and numerical quantifiers
Dawn Liu, Marie Juanchich, Miroslav Sirota, and Sheina Orbell. 2020 · 2020
Cited alongside, same era.
2020 Census of Population and Housing
United States Census Bureau. 2020 · 2020
Cited alongside, same era.
Effect of confidence and explanation on accuracy and trust calibration in AI-assisted decision making. In Proceedings of the 2020 conference on fairness, accountability, and transparency . 295–305
Yunfeng Zhang, Q Vera Liao, and Rachel KE Bellamy. 2020 · 2020
Cited alongside, same era.
Does the Whole Exceed Its Parts? The Effect of AI Explanations on Complementary Team Performance. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems (Yokohama, Japan) (CHI ’21) . Association for Computing Machinery, New York, NY, USA, Article 81, 16 pages
Gagan Bansal, Tongshuang Wu, Joyce Zhou, Raymond Fok, Besmira Nushi, Ece Kamar, Marco Tulio Ribeiro, and Daniel Weld. 2021 · 2021
Cited alongside, same era.
On the Dangers of Stochastic Parrots: Can Language Models Be Too Big?. In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency (Virtual Event, Canada) (FAccT ’21) . Association for Computing Machinery, New York, NY, USA, 610–623
Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021 · 2021
Cited alongside, same era.
Uncertainty as a Form of Transparency: Measuring, Communicating, and Using Uncertainty. In Proceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society (Virtual Event, USA) (AIES ’21) . Association for Computing Machinery, New York, NY, USA, 401–413
Umang Bhatt, Javier Antorán, Yunfeng Zhang, Q. Vera Liao, Prasanna Sattigeri, Riccardo Fogliato, Gabrielle Melançon, Ranganath Krishnan, Jason Stanley, Omesh Tickoo, Lama Nachman, Rumi Chunara, Madhulika Srikumar, Adrian Weller, and Alice Xiang. 2021 · 2021
Cited alongside, same era.
Understanding the Role of Human Intuition on Reliance in Human-AI Decision-Making with Explanations
Valerie Chen, Q. Vera Liao, Jennifer Wortman Vaughan, and Gagan Bansal. 2023a · 2023
Later among the works it cites.
A Close Look into the Calibration of Pre-trained Language Models. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . Association for Computational Linguistics, Toronto, Canada, 1343–1367
Yangyi Chen, Lifan Yuan, Ganqu Cui, Zhiyuan Liu, and Heng Ji. 2023b · 2023
Later among the works it cites.
Selectively Answering Ambiguous Questions. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , Houda Bouamor, Juan Pino, and Kalika Bali (Eds.). Association for Computational Linguistics, Singapore, 530–543
Jeremy Cole, Michael Zhang, Daniel Gillick, Julian Eisenschlos, Bhuwan Dhingra, and Jacob Eisenstein. 2023 · 2023
Later among the works it cites.
Social Dynamics of AI Support in Creative Writing. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (Hamburg, Germany) (CHI ’23) . Association for Computing Machinery, New York, NY, USA, Article 245, 15 pages
Katy Ilonka Gero, Tao Long, and Lydia B Chilton. 2023 · 2023
Later among the works it cites.
Survey of Hallucination in Natural Language Generation
Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung. 2023 · 2023
Later among the works it cites.
Samia Kabir, David N. Udo-Imeh, Bonan Kou, and Tianyi Zhang. 2023 · 2023
Later among the works it cites.
Humans, AI, and Context: Understanding End-Users’ Trust in a Real-World Computer Vision Application. In Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency (Chicago, IL, USA) (FAccT ’23) . Association for Computing Machinery, New York, NY, USA, 77–88
Sunnie S. Y. Kim, Elizabeth Anne Watkins, Olga Russakovsky, Ruth Fong, and Andrés Monroy-Hernández. 2023 · 2023
Later among the works it cites.
Tracking public attitudes toward ChatGPT on Twitter using sentiment analysis and topic modeling
Ratanond Koonchanok, Yanling Pan, and Hyeju Jang. 2023 · 2023
Later among the works it cites.
Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language Generation. In The Eleventh International Conference on Learning Representations
Lorenz Kuhn, Yarin Gal, and Sebastian Farquhar. 2023 · 2023
Later among the works it cites.
Towards a Science of Human-AI Decision Making: An Overview of Design Space in Empirical Human-Subject Studies. In Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency (FAccT ’23) . Association for Computing Machinery, New York, NY, USA, 1369–1385
Vivian Lai, Chacha Chen, Alison Smith-Renner, Q. Vera Liao, and Chenhao Tan. 2023 · 2023
Later among the works it cites.
Generating with Confidence: Uncertainty Quantification for Black-box Large Language Models
Zhen Lin, Shubhendu Trivedi, and Jimeng Sun. 2023 · 2023
Later among the works it cites.
Evaluating Verifiability in Generative Search Engines. In Findings of the Association for Computational Linguistics: EMNLP 2023 , Houda Bouamor, Juan Pino, and Kalika Bali (Eds.). Association for Computational Linguistics, Singapore, 7001–7025
Nelson Liu, Tianyi Zhang, and Percy Liang. 2023 · 2023
Later among the works it cites.
Who Broke Amazon Mechanical Turk? An Analysis of Crowdsourcing Data Quality over Time. In Proceedings of the 15th ACM Web Science Conference 2023 (Austin, TX, USA) (WebSci ’23) . Association for Computing Machinery, New York, NY, USA, 335–345
Catherine C. Marshall, Partha S.R. Goguladinne, Mudit Maheshwari, Apoorva Sathe, and Frank M. Shipman. 2023 · 2023
Later among the works it cites.
The New Bing and Edge – Progress from Our First Month
Yusuf Mehdi. 2023 · 2023
Later among the works it cites.
OpenAI. 2023 · 2023
Later among the works it cites.
Understanding Uncertainty: How Lay Decision-Makers Perceive and Interpret Uncertainty in Human-AI Decision Making. In Proceedings of the 28th International Conference on Intelligent User Interfaces (Sydney, NSW, Australia) (IUI ’23) . Association for Computing Machinery, New York, NY, USA, 379–396
Snehal Prabhudesai, Leyao Yang, Sumit Asthana, Xun Huan, Q. Vera Liao, and Nikola Banovic. 2023 · 2023
Later among the works it cites.
From Copilot to Pilot: Towards AI Supported Software Development
Rohith Pudari and Neil A. Ernst. 2023 · 2023
Later among the works it cites.
“I Think You Might Like This”: Exploring Effects of Confidence Signal Patterns on Trust in and Reliance on Conversational Recommender Systems. In Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency (Chicago, IL, USA) (FAccT ’23) . Association for Computing Machinery, New York, NY, USA, 792–804
Marissa Radensky, Julie Anne Séguin, Jang Soo Lim, Kristen Olson, and Robert Geiger. 2023 · 2023
Later among the works it cites.
Appropriate Reliance on AI Advice: Conceptualization and the Effect of Explanations. In Proceedings of the 28th International Conference on Intelligent User Interfaces (IUI ’23) . Association for Computing Machinery, New York, NY, USA, 410–422
Max Schemmer, Niklas Kuehl, Carina Benz, Andrea Bartos, and Gerhard Satzger. 2023 · 2023
Later among the works it cites.
Sociotechnical Harms of Algorithmic Systems: Scoping a Taxonomy for Harm Reduction. In Proceedings of the 2023 AAAI/ACM Conference on AI, Ethics, and Society (AIES ’23) . Association for Computing Machinery, New York, NY, USA, 723–741
Renee Shelby, Shalaleh Rismani, Kathryn Henne, AJung Moon, Negar Rostamzadeh, Paul Nicholas, N’Mah Yilla-Akbari, Jess Gallegos, Andrew Smart, Emilio Garcia, and Gurleen Virk. 2023 · 2023
Later among the works it cites.
Comparing Traditional and LLM-based Search for Consumer Choice: A Randomized Experiment
Sofia Eleni Spatharioti, David M. Rothschild, Daniel G. Goldstein, and Jake M. Hofman. 2023 · 2023
Later among the works it cites.
Is Perplexity AI showing us the future of search?
Mark Sullivan. 2023 · 2023
Later among the works it cites.
Artificial Intelligence Risk Management Framework (AI RMF 1.0)
Elham Tabassi. 2023 · 2023
Later among the works it cites.
Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , Houda Bouamor, Juan Pino, and Kalika Bali (Eds.). Association for Computational Linguistics, Singapore, 5433–5442
Katherine Tian, Eric Mitchell, Allan Zhou, Archit Sharma, Rafael Rafailov, Huaxiu Yao, Chelsea Finn, and Christopher Manning. 2023 · 2023
Later among the works it cites.
Educational Attainment in the United States: 2022
United States Census Bureau. 2022 · 2023
Later among the works it cites.
Explanations Can Reduce Overreliance on AI Systems During Decision-Making
Helena Vasconcelos, Matthew Jörke, Madeleine Grunde-McLaughlin, Tobias Gerstenberg, Michael S. Bernstein, and Ranjay Krishna. 2023b · 2023
Later among the works it cites.
Veniamin Veselovsky, Manoel Horta Ribeiro, and Robert West. 2023 · 2023
Later among the works it cites.
The ChatGPT Lawyer Explains Himself
Benjamin Weiser and Nate Schweber. 2023 · 2023
Later among the works it cites.
What Do We Mean When We Talk about Trust in Social Media? A Systematic Review. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (CHI ’23) . Association for Computing Machinery, New York, NY, USA, Article 670, 22 pages
Yixuan Zhang, Joseph D Gaggiano, Nutchanon Yongsatianchot, Nurul M Suhaimi, Miso Kim, Yifan Sun, Jacqueline Griffin, and Andrea G Parker. 2023 · 2023
Later among the works it cites.
Navigating the Grey Area: How Expressions of Uncertainty and Overconfidence Affect Language Models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , Houda Bouamor, Juan Pino, and Kalika Bali (Eds.). Association for Computational Linguistics, Singapore, 5506–5524
Kaitlyn Zhou, Dan Jurafsky, and Tatsunori Hashimoto. 2023 · 2023
Later among the works it cites.
AI Transparency in the Age of LLMs: A Human-Centered Research Roadmap
Q. Vera Liao and Jennifer Wortman Vaughan. 2024 · 2024
Closest in time.
When to Show a Suggestion? Integrating Human Feedback in AI-Assisted Programming
Hussein Mozannar, Gagan Bansal, Adam Fourney, and Eric Horvitz. 2024 · 2024
Closest in time.
European Union Artificial Intelligence Act Corrigendum
European Parliament. 2024 · 2024
Closest in time.
Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs. In The Twelfth International Conference on Learning Representations
Miao Xiong, Zhiyuan Hu, Xinyang Lu, YIFEI LI, Jie Fu, Junxian He, and Bryan Hooi. 2024 · 2024
Closest in time.
Relying on the Unreliable: The Impact of Language Models’ Reluctance to Express Uncertainty
Kaitlyn Zhou, Jena D. Hwang, Xiang Ren, and Maarten Sap. 2024 · 2024
Closest in time.