Fetching the paper…
Reading the bibliography…
We investigate the impact of hallucinations and Cognitive Forcing Functions in human-AI collaborative content-grounded data generation, focusing on the use of Large Language Models (LLMs) to assist in generating high quality conversational data.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 1901
Earlier work this paper cites.
Computational interpretations of the Gricean maxims in the generation of referring expressions
Robert Dale and Ehud Reiter. 1995 · 1995
Earlier work this paper cites.
Automation bias and errors: are crews better than individuals?
Linda J Skitka, Kathleen L Mosier, Mark Burdick, and Bonnie Rosenblatt. 2000 · 2000
Earlier work this paper cites.
Reactance to recommendations: When unsolicited advice yields contrary responses
Gavan J Fitzsimons and Donald R Lehmann. 2004 · 2004
Earlier work this paper cites.
ROUGE: A Package for Automatic Evaluation of Summaries. In Text Summarization Branches Out . Association for Computational Linguistics, Barcelona, Spain, 74–81
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Transforming HR for strategic impact at AstraZeneca: How centralized HR support improved internal customer service
Simon Hayward. 2006 · 2006
Earlier work this paper cites.
The evolution of HR: Developing HR as an internal consulting organization
Richard M Vosburgh. 2007 · 2007
Earlier work this paper cites.
Cheap and Fast – But is it Good? Evaluating Non-Expert Annotations for Natural Language Tasks. In Proceedings of the 2008 Conference on Empirical Methods in Natural Language Processing , Mirella Lapata and Hwee Tou Ng (Eds.). Association for Computational Linguistics, Honolulu, Hawaii, 254–263
Rion Snow, Brendan O’Connor, Daniel Jurafsky, and Andrew Ng. 2008 · 2008
Earlier work this paper cites.
Thinking, Fast and Slow
Daniel Kahneman. 2011 · 2011
Earlier work this paper cites.
Automation bias: a systematic review of frequency, effect mediators, and mitigators
Kate Goddard, Abdul Roudsari, and Jeremy C Wyatt. 2012 · 2012
Earlier work this paper cites.
Nested by design: model fitting and interpretation in a mixed model era
Holger Schielzeth and Shinichi Nakagawa. 2013 · 2013
Earlier work this paper cites.
Linear and generalized linear mixed models
Benjamin M Bolker. 2015 · 2015
Earlier work this paper cites.
Dual-process cognitive interventions to enhance diagnostic reasoning: a systematic review
Kathryn Ann Lambe, Gary O’Reilly, Brendan D Kelly, and Sarah Curristan. 2016 · 2016
Earlier work this paper cites.
Go/No Go Criteria in Formative E-Rubrics. In Learning and Collaboration Technologies. Learning and Teaching , Panayiotis Zaphiris and Andri Ioannou (Eds.). Springer International Publishing, Cham, 254–264
Pedro Company, Jeffrey Otey, María-Jesús Agost, Manuel Contero, and Jorge D. Camba. 2018 · 2018
Earlier work this paper cites.
Expert, Crowdsourced, and Machine Assessment of Suicide Risk via Online Postings. In Proceedings of the Fifth Workshop on Computational Linguistics and Clinical Psychology: From Keyboard to Clinic , Kate Loveys, Kate Niederhoffer, Emily Prud’hommeaux, Rebecca Resnik, and Philip Resnik (Eds.). Association for Computational Linguistics, New Orleans, LA, 25–36
Han-Chin Shing, Suraj Nair, Ayah Zirikly, Meir Friedenberg, Hal Daumé III, and Philip Resnik. 2018 · 2018
Earlier work this paper cites.
The Fact Extraction and VERification (FEVER) Shared Task. In Proceedings of the First Workshop on Fact Extraction and VERification (FEVER) , James Thorne, Andreas Vlachos, Oana Cocarascu, Christos Christodoulopoulos, and Arpit Mittal (Eds.). Association for Computational Linguistics, Brussels, Belgium, 1–9
James Thorne, Andreas Vlachos, Oana Cocarascu, Christos Christodoulopoulos, and Arpit Mittal. 2018 · 2018
Earlier work this paper cites.
Explaining Decision-Making Algorithms through UI: Strategies to Help Non-Expert Stakeholders. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems (Glasgow, Scotland Uk) (CHI ’19) . Association for Computing Machinery, New York, NY, USA, 1–12
Hao-Fei Cheng, Ruotong Wang, Zheng Zhang, Fiona O’Connell, Terrance Gray, F. Maxwell Harper, and Haiyi Zhu. 2019 · 2019
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers) , Jill Burstein, Christy Doran, and Thamar Solorio (Eds.). Association for Computational Linguistics, Minneapolis, Minnesota, 4171–4186
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
The principles and limits of algorithm-in-the-loop decision making
Ben Green and Yiling Chen. 2019 · 2019
Earlier work this paper cites.
The Impact of the Gricean Maxims of Quality, Quantity and Manner in Chatbots. In 2019 International Conference on Information and Digital Technologies (IDT) . IEEE, Zilina, Slovakia, 180–189
Baptiste Jacquet, Alexandre Hullin, Jean Baratgin, and Frank Jamet. 2019 · 2019
Earlier work this paper cites.
Let Me Explain: Impact of Personal and Impersonal Explanations on Trust in Recommender Systems. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems (Glasgow, Scotland Uk) (CHI ’19) . Association for Computing Machinery, New York, NY, USA, 1–12
Johannes Kunkel, Tim Donkers, Lisa Michael, Catalin-Mihai Barbu, and Jürgen Ziegler. 2019 · 2019
Cited alongside, same era.
A slow algorithm improves users’ assessments of the algorithm’s accuracy
Joon Sung Park, Rick Barber, Alex Kirlik, and Karrie Karahalios. 2019 · 2019
Cited alongside, same era.
Proxy Tasks and Subjective Measures Can Be Misleading in Evaluating Explainable AI Systems. In Proceedings of the 25th International Conference on Intelligent User Interfaces (Cagliari, Italy) (IUI ’20) . Association for Computing Machinery, New York, NY, USA, 454–464
Zana Buçinca, Phoebe Lin, Krzysztof Z. Gajos, and Elena L. Glassman. 2020 · 2020
Cited alongside, same era.
A Case for Humans-in-the-Loop: Decisions in the Presence of Erroneous Algorithmic Scores. In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI ’20) . Association for Computing Machinery, New York, NY, USA, 1–12
Bartscore: Evaluating generated text as text generation
Weizhe Yuan, Graham Neubig, and Pengfei Liu. 2021 · 2021
Later among the works it cites.
A test for evaluating performance in human-computer systems
Andres Campero, Michelle Vaccaro, Jaeyoon Song, Haoran Wen, Abdullah Almaatouq, and Thomas W Malone. 2022 · 2022
Later among the works it cites.
Scaling Instruction-Finetuned Language Models
Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Eric Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, Albert Webson, Shixiang Shane Gu, Zhuyun Dai, Mirac Suzgun, Xinyun Chen, Aakanksha Chowdhery, Sharan Narang, Gaurav Mishra, Adams Yu, Vincent Zhao, Yanping Huang, Andrew Dai, Hongkun Yu, Slav Petrov, Ed H. Chi, Jeff Dean, Jacob Devlin, Adam Roberts, Denny Zhou, Quoc V. Le, and Jason Wei. 2022 · 2022
Later among the works it cites.
Measuring Faithfulness of Abstractive Summaries. In Proceedings of the 18th Conference on Natural Language Processing (KONVENS 2022) , Robin Schaefer, Xiaoyu Bai, Manfred Stede, and Torsten Zesch (Eds.). KONVENS 2022 Organizers, Potsdam, Germany, 63–73
Tim Fischer, Steffen Remus, and Chris Biemann. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Maria De-Arteaga, Riccardo Fogliato, and Alexandra Chouldechova. 2020 · 2020
Cited alongside, same era.
Don‘t Stop Pretraining: Adapt Language Models to Domains and Tasks. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel Tetreault (Eds.). Association for Computational Linguistics, Online, 8342–8360
Suchin Gururangan, Ana Marasović, Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, and Noah A. Smith. 2020 · 2020
Cited alongside, same era.
Interpreting Interpretability: Understanding Data Scientists’ Use of Interpretability Tools for Machine Learning. In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI ’20) . Association for Computing Machinery, New York, NY, USA, 1–14
Harmanpreet Kaur, Harsha Nori, Samuel Jenkins, Rich Caruana, Hanna Wallach, and Jennifer Wortman Vaughan. 2020 · 2020
Cited alongside, same era.
On Faithfulness and Factuality in Abstractive Summarization. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL) . Association for Computational Linguistics, Online, 5075–5086
Joshua Maynez, Shashi Narayan, Bernd Bohnet, and Ryan Thomas McDonald. 2020 · 2020
Cited alongside, same era.
Rapid Trust Calibration through Interpretable and Uncertainty-Aware AI
Richard Tomsett, Alun Preece, Dave Braines, Federico Cerutti, Supriyo Chakraborty, Mani Srivastava, Gavin Pearson, and Lance Kaplan. 2020 · 2020
Cited alongside, same era.
How Do Visual Explanations Foster End Users’ Appropriate Trust in Machine Learning?. In Proceedings of the 25th International Conference on Intelligent User Interfaces (Cagliari, Italy) (IUI ’20) . Association for Computing Machinery, New York, NY, USA, 189–201
Fumeng Yang, Zhuanyi Huang, Jean Scholtz, and Dustin L. Arendt. 2020 · 2020
Cited alongside, same era.
Effect of confidence and explanation on accuracy and trust calibration in AI-assisted decision making. In Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency (Barcelona, Spain) (FAT* ’20) . Association for Computing Machinery, New York, NY, USA, 295–305
Yunfeng Zhang, Q. Vera Liao, and Rachel K. E. Bellamy. 2020 · 2020
Cited alongside, same era.
Does explainable artificial intelligence improve human decision-making?. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 35. AAAI Press, Palo Alto, California, USA, 6618–6626
Yasmeen Alufaisan, Laura R Marusich, Jonathan Z Bakdash, Yan Zhou, and Murat Kantarcioglu. 2021 · 2021
Cited alongside, same era.
Ai-assisted human labeling: Batching for efficiency without overreliance
Zahra Ashktorab, Michael Desmond, Josh Andres, Michael Muller, Narendra Nath Joshi, Michelle Brachman, Aabhas Sharma, Kristina Brimijoin, Qian Pan, Christine T Wolf, et al · 2021
Cited alongside, same era.
Analyzing and Evaluating Faithfulness in Dialogue Summarization. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing , Yoav Goldberg, Zornitsa Kozareva, and Yue Zhang (Eds.). Association for Computational Linguistics, Abu Dhabi, United Arab Emirates, 4897–4908
Bin Wang, Chen Zhang, Yan Zhang, Yiming Chen, and Haizhou Li. 2022 · 2022
Later among the works it cites.
Advancing Human-AI Complementarity: The Impact of User Expertise and Algorithmic Tuning on Joint Decision Making
Kori Inkpen, Shreya Chappidi, Keri Mallari, Besmira Nushi, Divya Ramesh, Pietro Michelucci, Vani Mandava, LibuŠe Hannah VepŘek, and Gabrielle Quinn. 2023 · 2023
Later among the works it cites.
Co-Writing with Opinionated Language Models Affects Users’ Views. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (Hamburg, Germany) (CHI ’23) . Association for Computing Machinery, New York, NY, USA, Article 111, 15 pages
Maurice Jakesch, Advait Bhat, Daniel Buschek, Lior Zalmanson, and Mor Naaman. 2023a · 2023
Later among the works it cites.
Human heuristics for AI-generated language are flawed
Maurice Jakesch, Jeffrey T Hancock, and Mor Naaman. 2023b · 2023
Later among the works it cites.
An Empirical Evaluation of Predicted Outcomes as Explanations in Human-AI Decision-Making. In Machine Learning and Principles and Practice of Knowledge Discovery in Databases , Irena Koprinska, Paolo Mignone, Riccardo Guidotti, Szymon Jaroszewicz, Holger Fröning, Francesco Gullo, Pedro M. Ferreira, Damian Roqueiro, Gaia Ceddia, Slawomir Nowaczyk, João Gama, Rita Ribeiro, Ricard Gavaldà, Elio Masciari, Zbigniew Ras, Ettore Ritacco, Francesca Naretto, Andreas Theissler, Przemyslaw Biecek, Wouter Verbeke, Gregor Schiele, Franz Pernkopf, Michaela Blott, Ilaria Bordino, Ivan Luciano Danesi, Giovanni Ponti, Lorenzo Severini, Annalisa Appice, Giuseppina Andresini, Ibéria Medeiros, Guilherme Graça, Lee Cooper, Naghmeh Ghazaleh, Jonas Richiardi, Diego Saldana, Konstantinos Sechidis, Arif Canakoglu, Sara Pido, Pietro Pinoli, Albert Bifet, and Sepideh Pashami (Eds.). Springer Nature Switzerland, Cham, 353–368
Johannes Jakubik, Jakob Schöffer, Vincent Hoge, Michael Vössing, and Niklas Kühl. 2023 · 2023
Later among the works it cites.
From ChatGPT to FactGPT: A Participatory Design Study to Mitigate the Effects of Large Language Model Hallucinations on Users. In Proceedings of Mensch Und Computer 2023 (Rapperswil, Switzerland) (MuC ’23) . Association for Computing Machinery, New York, NY, USA, 81–90
Florian Leiser, Sven Eckhardt, Merlin Knaeble, Alexander Maedche, Gerhard Schwabe, and Ali Sunyaev. 2023 · 2023
Later among the works it cites.
Understanding Uncertainty: How Lay Decision-makers Perceive and Interpret Uncertainty in Human-AI Decision Making. In Proceedings of the 28th International Conference on Intelligent User Interfaces (Sydney, NSW, Australia) (IUI ’23) . Association for Computing Machinery, New York, NY, USA, 379–396
Snehal Prabhudesai, Leyao Yang, Sumit Asthana, Xun Huan, Q. Vera Liao, and Nikola Banovic. 2023 · 2023
Later among the works it cites.
Appropriate Reliance on AI Advice: Conceptualization and the Effect of Explanations. In Proceedings of the 28th International Conference on Intelligent User Interfaces (Sydney, NSW, Australia) (IUI ’23) . Association for Computing Machinery, New York, NY, USA, 410–422
Max Schemmer, Niklas Kuehl, Carina Benz, Andrea Bartos, and Gerhard Satzger. 2023 · 2023
Later among the works it cites.
Defamation in the Age of Artificial Intelligence
Leslie Y Garfield Tenzer. 2023 · 2023
Later among the works it cites.
Explanations Can Reduce Overreliance on AI Systems During Decision-Making
Helena Vasconcelos, Matthew Jörke, Madeleine Grunde-McLaughlin, Tobias Gerstenberg, Michael S. Bernstein, and Ranjay Krishna. 2023 · 2023
Later among the works it cites.
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al · 2023
Later among the works it cites.
Synthetic Lies: Understanding AI-Generated Misinformation and Evaluating Algorithmic and Human Solutions. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (Hamburg, Germany) (CHI ’23) . Association for Computing Machinery, New York, NY, USA, Article 436, 20 pages
Jiawei Zhou, Yixuan Zhang, Qianni Luo, Andrea G Parker, and Munmun De Choudhury. 2023 · 2023
Later among the works it cites.
The Who in XAI: How AI Background Shapes Perceptions of AI Explanations. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI ’24) . Association for Computing Machinery, New York, NY, USA, Article 316, 32 pages
Upol Ehsan, Samir Passi, Q. Vera Liao, Larry Chan, I-Hsiang Lee, Michael Muller, and Mark O Riedl. 2024 · 2024
Closest in time.
The imitation game: Detecting human and AI-generated texts in the era of ChatGPT and BARD
Khaled Hayawi, Hossain Shahriar, and Sibi S Mathew. 2024 · 2024
Closest in time.
Recommender Systems in the Era of Large Language Models (LLMs)
Zihuai Zhao, Wenqi Fan, Jiatong Li, Yunqing Liu, Xiaowei Mei, Yiqi Wang, Zhen Wen, Fei Wang, Xiangyu Zhao, Jiliang Tang, and Qing Li. 2024 · 2024
Closest in time.