Fetching the paper…
Reading the bibliography…
The Wizard of Oz (WoZ) method is a widely adopted research approach where a human Wizard ``role-plays'' a not readily available technology and interacts with participants to elicit user behaviors and probe the design space.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 1901
Earlier work this paper cites.
A new readability yardstick
Rudolph Flesch. 1948 · 1948
Earlier work this paper cites.
An Empirical Methodology for Writing User-Friendly Natural Language Computer Applications. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems (Boston, Massachusetts, USA) (CHI ’83) . Association for Computing Machinery, New York, NY, USA, 193–196
J. F. Kelley. 1983 · 1983
Earlier work this paper cites.
An Iterative Design Methodology for User-Friendly Natural Language Office Information Applications
John F. Kelley. 1984 · 1984
Earlier work this paper cites.
The Rapid Development of User Interfaces: Experience with the Wizard of OZ Method
Paul Green and Lisa Wei-Haas. 1985 · 1985
Earlier work this paper cites.
Wizard of Oz Studies: Why and How. In Proceedings of the 1st International Conference on Intelligent User Interfaces (Orlando, Florida, USA) (IUI ’93) . Association for Computing Machinery, New York, NY, USA, 193–200
Nils Dahlbäck, Arne Jönsson, and Lars Ahrenberg. 1993 · 1993
Earlier work this paper cites.
Prototyping an Intelligent Agent through Wizard of Oz. In Proceedings of the INTERACT ’93 and CHI ’93 Conference on Human Factors in Computing Systems (Amsterdam, The Netherlands) (CHI ’93) . Association for Computing Machinery, New York, NY, USA, 277–284
David Maulsby, Saul Greenberg, and Richard Mander. 1993 · 1993
Earlier work this paper cites.
MAC/FAC: A model of similarity-based retrieval
Kenneth D Forbus, Dedre Gentner, and Keith Law. 1995 · 1995
Earlier work this paper cites.
Suede: A Wizard of Oz Prototyping Tool for Speech User Interfaces. In Proceedings of the 13th Annual ACM Symposium on User Interface Software and Technology (San Diego, California, USA) (UIST ’00) . Association for Computing Machinery, New York, NY, USA, 1–10
Scott R. Klemmer, Anoop K. Sinha, Jack Chen, James A. Landay, Nadeem Aboobaker, and Annie Wang. 2000 · 2000
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting of the Association for Computational Linguistics . Association for Computational Linguistics, Philadelphia, Pennsylvania, USA, 311–318
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
ROUGE: A Package for Automatic Evaluation of summaries. In Proceedings of the ACL Workshop: Text Summarization Braches Out . Association for Computational Linguistics, Barcelona, Spain, 10
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Guanhua Zhang, Bing Bai, Junqi Zhang, Kun Bai, Conghui Zhu, and Tiejun Zhao. 2020a · 2004
Earlier work this paper cites.
Wizard of Oz Interfaces for Mixed Reality Applications. In CHI ’05 Extended Abstracts on Human Factors in Computing Systems (Portland, OR, USA) (CHI EA ’05) . Association for Computing Machinery, New York, NY, USA, 1339–1342
Steven Dow, Jaemin Lee, Christopher Oezbek, Blair MacIntyre, Jay David Bolter, and Maribeth Gandy. 2005 · 2005
Earlier work this paper cites.
Group Attention Control for Communication Robots with Wizard of OZ Approach. In Proceedings of the ACM/IEEE International Conference on Human-Robot Interaction (Arlington, Virginia, USA) (HRI ’07) . Association for Computing Machinery, New York, NY, USA, 121–128
Masahiro Shiomi, Takayuki Kanda, Satoshi Koizumi, Hiroshi Ishiguro, and Norihiro Hagita. 2007 · 2007
Earlier work this paper cites.
The Rating of Chessplayers: Past and Present
A.E. Elo. 2008 · 2008
Earlier work this paper cites.
RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models
Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A. Smith. 2020 · 2009
Earlier work this paper cites.
Design thinking
Hasso Plattner, Christoph Meinel, and Ulrich Weinberg. 2009 · 2009
Earlier work this paper cites.
The Oz of Wizard: Simulating the Human for Interaction Research. In Proceedings of the 4th ACM/IEEE International Conference on Human Robot Interaction (La Jolla, California, USA) (HRI ’09) . Association for Computing Machinery, New York, NY, USA, 101–108
Aaron Steinfeld, Odest Chadwicke Jenkins, and Brian Scassellati. 2009 · 2009
Earlier work this paper cites.
Wizard of Oz Experiments for a companion dialogue system: eliciting companionable conversation.. In In Proceedings of the Seventh conference on International Language Resources and Evaluation (LREC ’10) . European Language Resources Association (ELRA), Valletta, Malta, 5 pages
Nick Webb, David Benyon, Jay Bradley, Preben Hansen, and Oli Mival. 2010 · 2010
Earlier work this paper cites.
Professor–student rapport scale predicts student outcomes
Janie H Wilson, Rebecca G Ryan, and James L Pugh. 2010 · 2010
Earlier work this paper cites.
VADER: A Parsimonious Rule-Based Model for Sentiment Analysis of Social Media Text
C. Hutto and Eric Gilbert. 2014 · 2014
Earlier work this paper cites.
In-car multi-domain spoken dialogs: A wizard of oz study. In Proceedings of the EACL 2014 Workshop on Dialogue in Motion . Association for Computational Linguistics, Gothenburg, Sweden, 1–9
Sven Reichel, Ute Ehrlich, André Berton, and Michael Weber. 2014 · 2014
Earlier work this paper cites.
Towards a Chatbot for Digital Counselling. In Proceedings of the 31st British Computer Society Human Computer Interaction Conference (Sunderland, UK) (HCI ’17) . BCS Learning & Development Ltd., Swindon, GBR, Article 24, 7 pages
Gillian Cameron, David Cameron, Gavin Megaw, Raymond Bond, Maurice Mulvenna, Siobhan O’Neill, Cherie Armour, and Michael McTear. 2017 · 2017
Earlier work this paper cites.
How do you want your chatbot? An exploratory Wizard-of-Oz study with young, urban Indians. In Human-Computer Interaction-INTERACT 2017: 16th IFIP TC 13 International Conference, Mumbai, India, September 25–29, 2017, Proceedings, Part I 16 . Springer, Springer, Mumbai, India, 441–459
Indrani Medhi Thies, Nandita Menon, Sneha Magapu, Manisha Subramony, and Jacki O’neill. 2017 · 2017
Earlier work this paper cites.
Composite Task-Completion Dialogue Policy Learning via Hierarchical Deep Reinforcement Learning. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing . Association for Computational Linguistics, Stroudsburg, Pennsylvania, 2231––2240
Baolin Peng, Xiujun Li, Lihong Li, Jianfeng Gao, Asli Celikyilmaz, Sungjin Lee, and Kam-Fai Wong. 2017 · 2017
Earlier work this paper cites.
A New Chatbot for Customer Service on Social Media. In Proceedings of the 2017 CHI Conference on Human Factors in Computing Systems (Denver, Colorado, USA) (CHI ’17) . Association for Computing Machinery, New York, NY, USA, 3506–3510
Anbang Xu, Zhe Liu, Yufan Guo, Vibha Sinha, and Rama Akkiraju. 2017 · 2017
Earlier work this paper cites.
Automated Facilitation for Idea Platforms: Design and Evaluation of a Chatbot Prototype. In Proceedings of the International Conference on Information Systems - Bridging the Internet of People, Data, and Things 2018 (ICIS 2018) , Jan Pries-Heje, Sudha Ram, and Michael Rosemann (Eds.). Association for Information Systems, San Francisco, CA, USA, 9 pages
Navid Tavanapour and Eva A. C. Bittner. 2018 · 2018
Cited alongside, same era.
Understanding Chatbot-Mediated Task Management. In Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems (Montreal QC, Canada) (CHI ’18) . Association for Computing Machinery, New York, NY, USA, 1–6
Carlos Toxtli, Andrés Monroy-Hernández, and Justin Cranshaw. 2018 · 2018
Cited alongside, same era.
Supporting Workplace Detachment and Reattachment with Conversational Intelligence. In Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems (Montreal QC, Canada) (CHI ’18) . Association for Computing Machinery, New York, NY, USA, 1–13
Alex C. Williams, Harmanpreet Kaur, Gloria Mark, Anne Loomis Thompson, Shamsi T. Iqbal, and Jaime Teevan. 2018 · 2018
Cited alongside, same era.
Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90%* ChatGPT Quality
Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. 2023 · 2023
Later among the works it cites.
Can AI language models replace human participants?
Danica Dillion, Niket Tandon, Yuling Gu, and Kurt Gray. 2023 · 2023
Later among the works it cites.
Bias of AI-Generated Content: An Examination of News Produced by Large Language Models
Xiao Fang, Shangkun Che, Minjia Mao, Hongzhe Zhang, Ming Zhao, and Xiaohang Zhao. 2023 · 2023
Later among the works it cites.
Gptscore: Evaluate as you desire
Jinlan Fu, See-Kiong Ng, Zhengbao Jiang, and Pengfei Liu. 2023 · 2023
Later among the works it cites.
Evaluating Large Language Models in Generating Synthetic HCI Research Data: A Case Study. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (Hamburg, Germany) (CHI ’23) . Association for Computing Machinery, New York, NY, USA, Article 433, 19 pages
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Wizard of Oz Prototyping for Machine Learning Experiences. In Extended Abstracts of the 2019 CHI Conference on Human Factors in Computing Systems (Glasgow, Scotland Uk) (CHI EA ’19) . Association for Computing Machinery, New York, NY, USA, 1–6
Jacob T. Browne. 2019 · 2019
Cited alongside, same era.
Comparing Data from Chatbot and Web Surveys: Effects of Platform and Conversational Style on Survey Response Quality. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems (Glasgow, Scotland Uk) (CHI ’19) . Association for Computing Machinery, New York, NY, USA, 1–12
Soomin Kim, Joonhwan Lee, and Gahgene Gweon. 2019 · 2019
Cited alongside, same era.
The Woman Worked as a Babysitter: On Biases in Language Generation. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) , Kentaro Inui, Jing Jiang, Vincent Ng, and Xiaojun Wan (Eds.). Association for Computational Linguistics, Hong Kong, China, 3407–3412
Emily Sheng, Kai-Wei Chang, Premkumar Natarajan, and Nanyun Peng. [n. d.] · 2019
Cited alongside, same era.
Who Should Be My Teammates: Using a Conversational Agent to Understand Individuals and Help Teaming. In Proceedings of the 24th International Conference on Intelligent User Interfaces (Marina del Ray, California) (IUI ’19) . Association for Computing Machinery, New York, NY, USA, 437–447
Ziang Xiao, Michelle X. Zhou, and Wat-Tat Fu. 2019 · 2019
Cited alongside, same era.
Challenges in building intelligent open-domain dialog systems
Minlie Huang, Xiaoyan Zhu, and Jianfeng Gao. 2020 · 2020
Cited alongside, same era.
"I Hear You, I Feel You": Encouraging Deep Self-Disclosure through a Chatbot. In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI ’20) . Association for Computing Machinery, New York, NY, USA, 1–12
Yi-Chieh Lee, Naomi Yamashita, Yun Huang, and Wai Fu. 2020 · 2020
Cited alongside, same era.
Effects of Persuasive Dialogues: Testing Bot Identities and Inquiry Strategies. In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI ’20) . Association for Computing Machinery, New York, NY, USA, 1–13
Weiyan Shi, Xuewei Wang, Yoo Jung Oh, Jingwen Zhang, Saurav Sahay, and Zhou Yu. 2020 · 2020
Cited alongside, same era.
Tell me about yourself: Using an AI-powered chatbot to conduct conversational surveys with open-ended questions
Ziang Xiao, Michelle X Zhou, Q Vera Liao, Gloria Mark, Changyan Chi, Wenxi Chen, and Huahai Yang. 2020 · 2020
Cited alongside, same era.
Artificial intelligence chatbot behavior change model for designing artificial intelligence chatbots to promote physical activity and a healthy diet
Jingwen Zhang, Yoo Jung Oh, Patrick Lange, Zhou Yu, and Yoshimi Fukuoka. 2020b · 2020
Cited alongside, same era.
Perttu Hämäläinen, Mikke Tavast, and Anton Kunnari. 2023 · 2023
Later among the works it cites.
Ai language models cannot replace human research participants
Jacqueline Harding, William D’Alessandro, NG Laskowski, and Robert Long. 2023 · 2023
Later among the works it cites.
The use of ChatGPT and other large language models in surgical science
Boris V Janssen, Geert Kazemier, and Marc G Besselink. 2023 · 2023
Later among the works it cites.
Understanding the Benefits and Challenges of Deploying Conversational AI Leveraging Large Language Models for Public Health Intervention. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (Hamburg, Germany) (CHI ’23) . Association for Computing Machinery, New York, NY, USA, Article 18, 16 pages
Eunkyung Jo, Daniel A. Epstein, Hyunhoon Jung, and Young-Ho Kim. 2023 · 2023
Later among the works it cites.
Working With AI to Persuade: Examining a Large Language Model’s Ability to Generate Pro-Vaccination Messages
Elise Karinshak, Sunny Xun Liu, Joon Sung Park, and Jeffrey T Hancock. 2023 · 2023
Later among the works it cites.
Understanding People’s Perception and Usage of Plug-in Electric Hybrids. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (Hamburg, Germany) (CHI ’23) . Association for Computing Machinery, New York, NY, USA, Article 201, 21 pages
Matthew L Lee, Scott Carter, Rumen Iliev, Nayeli Suseth Bravo, Monica P Van, Laurent Denoue, Everlyne Kimani, Alexandre L. S. Filipowicz, David A. Shamma, Kate A Sieck, Candice Hogan, and Charlene C. Wu. 2023 · 2023
Later among the works it cites.
G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment
Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023 · 2023
Later among the works it cites.
Wizard of Oz Prototyping for Interactive Spatial Augmented Reality in HCI Education: Experiences with Rapid Prototyping for Interactive Spatial Augmented Reality. In Extended Abstracts of the 2023 CHI Conference on Human Factors in Computing Systems (Hamburg, Germany) (CHI EA ’23) . Association for Computing Machinery, New York, NY, USA, Article 407, 10 pages
Danica Mast, Alex Roidl, and Antti Jylha. 2023 · 2023
Later among the works it cites.
Using In-Context Learning to Improve Dialogue Safety
Nicholas Meade, Spandana Gella, Devamanyu Hazarika, Prakhar Gupta, Di Jin, Siva Reddy, Yang Liu, and Dilek Hakkani-Tür. 2023 · 2023
Later among the works it cites.
API Reference
OpenAI. 2023 · 2023
Later among the works it cites.
Generative Agents: Interactive Simulacra of Human Behavior
Joon Sung Park, Joseph C. O’Brien, Carrie J. Cai, Meredith Ringel Morris, Percy Liang, and Michael S. Bernstein. 2023 · 2023
Later among the works it cites.
Personality Traits in Large Language Models
Greg Serapio-García, Mustafa Safdari, Clément Crepy, Luning Sun, Stephen Fitz, Peter Romero, Marwa Abdulhai, Aleksandra Faust, and Maja Matarić. 2023 · 2023
Later among the works it cites.
Role-Play with Large Language Models
Murray Shanahan, Kyle McDonell, and Laria Reynolds. 2023 · 2023
Later among the works it cites.
"Kelly is a Warm Person, Joseph is a Role Model": Gender Biases in LLM-Generated Reference Letters
Yixin Wan, George Pu, Jiao Sun, Aparna Garimella, Kai-Wei Chang, and Nanyun Peng. 2023 · 2023
Later among the works it cites.
Large Language Models are not Fair Evaluators
Peiyi Wang, Lei Li, Liang Chen, Zefan Cai, Dawei Zhu, Binghuai Lin, Yunbo Cao, Qi Liu, Tianyu Liu, and Zhifang Sui. 2023 · 2023
Later among the works it cites.
Inform the Uninformed: Improving Online Informed Consent Reading with an AI-Powered Chatbot. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (Hamburg, Germany, </conf-loc>) (CHI ’23) . Association for Computing Machinery, New York, NY, USA, Article 112, 17 pages
Ziang Xiao, Tiffany Wenting Li, Karrie Karahalios, and Hari Sundaram. 2023 · 2023
Later among the works it cites.
Tree of Thoughts: Deliberate Problem Solving with Large Language Models
Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L. Griffiths, Yuan Cao, and Karthik Narasimhan. 2023 · 2023
Later among the works it cites.
Evaluating Large Language Models at Evaluating Instruction Following
Zhiyuan Zeng, Jiatong Yu, Tianyu Gao, Yu Meng, Tanya Goyal, and Danqi Chen. 2023 · 2023
Later among the works it cites.
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric. P Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2023 · 2023
Later among the works it cites.
Public Perceptions of Gender Bias in Large Language Models: Cases of ChatGPT and Ernie
Kyrie Zhixuan Zhou and Madelyn Rose Sanfilippo. 2023 · 2023
Later among the works it cites.
The illusion of artificial inclusion
William Agnew, A Stevie Bergman, Jennifer Chien, Mark Díaz, Seliem El-Sayed, Jaylen Pittman, Shakir Mohamed, and Kevin R McKee. 2024 · 2024
Closest in time.