Fetching the paper…
Reading the bibliography…
The role-play ability of Large Language Models (LLMs) has emerged as a popular research direction.
Computing machinery and intelligence
Alan M Turing · 1950
Earlier work this paper cites.
Can gpt-3 pass a writer’s turing test?
Katherine Elkins and Jon Chun · 2020
Earlier work this paper cites.
Turingbench: A benchmark environment for turing test in the age of neural text generation
Adaku Uchendu, Zeyu Ma, Thai Le, Rui Zhang, and Dongwon Lee · 2021
Earlier work this paper cites.
Calibrate before use: Improving few-shot performance of language models
Zihao Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh · 2021
Earlier work this paper cites.
Education in the era of generative artificial intelligence (ai): Understanding the potential benefits of chatgpt in promoting teaching and learning
David Baidoo-Anu and Leticia Owusu Ansah · 2023
Earlier work this paper cites.
Chateval: Towards better llm-based evaluators through multi-agent debate
Chi-Min Chan, Weize Chen, Yusheng Su, Jianxuan Yu, Wei Xue, Shanghang Zhang, Jie Fu, and Zhiyuan Liu · 2023
Earlier work this paper cites.
Emotionally numb or empathetic? evaluating how llms feel using emotionbench
Jen-tse Huang, Man Ho Lam, Eric John Li, Shujie Ren, Wenxuan Wang, Wenxiang Jiao, Zhaopeng Tu, and Michael R Lyu · 2023
Earlier work this paper cites.
Human or not? a gamified approach to the turing test
Daniel Jannai, Amos Meron, Barak Lenz, Yoav Levine, and Yoav Shoham · 2023
Earlier work this paper cites.
Is chatgpt a good translator? a preliminary study
Wenxiang Jiao, Wenxuan Wang, Jen-tse Huang, Xing Wang, and Zhaopeng Tu · 2023
Earlier work this paper cites.
Assessing the accuracy and reliability of ai-generated medical responses: an evaluation of the chat-gpt model
Douglas Johnson, Rachel Goodman, J Patrinely, Cosby Stone, Eli Zimmerman, Rebecca Donald, Sam Chang, Sean Berkowitz, Avni Finn, Eiman Jahangir, et al · 2023
Cited alongside, same era.
Does gpt-4 pass the turing test?
Cameron Jones and Benjamin Bergen · 2023
Cited alongside, same era.
Better zero-shot reasoning with role-play prompting
Aobo Kong, Shiwan Zhao, Hao Chen, Qicheng Li, Yong Qin, Ruiqi Sun, and Xin Zhou · 2023
Cited alongside, same era.
Chatgpt and other large language models as evolutionary engines for online interactive collaborative game design
Pier Luca Lanzi and Daniele Loiacono · 2023
Cited alongside, same era.
Chatharuhi: Reviving anime character in reality via large language model
Use chat gpt to solve programming bugs
Nigar M Shafiq Surameery and Mohammed Y Shakor · 2023
Later among the works it cites.
Chatgpt or grammarly? evaluating chatgpt on grammatical error correction benchmark
Haoran Wu, Wenxuan Wang, Yuxuan Wan, Wenxiang Jiao, and Michael Lyu · 2023
Later among the works it cites.
Characterglm: Customizing chinese conversational ai characters with large language models
Jinfeng Zhou, Zhuang Chen, Dazhen Wan, Bosi Wen, Yi Song, Jifan Yu, Yongkang Huang, Libiao Peng, Jiaming Yang, Xiyao Xiao, et al · 2023
Later among the works it cites.
Large language models for information retrieval: A survey
Yutao Zhu, Huaying Yuan, Shuting Wang, Jiongnan Liu, Wenhan Liu, Chenlong Deng, Zhicheng Dou, and Ji-Rong Wen · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cheng Li, Ziang Leng, Chenxi Yan, Junyi Shen, Hao Wang, Weishi Mi, Yaying Fei, Xiaoyang Feng, Song Yan, HaoSheng Wang, et al · 2023
Cited alongside, same era.
Introducing gpts
OpenAI · 2023
Cited alongside, same era.
Large language models and the reverse turing test
Terrence J Sejnowski · 2023
Cited alongside, same era.
Role play with large language models
Murray Shanahan, Kyle McDonell, and Laria Reynolds · 2023
Cited alongside, same era.
Character-llm: A trainable agent for role-playing
Yunfan Shao, Linyang Li, Junqi Dai, and Xipeng Qiu · 2023
Cited alongside, same era.
Jen-tse Huang, Eric John Li, Man Ho Lam, Tian Liang, Wenxuan Wang, Youliang Yuan, Wenxiang Jiao, Xing Wang, Zhaopeng Tu, and Michael R Lyu
Cited in the paper.
On the humanity of conversational ai: Evaluating the psychological portrayal of llms
Jen-tse Huang, Wenxuan Wang, Eric John Li, Man Ho LAM, Shujie Ren, Youliang Yuan, Wenxiang Jiao, Zhaopeng Tu, and Michael Lyu
Cited in the paper.
Incharacter: Evaluating personality fidelity in role-playing agents through psychological interviews
Xintao Wang, Yunze Xiao, Jen-tse Huang, Siyu Yuan, Rui Xu, Haoran Guo, Quan Tu, Yaying Fei, Ziang Leng, Wei Wang, Jiangjie Chen, Cheng Li, and Yanghua Xiao
Cited in the paper.
Guiming Hardy Chen, Shunian Chen, Ziche Liu, Feng Jiang, and Benyou Wang · 2024
Closest in time.
Evalullm: Llm assisted evaluation of generative outputs
Michael Desmond, Zahra Ashktorab, Qian Pan, Casey Dugan, and James M Johnson · 2024
Closest in time.
Legalbench: A collaboratively built benchmark for measuring legal reasoning in large language models
Neel Guha, Julian Nyarko, Daniel Ho, Christopher Ré, Adam Chilton, Alex Chohlas-Wood, Austin Peters, Brandon Waldon, Daniel Rockmore, Diego Zambrano, et al · 2024
Closest in time.
A & b== b & a: Triggering logical reasoning failures in large language models
Yuxuan Wan, Wenxuan Wang, Yiliu Yang, Youliang Yuan, Jen-tse Huang, Pinjia He, Wenxiang Jiao, and Michael R Lyu · 2024
Closest in time.
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al · 2024
Closest in time.