Fetching the paper…
Reading the bibliography…
Assessing how well a large language model (LLM) understands human, rather than merely text, remains an open challenge.
Client-centered/person-centered approach to therapy
Carl R Rogers · 2001
Earlier work this paper cites.
Congruence/genuineness
Gregory G Kolden, Marjorie H Klein, Chia-Chiang Wang, and Sara B Austin · 2011
Earlier work this paper cites.
The relationship inventory: A complete resource and guide
Godfrey T Barrett-Lennard · 2015
Earlier work this paper cites.
Towards emotional support dialog systems
Siyang Liu, Chujie Zheng, Orianna Demasi, Sahand Sabour, Yu Li, Zhou Yu, Yong Jiang, and Minlie Huang · 2021
Earlier work this paper cites.
Knowledge bridging for empathetic dialogue generation
Qintong Li, Piji Li, Zhaochun Ren, Pengjie Ren, and Zhumin Chen · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Earlier work this paper cites.
Neural theory-of-mind? on the limits of social intelligence in large LMs
Maarten Sap, Ronan Le Bras, Daniel Fried, and Yejin Choi · 2022
Earlier work this paper cites.
Chateval: Towards better llm-based evaluators through multi-agent debate
Chi-Min Chan, Weize Chen, Yusheng Su, Jianxuan Yu, Wei Xue, Shanghang Zhang, Jie Fu, and Zhiyuan Liu · 2023
Earlier work this paper cites.
Development and validation of a 12-item version of the barrett-lennard relationship inventory (bl ri: mini) using item response theory
Shun Chen, Faith Liao, David Murphy, and Stephen Joseph · 2023
Earlier work this paper cites.
Alpacafarm: A simulation framework for methods that learn from human feedback
Yann Dubois, Chen Xuechen Li, Rohan Taori, Tianyi Zhang, Ishaan Gulrajani, Jimmy Ba, Carlos Guestrin, Percy S Liang, and Tatsunori B Hashimoto · 2023
Earlier work this paper cites.
Hi-tom: A benchmark for evaluating higher-order theory of mind reasoning in large language models
Yinghui He, Yufan Wu, Yilin Jia, Rada Mihalcea, Yulong Chen, and Naihao Deng · 2023
Earlier work this paper cites.
Fantom: A benchmark for stress-testing machine theory of mind in interactions
Hyunwoo Kim, Melanie Sclar, Xuhui Zhou, Ronan Le Bras, Gunhee Kim, Yejin Choi, and Maarten Sap · 2023
Earlier work this paper cites.
Rlaif vs. rlhf: Scaling reinforcement learning from human feedback with ai feedback
Harrison Lee, Samrat Phatale, Hassan Mansoor, Thomas Mesnard, Johan Ferret, Kellie Lu, Colton Bishop, Ethan Hall, Victor Carbune, Abhinav Rastogi, et al · 2023
Earlier work this paper cites.
Let’s verify step by step
Hunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe · 2023
Earlier work this paper cites.
Training models to generate, recognize, and reframe unhelpful thoughts
Mounica Maddela, Megan Ung, Jing Xu, Andrea Madotto, Heather Foran, and Y-Lan Boureau · 2023
Earlier work this paper cites.
Orca: Progressive learning from complex explanation traces of gpt-4
Subhabrata Mukherjee, Arindam Mitra, Ganesh Jawahar, Sahaj Agarwal, Hamid Palangi, and Ahmed Awadallah · 2023
Earlier work this paper cites.
Eq-bench: An emotional intelligence benchmark for large language models
Samuel J Paech · 2023
Earlier work this paper cites.
Large language models are effective text rankers with pairwise ranking prompting
Zhen Qin, Rolf Jagerman, Kai Hui, Honglei Zhuang, Junru Wu, Le Yan, Jiaming Shen, Tianqi Liu, Jialu Liu, Donald Metzler, et al · 2023
Earlier work this paper cites.
Branch-solve-merge improves large language model evaluation and generation
Swarnadeep Saha, Omer Levy, Asli Celikyilmaz, Mohit Bansal, Jason Weston, and Xian Li · 2023
Earlier work this paper cites.
Clever hans or neural theory of mind? stress testing social reasoning in large language models
Natalie Shapira, Mosh Levy, Seyed Hossein Alavi, Xuhui Zhou, Yejin Choi, Yoav Goldberg, Maarten Sap, and Vered Shwartz · 2023
Cited alongside, same era.
Reflexion: Language agents with verbal reinforcement learning
Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao · 2023
Cited alongside, same era.
Stanford alpaca: An instruction-following llama model
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto · 2023
Cited alongside, same era.
Macgyver: Are large language models creative problem solvers?
Yufei Tian, Abhilasha Ravichander, Lianhui Qin, Ronan Le Bras, Raja Marjieh, Nanyun Peng, Yejin Choi, Thomas L Griffiths, and Faeze Brahman · 2023
Cited alongside, same era.
Wizardlm: Empowering large language models to follow complex instructions
Infobench: Evaluating instruction following ability in large language models
Yiwei Qin, Kaiqiang Song, Yebowen Hu, Wenlin Yao, Sangwoo Cho, Xiaoyang Wang, Xuansheng Wu, Fei Liu, Pengfei Liu, and Dong Yu · 2024
Later among the works it cites.
Emobench: Evaluating the emotional intelligence of large language models
Sahand Sabour, Siyang Liu, Zheyuan Zhang, June M Liu, Jinfeng Zhou, Alvionna S Sunaryo, Juanzi Li, Tatia Lee, Rada Mihalcea, and Minlie Huang · 2024
Later among the works it cites.
Personagym: Evaluating persona agents and llms
Vinay Samuel, Henry Peng Zou, Yue Zhou, Shreyas Chaudhari, Ashwin Kalyan, Tanmay Rajpurohit, Ameet Deshpande, Karthik Narasimhan, and Vishvak Murahari · 2024
Later among the works it cites.
Rehearsal: Simulating conflict to teach conflict resolution
Omar Shaikh, Valentino Emil Chai, Michele Gelfand, Diyi Yang, and Michael S Bernstein · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Can Xu, Qingfeng Sun, Kai Zheng, Xiubo Geng, Pu Zhao, Jiazhan Feng, Chongyang Tao, and Daxin Jiang · 2023
Cited alongside, same era.
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al · 2023
Cited alongside, same era.
Facilitating multi-turn emotional support conversation with positive emotion elicitation: A reinforcement learning approach
Jinfeng Zhou, Zhuang Chen, Bo Wang, and Minlie Huang · 2023
Cited alongside, same era.
CASE: Aligning coarse-to-fine cognition and affection for empathetic response generation
Jinfeng Zhou, Chujie Zheng, Bo Wang, Zheng Zhang, and Minlie Huang · 2023
Cited alongside, same era.
Judgelm: Fine-tuned large language models are scalable judges
Lianghui Zhu, Xinggang Wang, and Xinlong Wang · 2023
Cited alongside, same era.
Chatbot arena: An open platform for evaluating llms by human preference
Wei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos, Tianle Li, Dacheng Li, Banghua Zhu, Hao Zhang, Michael Jordan, Joseph E Gonzalez, et al · 2024
Cited alongside, same era.
V-star: Training verifiers for self-taught reasoners
Arian Hosseini, Xingdi Yuan, Nikolay Malkin, Aaron Courville, Alessandro Sordoni, and Rishabh Agarwal · 2024
Cited alongside, same era.
Trustagent: Towards safe and trustworthy llm-based agents through agent constitution
Wenyue Hua, Xianjun Yang, Mingyu Jin, Zelong Li, Wei Cheng, Ruixiang Tang, and Yongfeng Zhang · 2024
Cited alongside, same era.
James WA Strachan, Dalila Albergo, Giulia Borghini, Oriana Pansardi, Eugenio Scaliti, Saurabh Gupta, Krati Saxena, Alessandro Rufo, Stefano Panzeri, Guido Manzi, et al · 2024
Later among the works it cites.
Charactereval: A chinese benchmark for role-playing conversational agent evaluation
Quan Tu, Shilong Fan, Zihang Tian, and Rui Yan · 2024
Later among the works it cites.
Sotopia- π \pi : Interactive learning of socially intelligent language agents
Ruiyi Wang, Haofei Yu, Wenxin Zhang, Zhengyang Qi, Maarten Sap, Graham Neubig, Yonatan Bisk, and Hao Zhu · 2024
Later among the works it cites.
Meta-rewarding language models: Self-improving alignment with llm-as-a-meta-judge
Tianhao Wu, Weizhe Yuan, Olga Golovneva, Jing Xu, Yuandong Tian, Jiantao Jiao, Jason Weston, and Sainbayar Sukhbaatar · 2024
Later among the works it cites.
Enhancing llm reasoning via critique models with test-time and training-time supervision
Zhiheng Xi, Dingwen Yang, Jixuan Huang, Jiafu Tang, Guanyu Li, Yiwen Ding, Wei He, Boyang Hong, Shihan Do, Wenyu Zhan, et al · 2024
Later among the works it cites.
Academically intelligent llms are not necessarily socially intelligent
Ruoxi Xu, Hongyu Lin, Xianpei Han, Le Sun, and Yingfei Sun · 2024
Later among the works it cites.
Social skill training with large language models
Diyi Yang, Caleb Ziems, William Held, Omar Shaikh, Michael S Bernstein, and John Mitchell · 2024
Later among the works it cites.
Agent-as-a-judge: Evaluate agents with agents
Mingchen Zhuge, Changsheng Zhao, Dylan Ashley, Wenyi Wang, Dmitrii Khizbullin, Yunyang Xiong, Zechun Liu, Ernie Chang, Raghuraman Krishnamoorthi, Yuandong Tian, et al · 2024
Later among the works it cites.
Spc: Evolving self-play critic via adversarial games for llm reasoning
Jiaqi Chen, Bang Zhang, Ruotian Ma, Peisong Wang, Xiaodan Liang, Zhaopeng Tu, Xiaolong Li, and Kwan-Yee K Wong · 2025
Closest in time.
Are autonomous web agents good testers?
Antoine Chevrot, Alexandre Vernotte, Jean-Rémy Falleri, Xavier Blanc, and Bruno Legeard · 2025
Closest in time.
Competing large language models in multi-agent gaming environments
Jen-tse Huang, Eric John Li, Man Ho LAM, Tian Liang, Wenxuan Wang, Youliang Yuan, Wenxiang Jiao, Xing Wang, Zhaopeng Tu, and Michael Lyu · 2025
Closest in time.
Agent-as-judge for factual summarization of long narratives
Yeonseok Jeong, Minsoo Kim, Seung-won Hwang, and Byung-Hak Kim · 2025
Closest in time.
Dancing with critiques: Enhancing llm reasoning with stepwise natural language self-critique
Yansi Li, Jiahao Xu, Tian Liang, Xingyu Chen, Zhiwei He, Qiuzhi Liu, Rui Wang, Zhuosheng Zhang, Zhaopeng Tu, Haitao Mi, et al · 2025
Closest in time.
Scaling test-time compute optimally can be more effective than scaling LLM parameters
Charlie Victor Snell, Jaehoon Lee, Kelvin Xu, and Aviral Kumar · 2025
Closest in time.
Shenghan Wu, Yang Deng, Yimo Zhu, Wynne Hsu, and Mong Li Lee · 2025
Closest in time.