Fetching the paper…
Reading the bibliography…
To evaluate Large Language Models (LLMs) for question answering (QA), traditional methods typically focus on assessing single-turn responses to given questions.
Artificial life meets entertainment: lifelike autonomous agents
P. Maes · 1995
Earlier work this paper cites.
Intelligent agents: theory and practice
M. Wooldridge and N. R. Jennings · 1995
Earlier work this paper cites.
About face 2: The essentials of interaction design, 2014
A. Cooper, R. Reimann, D. Cronin, and C. Noessel · 2014
Earlier work this paper cites.
People-centric natural language processing
D. Bamman · 2015
Earlier work this paper cites.
O. Vinyals and Q. V. Le · 2015
Earlier work this paper cites.
A persona-based neural conversation model
J. Li, M. Galley, C. Brockett, G. P. Spithourakis, J. Gao, and B. Dolan · 2016
Earlier work this paper cites.
How not to evaluate your dialogue system: An empirical study of unsupervised evaluation metrics for dialogue response generation
C.-W. Liu, R. Lowe, I. Serban, M. Noseworthy, L. Charlin, and J. Pineau · 2016
Earlier work this paper cites.
On evaluating and comparing conversational agents
A. Venkatesh, C. Khatri, A. Ram, F. Guo, R. Gabriel, A. Nagar, R. Prasad, M. Cheng, B. Hedayatnia, A. Metallinou, et al · 2017
Earlier work this paper cites.
QuAC: Question answering in context
E. Choi, H. He, M. Iyyer, M. Yatskar, W.-t. Yih, Y. Choi, P. Liang, and L. Zettlemoyer · 2018
Earlier work this paper cites.
Hotpotqa: A dataset for diverse, explainable multi-hop question answering
Z. Yang, P. Qi, S. Zhang, Y. Bengio, W. W. Cohen, R. Salakhutdinov, and C. D. Manning · 2018
Earlier work this paper cites.
Personalizing dialogue agents: I have a dog, do you have pets too?
S. Zhang, E. Dinan, J. Urbanek, A. Szlam, D. Kiela, and J. Weston · 2018
Earlier work this paper cites.
Modeling personalization in continuous space for response generation via augmented wasserstein autoencoders
Z. Chan, J. Li, X. Yang, X. Chen, W. Hu, D. Zhao, and R. Yan · 2019
Earlier work this paper cites.
Approximating interactive human evaluation with self-play for open-domain dialog systems
A. Ghandeharioun, J. H. Shen, N. Jaques, C. Ferguson, N. Jones, A. Lapedriza, and R. Picard · 2019
Earlier work this paper cites.
Acute-eval: Improved dialogue evaluation with optimized questions and multi-turn comparisons
M. Li, J. Weston, and S. Roller · 2019
Earlier work this paper cites.
Personalizing dialogue agents via meta-learning
A. Madotto, Z. Lin, C.-S. Wu, and P. Fung · 2019
Earlier work this paper cites.
What makes a good conversation? how controllable attributes affect human judgments
A. See, S. Roller, D. Kiela, and J. Weston · 2019
Earlier work this paper cites.
A pre-training based personalized dialogue generation model with persona-sparse data
Y. Zheng, R. Zhang, X. Mao, and M. Huang · 2019
Earlier work this paper cites.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Earlier work this paper cites.
Realtoxicityprompts: Evaluating neural toxic degeneration in language models
S. Gehman, S. Gururangan, M. Sap, Y. Choi, and N. A. Smith · 2020
Earlier work this paper cites.
Measuring massive multitask language understanding
D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt · 2020
Earlier work this paper cites.
AmbigQA: Answering ambiguous open-domain questions
S. Min, J. Michael, H. Hajishirzi, and L. Zettlemoyer · 2020
Cited alongside, same era.
Survey on evaluation methods for dialogue systems
J. Deriu, Á. Rodrigo, A. Otegi, G. Echegoyen, S. Rosset, E. Agirre, and M. Cieliebak · 2021
Cited alongside, same era.
Revealing persona biases in dialogue systems
E. Sheng, J. Arnold, Z. Yu, K.-W. Chang, and N. Peng · 2021
Cited alongside, same era.
Out of One, Many: Using Language Models to Simulate Human Samples
L. P. Argyle, E. C. Busby, N. Fulda, J. Gubler, C. Rytting, and D. Wingate · 2022
Cited alongside, same era.
PaLM: Scaling Language Modeling with Pathways, Oct. 2022
A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann, P. Schuh, K. Shi, S. Tsvyashchenko, J. Maynez, A. Rao, P. Barnes, Y. Tay, N. Shazeer, V. Prabhakaran, E. Reif, N. Du, B. Hutchinson, R. Pope, J. Bradbury, J. Austin, M. Isard, G. Gur-Ari, P. Yin, T. Duke, A. Levskaya, S. Ghemawat, S. Dev, H. Michalewski, X. Garcia, V. Misra, K. Robinson, L. Fedus, D. Zhou, D. Ippolito, D. Luan, H. Lim, B. Zoph, A. Spiridonov, R. Sepassi, D. Dohan, S. Agrawal, M. Omernick, A. M. Dai, T. S. Pillai, M. Pellat, A. Lewkowycz, E. Moreira, R. Child, O. Polozov, K. Lee, Z. Zhou, X. Wang, B. Saeta, M. Diaz, O. Firat, M. Catasta, J. Wei, K. Meier-Hellstern, D. Eck, J. Dean, S. Petrov, and N. Fiedel · 2022
X. Li, Y. Li, S. Joty, L. Liu, F. Huang, L. Qiu, and L. Bing · 2023
Later among the works it cites.
Y.-T. Lin and Y.-N. Chen · 2023
Later among the works it cites.
Large language model guided tree-of-thought
J. Long · 2023
Later among the works it cites.
Generative agents: Interactive simulacra of human behavior
J. S. Park, J. C. O’Brien, C. J. Cai, M. R. Morris, P. Liang, and M. S. Bernstein · 2023
Later among the works it cites.
Can ChatGPT Assess Human Personalities? A General Evaluation Framework, Mar. 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Par: Persona aware response in conversational systems
A. Nargund, S. Pandey, and J. Ham · 2022
Cited alongside, same era.
Social Simulacra: Creating Populated Prototypes for Social Computing Systems
J. S. Park, L. Popowski, C. Cai, M. R. Morris, P. Liang, and M. S. Bernstein · 2022
Cited alongside, same era.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
B. bench authors · 2023
Cited alongside, same era.
A survey on evaluation of large language models, 2023
Y. Chang, X. Wang, J. Wang, Y. Wu, L. Yang, K. Zhu, H. Chen, X. Yi, C. Wang, Y. Wang, W. Ye, Y. Zhang, Y. Chang, P. S. Yu, Q. Yang, and X. Xie · 2023
Cited alongside, same era.
Toxicity in chatgpt: Analyzing persona-assigned language models
A. Deshpande, V. Murahari, T. Rajpurohit, A. Kalyan, and K. Narasimhan · 2023
Cited alongside, same era.
Improving language model negotiation with self-play and in-context learning from ai feedback, 2023
Y. Fu, H. Peng, T. Khot, and M. Lapata · 2023
Cited alongside, same era.
S$^3$: Social-network Simulation System with Large Language Model-Empowered Agents, July 2023
C. Gao, X. Lan, Z. Lu, J. Mao, J. Piao, H. Wang, D. Jin, and Y. Li · 2023
Cited alongside, same era.
H. Rao, C. Leung, and C. Miao · 2023
Later among the works it cites.
LLaMA: Open and Efficient Foundation Language Models, Feb. 2023
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, A. Rodriguez, A. Joulin, E. Grave, and G. Lample · 2023
Later among the works it cites.
The Rise and Potential of Large Language Model Based Agents: A Survey, Sept. 2023
Z. Xi, W. Chen, X. Guo, W. He, Y. Ding, B. Hong, M. Zhang, J. Wang, S. Jin, E. Zhou, R. Zheng, X. Fan, X. Wang, L. Xiong, Q. Liu, Y. Zhou, W. Wang, C. Jiang, Y. Zou, X. Liu, Z. Yin, S. Dou, R. Weng, W. Cheng, Q. Zhang, W. Qin, Y. Zheng, X. Qiu, X. Huan, and T. Gui · 2023
Later among the works it cites.
Y. Xiao, Y. Cheng, J. Fu, J. Wang, W. Li, and P. Liu · 2023
Later among the works it cites.
Intercode: Standardizing and benchmarking interactive coding with execution feedback
J. Yang, A. Prabhakar, K. Narasimhan, and S. Yao · 2023
Later among the works it cites.
A survey of large language models
W. X. Zhao, K. Zhou, J. Li, T. Tang, X. Wang, Y. Hou, Y. Min, B. Zhang, J. Zhang, Z. Dong, et al · 2023
Later among the works it cites.
Judging llm-as-a-judge with mt-bench and chatbot arena
L. Zheng, W.-L. Chiang, Y. Sheng, S. Zhuang, Z. Wu, Y. Zhuang, Z. Lin, Z. Li, D. Li, E. Xing, et al · 2023
Later among the works it cites.
Interactive question answering systems: Literature review
G. M. Biancofiore, Y. Deldjoo, T. D. Noia, E. Di Sciascio, and F. Narducci · 2024
Closest in time.
Thinking fair and slow: On the efficacy of structured prompts for debiasing language models
S. Furniturewala, S. Jandial, A. Java, P. Banerjee, S. Shahid, S. Bhatia, and K. Jaidka · 2024
Closest in time.
Dsbench: How far are data science agents to becoming data science experts?, 2024
L. Jing, Z. Huang, X. Wang, W. Yao, W. Yu, K. Ma, H. Zhang, X. Du, and D. Yu · 2024
Closest in time.
PRD: Peer rank and discussion improve large language model based evaluations
R. Li, T. Patel, and X. Du · 2024
Closest in time.
Llm evaluators recognize and favor their own generations, 2024
A. Panickssery, S. R. Bowman, and S. Feng · 2024
Closest in time.
A survey on large language model based autonomous agents
L. Wang, C. Ma, X. Feng, Z. Zhang, H. Yang, J. Zhang, Z. Chen, J. Tang, X. Chen, Y. Lin, et al · 2024
Closest in time.
Logical reasoning over natural language as knowledge representation: A survey, 2024
Z. Yang, X. Du, R. Mao, J. Ni, and E. Cambria · 2024
Closest in time.
Towards persona-based empathetic conversational models
P. Zhong, C. Zhang, H. Wang, Y. Liu, and C. Miao · 2024
Closest in time.