Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have increasingly enhanced or replaced traditional Non-Player Characters (NPCs) in video games.
Game design for classical AI
Horswill, I. 2014 · 2014
Earlier work this paper cites.
On Measuring Social Biases in Sentence Encoders
May, C.; Wang, A.; Bordia, S.; Bowman, S.; and Rudinger, R. 2019 · 2019
Earlier work this paper cites.
Behavior tree design of intelligent behavior of non-player character (NPC) based on Unity3D
Zhu, X. 2019 · 2019
Earlier work this paper cites.
Auto-debias: Debiasing masked language models with automated biased prompts
Guo, Y.; Yang, Y.; and Abbasi, A. 2022 · 2022
Earlier work this paper cites.
Education in the era of generative artificial intelligence (AI): Understanding the potential benefits of ChatGPT in promoting teaching and learning
Baidoo-Anu, D.; and Ansah, L. O. 2023 · 2023
Earlier work this paper cites.
Gamegpt: Multi-agent collaborative framework for game development
Chen, D.; Wang, H.; Huo, Y.; Li, Y.; and Zhang, H. 2023 · 2023
Earlier work this paper cites.
Conversational Interactions with NPCs in LLM-Driven Gaming: Guidelines from a Content Analysis of Player Feedback
Cox, S. R.; and Ooi, W. T. 2023 · 2023
Earlier work this paper cites.
Fft: Towards harmlessness evaluation and analysis for llms with factuality, fairness, toxicity
Cui, S.; Zhang, Z.; Chen, Y.; Zhang, W.; Liu, T.; Wang, S.; and Liu, T. 2023 · 2023
Earlier work this paper cites.
WinoQueer: A Community-in-the-Loop Benchmark for Anti-LGBTQ+ Bias in Large Language Models
Felkner, V.; Chang, H.-C. H.; Jang, E.; and May, J. 2023 · 2023
Earlier work this paper cites.
LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models
Guha, N.; Nyarko, J.; Ho, D. E.; Re, C.; Chilton, A.; Narayana, A.; Chohlas-Wood, A.; Peters, A.; Waldon, B.; Rockmore, D.; et al. 2023 · 2023
Earlier work this paper cites.
Is ChatGPT a good translator? A preliminary study
Jiao, W.; Wang, W.; Huang, J.-t.; Wang, X.; and Tu, Z. 2023 · 2023
Earlier work this paper cites.
Assessing the accuracy and reliability of AI-generated medical responses: an evaluation of the Chat-GPT model
Johnson, D.; Goodman, R.; Patrinely, J.; Stone, C.; Zimmerman, E.; Donald, R.; Chang, S.; Berkowitz, S.; Finn, A.; Jahangir, E.; et al. 2023 · 2023
Earlier work this paper cites.
Split and merge: Aligning position biases in large language model based evaluators
Li, Z.; Wang, C.; Ma, P.; Wu, D.; Wang, S.; Gao, C.; and Liu, Y. 2023 · 2023
Earlier work this paper cites.
OpenAI. 2023 · 2023
Earlier work this paper cites.
Gameeval: Evaluating llms on conversational games
Qiao, D.; Wu, C.; Liang, Y.; Li, J.; and Duan, N. 2023 · 2023
Earlier work this paper cites.
Is ChatGPT a General-Purpose Natural Language Processing Task Solver?
Qin, C.; Zhang, A.; Zhang, Z.; Chen, J.; Yasunaga, M.; and Yang, D. 2023 · 2023
Earlier work this paper cites.
Voyager: An open-ended embodied agent with large language models
Wang, G.; Xie, Y.; Jiang, Y.; Mandlekar, A.; Xiao, C.; Zhu, Y.; Fan, L.; and Anandkumar, A. 2023 · 2023
Cited alongside, same era.
Is chatgpt fair for recommendation? evaluating fairness in large language model recommendation
Zhang, J.; Bao, K.; Zhang, Y.; Wang, W.; Feng, F.; and He, X. 2023 · 2023
Cited alongside, same era.
Judging llm-as-a-judge with mt-bench and chatbot arena
Zheng, L.; Chiang, W.-L.; Sheng, Y.; Zhuang, S.; Wu, Z.; Zhuang, Y.; Lin, Z.; Li, Z.; Li, D.; Xing, E.; et al. 2023 · 2023
Cited alongside, same era.
Causal-debias: Unifying debiasing in pretrained language models and fine-tuning via causal invariant learning
Zhou, F.; Mao, Y.; Yu, L.; Yang, Y.; and Zhong, T. 2023 · 2023
Cited alongside, same era.
Cooperation, competition, and maliciousness: LLM-stakeholders interactive negotiation
Abdelnabi, S.; Gomaa, A.; Sivaprasad, S.; Schönherr, L.; and Fritz, M. 2024 · 2024
Cited alongside, same era.
Liu, A.; Feng, B.; Xue, B.; Wang, B.; Wu, B.; Lu, C.; Zhao, C.; Deng, C.; Zhang, C.; Ruan, C.; et al. 2024 · 2024
Later among the works it cites.
LLM Comparative Assessment: Zero-shot NLG Evaluation through Pairwise Comparisons using Large Language Models
Liusie, A.; Manakul, P.; and Gales, M. 2024 · 2024
Later among the works it cites.
Luo, H.; Huang, H.; Deng, Z.; Liu, X.; Chen, R.; and Liu, Z. 2024 · 2024
Later among the works it cites.
The Effect of LLM-Based NPC Emotional States on Player Emotions: An Analysis of Interactive Game Play
Marincioni, A.; Miltiadous, M.; Zacharia, K.; Heemskerk, R.; Doukeris, G.; Preuss, M.; and Barbero, G. 2024 · 2024
Later among the works it cites.
Having Beer after Prayer? Measuring Cultural Bias in Large Language Models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Game Plot Design with an LLM-powered Assistant: An Empirical Study with Game Designers
Alavi, S. H.; Xu, W.; Jojic, N.; Kennett, D.; Ng, R. T.; Rao, S.; Zhang, H.; Dolan, B.; and Shwartz, V. 2024 · 2024
Cited alongside, same era.
Towards Implicit Bias Detection and Mitigation in Multi-Agent LLM Interactions
Borah, A.; and Mihalcea, R. 2024 · 2024
Cited alongside, same era.
Revisiting Zero-Shot Abstractive Summarization in the Era of Large Language Models from the Perspective of Position Bias
Chhabra, A.; Askari, H.; and Mohapatra, P. 2024 · 2024
Cited alongside, same era.
Bias and unfairness in information retrieval systems: New challenges in the llm era
Dai, S.; Xu, C.; Xu, S.; Pang, L.; Dong, Z.; and Xu, J. 2024 · 2024
Cited alongside, same era.
Gtbench: Uncovering the strategic reasoning limitations of llms via game-theoretic evaluations
Duan, J.; Zhang, R.; Diffenderfer, J.; Kailkhura, B.; Sun, L.; Stengel-Eskin, E.; Bansal, M.; Chen, T.; and Xu, K. 2024 · 2024
Cited alongside, same era.
Dubey, A.; Jauhri, A.; Pandey, A.; Kadian, A.; Al-Dahle, A.; Letman, A.; Mathur, A.; Schelten, A.; Yang, A.; Fan, A.; et al. 2024 · 2024
Cited alongside, same era.
Can large language models serve as rational players in game theory? a systematic analysis
Fan, C.; Chen, J.; Jin, Y.; and He, H. 2024 · 2024
Cited alongside, same era.
Naous, T.; Ryan, M. J.; Ritter, A.; and Xu, W. 2024 · 2024
Later among the works it cites.
LLM-based methods for the creation of unit tests in game development
Paduraru, C.; Staicu, A.; and Stefanescu, A. 2024 · 2024
Later among the works it cites.
Player-driven emergence in llm-driven game narrative
Peng, X.; Quaye, J.; Rao, S.; Xu, W.; Botchway, P.; Brockett, C.; Jojic, N.; DesGarennes, G.; Lobb, K.; Xu, M.; et al. 2024 · 2024
Later among the works it cites.
Llm economicus? mapping the behavioral biases of llms via utility theory
Ross, J.; Kim, Y.; and Lo, A. W. 2024 · 2024
Later among the works it cites.
General phrase debiaser: Debiasing masked language models at a multi-token level
Shi, B.; Zhang, X.; Kong, D.; Wu, Y.; Liu, Z.; Lyu, H.; and Huang, L. 2024a · 2024
Later among the works it cites.
Large language models are inconsistent and biased evaluators
Stureborg, R.; Alikaniotis, D.; and Suhara, Y. 2024 · 2024
Later among the works it cites.
Systematic Biases in LLM Simulations of Debates
Taubenfeld, A.; Dover, Y.; Reichart, R.; and Goldstein, A. 2024 · 2024
Later among the works it cites.
Fake Alignment: Are LLMs Really Aligned Well?
Wang, Y.; Teng, Y.; Huang, K.; Lyu, C.; Zhang, S.; Zhang, W.; Ma, X.; Jiang, Y.-G.; Qiao, Y.; and Wang, Y. 2024b · 2024
Later among the works it cites.
MAgIC: Investigation of Large Language Model Powered Multi-Agent in Cognition, Adaptability, Rationality and Collaboration
Xu, L.; Hu, Z.; Zhou, D.; Ren, H.; Dong, Z.; Keutzer, K.; Ng, S. K.; and Feng, J. 2024a · 2024
Later among the works it cites.
Yang, A.; Yang, B.; Hui, B.; Zheng, B.; Yu, B.; Zhou, C.; Li, C.; Li, C.; Liu, D.; Huang, F.; et al. 2024 · 2024
Later among the works it cites.
Grok 3 beta — the age of reasoning agents
xAI. 2025 · 2025
Closest in time.