Fetching the paper…
Reading the bibliography…
Large-scale surveys are essential tools for informing social science research and policy, but running surveys is costly and time-intensive.
Stop measuring calibration when humans disagree
Joris Baan, Wilker Aziz, Barbara Plank, and Raquel Fernandez. 2022 · 1915
Earlier work this paper cites.
A metric for distributions with applications to image databases
Yossi Rubner, Carlo Tomasi, and Leonidas J Guibas. 1998 · 1998
Earlier work this paper cites.
World values survey wave 7 (2017-2022) cross-national data-set
C. Haerpfer, R. Inglehart, A. Moreno, C. Welzel, K. Kizilova, J. Diez-Medrano, M. Lagos, P. Norris, E. Ponarin, and B. Puranen. 2022 · 2022
Earlier work this paper cites.
LoRA: Low-rank adaptation of large language models
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022 · 2022
Earlier work this paper cites.
Using large language models to simulate multiple humans and replicate human subject studies
Gati V Aher, Rosa I Arriaga, and Adam Tauman Kalai. 2023 · 2023
Earlier work this paper cites.
Out of one, many: Using language models to simulate human samples
Lisa P Argyle, Ethan C Busby, Nancy Fulda, Joshua R Gubler, Christopher Rytting, and David Wingate. 2023 · 2023
Earlier work this paper cites.
Probing pre-trained language models for cross-cultural differences in values
Arnav Arora, Lucie-aimée Kaffee, and Isabelle Augenstein. 2023 · 2023
Earlier work this paper cites.
Assessing LLMs for moral value pluralism
Noam Benkler, Drisana Mosaphir, Scott Friedman, Andrew Smart, and Sonja Schmer-Galunder. 2023 · 2023
Earlier work this paper cites.
Synthetic replacements for human survey data? the perils of large language models
James Bisbee, Joshua D Clinton, Cassy Dorff, Brenton Kenkel, and Jennifer M Larson. 2023 · 2023
Earlier work this paper cites.
Assessing cross-cultural alignment between ChatGPT and human societies: An empirical study
Yong Cao, Li Zhou, Seolhwa Lee, Laura Cabello, Min Chen, and Daniel Hershcovich. 2023 · 2023
Earlier work this paper cites.
Vicuna: An open-source chatbot impressing GPT-4 with 90%* ChatGPT quality
Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. 2023 · 2023
Earlier work this paper cites.
Questioning the survey responses of large language models
Ricardo Dominguez-Olmedo, Moritz Hardt, and Celestine Mendler-Dünner. 2023 · 2023
Cited alongside, same era.
Towards measuring the representation of subjective global opinions in language models
Esin Durmus, Karina Nyugen, Thomas I Liao, Nicholas Schiefer, Amanda Askell, Anton Bakhtin, Carol Chen, Zac Hatfield-Dodds, Danny Hernandez, Nicholas Joseph, et al. 2023 · 2023
Cited alongside, same era.
Large language models as simulated economic agents: What can we learn from homo silicus?
John J Horton. 2023 · 2023
Cited alongside, same era.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023 · 2023
Cited alongside, same era.
Llama 3 model card
Simulating subjects: The promise and peril of ai stand-ins for social agents and interactions
Austin Kozlowski and James Evans. 2024 · 2024
Later among the works it cites.
Evaluating cultural adaptability of a large language model via simulation of synthetic personas
Louis Kwok, Michal Bravansky, and Lewis D Griffin. 2024 · 2024
Later among the works it cites.
Can multiple-choice questions really be useful in detecting the abilities of LLMs?
Wangyue Li, Liangzhi Li, Tong Xiang, Xiao Liu, Wei Deng, and Noa Garcia. 2024 · 2024
Later among the works it cites.
Automated social science: Language models as scientist and subjects
Benjamin S Manning, Kehang Zhu, and John J Horton. 2024 · 2024
Later among the works it cites.
Diminished diversity-of-thought in a standard large language model
Peter S Park, Philipp Schoenegger, and Chongyang Zhu. 2024 · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
AI@Meta. 2024 · 2024
Cited alongside, same era.
Investigating cultural alignment of large language models
Badr AlKhamissi, Muhammad ElNokrashy, Mai Alkhamissi, and Mona Diab. 2024 · 2024
Cited alongside, same era.
Interpreting predictive probabilities: Model confidence or human label variation?
Joris Baan, Raquel Fernández, Barbara Plank, and Wilker Aziz. 2024 · 2024
Cited alongside, same era.
Can generative ai improve social science?
Christopher A Bail. 2024 · 2024
Cited alongside, same era.
Investigating uncertainty calibration of aligned language models under the multiple-choice setting
Guande He, Peng Cui, Jianfei Chen, Wenbo Hu, and Jun Zhu. 2024 · 2024
Cited alongside, same era.
Predicting results of social science experiments using large language models
Luke Hewitt, Ashwini Ashokkumar, Isaias Ghezae, and Robb Willer. 2024 · 2024
Cited alongside, same era.
Seungjong Sun, Eungu Lee, Dongyan Nan, Xiangying Zhao, Wonbyung Lee, Bernard J. Jansen, and Jang Hyun Kim. 2024 · 2024
Later among the works it cites.
Revealing fine-grained values and opinions in large language models
Dustin Wright, Arnav Arora, Nadav Borenstein, Srishti Yadav, Serge Belongie, and Isabelle Augenstein. 2024 · 2024
Later among the works it cites.
On the calibration of multilingual question answering llms
Yahan Yang, Soham Dan, Dan Roth, and Insup Lee. 2024 · 2024
Later among the works it cites.
Worldvaluesbench: A large-scale benchmark dataset for multi-cultural value awareness of language models
Wenlong Zhao, Debanjan Mondal, Niket Tandon, Danica Dillion, Kurt Gray, and Yuling Gu. 2024 · 2024
Later among the works it cites.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et al. 2025 · 2025
Closest in time.