Fetching the paper…
Reading the bibliography…
With the rapid adoption of LLM-based chatbots, there is a pressing need to evaluate what humans and LLMs can achieve together.
Demographics of mechanical turk
Panos Ipeirotis. 2010 · 2010
Earlier work this paper cites.
HotpotQA: A dataset for diverse, explainable multi-hop question answering
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W. Cohen, Ruslan Salakhutdinov, and Christopher D. Manning. 2018 · 2018
Earlier work this paper cites.
Does the whole exceed its parts? the effect of ai explanations on complementary team performance
Gagan Bansal, Tongshuang Wu, Joyce Zhou, Raymond Fok, Besmira Nushi, Ece Kamar, Marco Tulio Ribeiro, and Daniel Weld. 2021 · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. 2021 · 2021
Earlier work this paper cites.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021 · 2021
Earlier work this paper cites.
Characterizing the global crowd workforce:a cross-country comparison of crowdworker demographics
Lisa Posch, Arnim Bleier, Fabian Flöck, Clemens M. Lechner, Katharina Kinder-Kurlanda, Denis Helic, and Markus Strohmaier. 2022 · 2022
Earlier work this paper cites.
Out of one, many: Using language models to simulate human samples
Lisa P. Argyle, Ethan C. Busby, Nancy Fulda, Joshua R. Gubler, Christopher Rytting, and David Wingate. 2023 · 2023
Earlier work this paper cites.
Compost: Characterizing and evaluating caricature in LLM simulations
Myra Cheng, Tiziano Piccardi, and Diyi Yang. 2023b · 2023
Earlier work this paper cites.
Alpacafarm: A simulation framework for methods that learn from human feedback
Yann Dubois, Xuechen Li, Rohan Taori, Tianyi Zhang, Ishaan Gulrajani, Jimmy Ba, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023 · 2023
Earlier work this paper cites.
Gemini: A family of highly capable multimodal models
Gemini Team Google. 2023 · 2023
Earlier work this paper cites.
Large language models as simulated economic agents: What can we learn from homo silicus?
John J. Horton. 2023 · 2023
Earlier work this paper cites.
Aligning language models to user opinions
EunJeong Hwang, Bodhisattwa Majumder, and Niket Tandon. 2023 · 2023
Earlier work this paper cites.
Evaluating human-language model interaction
Mina Lee, Megha Srivastava, Amelia Hardy, John Thickstun, Esin Durmus, Ashwin Paranjape, Ines Gerard-Ursin, Xiang Lisa Li, Faisal Ladhak, Frieda Rong, Rose E. Wang, Minae Kwon, Joon Sung Park, Hancheng Cao, Tony Lee, Rishi Bommasani, Michael Bernstein, and Percy Liang. 2023 · 2023
Earlier work this paper cites.
Holistic evaluation of language models
Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, et al. 2023 · 2023
Earlier work this paper cites.
OpenAI. 2023 · 2023
Cited alongside, same era.
Generative agents: Interactive simulacra of human behavior
Joon Sung Park, Joseph C. O’Brien, Carrie J. Cai, Meredith Ringel Morris, Percy Liang, and Michael S. Bernstein. 2023 · 2023
Cited alongside, same era.
Judging LLM-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2023 · 2023
Cited alongside, same era.
The rapid adoption of generative AI
Alexander Bick, Adam Blandin, and David Deming. 2024 · 2024
Cited alongside, same era.
Synthetic replacements for human survey data? the perils of large language models
James Bisbee, Joshua D. Clinton, Cassy Dorff, Brenton Kenkel, and Jennifer M. Larson. 2024 · 2024
Cited alongside, same era.
From live data to high-quality benchmarks: The arena-hard pipeline
Tianle Li, Wei-Lin Chiang, Evan Frick, Lisa Dunlap, Banghua Zhu, Joseph E. Gonzalez, and Ion Stoica. 2024c · 2024
Later among the works it cites.
WildBench: Benchmarking LLMs with challenging tasks from real users in the wild
Bill Yuchen Lin, Yuntian Deng, Khyathi Chandu, Faeze Brahman, Abhilasha Ravichander, Valentina Pyatkin, Nouha Dziri, Ronan Le Bras, and Yejin Choi. 2024 · 2024
Later among the works it cites.
Llama Team, AI@Meta. 2024 · 2024
Later among the works it cites.
Adding error bars to evals: A statistical approach to language model evaluations
Evan Miller. 2024 · 2024
Later among the works it cites.
BASES: Large-scale web search user simulation with large language model based agents
Ruiyang Ren, Peng Qiu, Yingqi Qu, Jing Liu, Xin Zhao, Hua Wu, Ji-Rong Wen, and Haifeng Wang. 2024 · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Serina Chang, Alicja Chaszczewicz, Emma Wang, Maya Josifovska, Emma Pierson, and Jure Leskovec. 2024 · 2024
Cited alongside, same era.
Chatbot Arena: An open platform for evaluating LLMs by human preference
Wei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos, Tianle Li, Dacheng Li, Hao Zhang, Banghua Zhu, Michael Jordan, Joseph E. Gonzalez, and Ion Stoica. 2024 · 2024
Cited alongside, same era.
Evaluating language models for mathematics through interactions
Katherine M. Collins, Albert Q. Jiang, Simon Frieder, Lionel Wong, Miri Zilka, Umang Bhatt, Thomas Lukasiewicz, Yuhuai Wu, Joshua B. Tenenbaum, William Hart, Timothy Gowers, Wenda Li, Adrian Weller, and Mateja Jamnik. 2024 · 2024
Cited alongside, same era.
Shaping human-ai collaboration: Varied scaffolding levels in co-writing with language models
Paramveer S. Dhillon, Somayeh Molaei, Jiaqi Li, Maximilian Golub, Shaochun Zheng, and Lionel Peter Robert. 2024 · 2024
Cited alongside, same era.
Aryo Pradipta Gema, Joshua Ong Jun Leang, Giwon Hong, Alessio Devoto, Alberto Carlo Maria Mancino, Rohit Saxena, Xuanli He, Yu Zhao, Xiaotang Du, Mohammad Reza Ghasemi Madani, Claire Barale, Robert McHardy, Joshua Harris, Jean Kaddour, Emile van Krieken, and Pasquale Minervini. 2024 · 2024
Cited alongside, same era.
More than marketing? on the information value of ai benchmarks for practitioners
Amelia Hardy, Anka Reuel, Kiana Jafari Meimandi, Lisa Soder, Allie Griffith, Dylan M. Asmar, Sanmi Koyejo, Michael S. Bernstein, and Mykel J. Kochenderfer. 2024 · 2024
Cited alongside, same era.
Predicting results of social science experiments using large language models
Luke Hewitt, Ashwini Ashokkumar, Isaias Ghezae1, and Robb Willer. 2024 · 2024
Cited alongside, same era.
Later among the works it cites.
Collaborative gym: A framework for enabling and evaluating human-agent collaboration
Yijia Shao, Vinay Samuel, Yucheng Jiang, John Yang, and Diyi Yang. 2024 · 2024
Later among the works it cites.
When combinations of humans and AI are useful: A systematic review and meta-analysis
Michelle Vaccaro, Abdullah Almaatouq, and Thomas Malone. 2024 · 2024
Later among the works it cites.
WildChat: 1m chatgpt interaction logs in the wild
Wenting Zhao, Xiang Ren, Jack Hessel, Claire Cardie, Yejin Choi, and Yuntian Deng. 2024 · 2024
Later among the works it cites.
LMSYS-Chat-1M: A large-scale real-world LLM conversation dataset
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Tianle Li, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zhuohan Li, Zi Lin, Eric P. Xing, Joseph E. Gonzalez, Ion Stoica, and Hao Zhang. 2024 · 2024
Later among the works it cites.
Language model fine-tuning on scaled survey data for predicting distributions of public opinions
Joseph Suh, Erfan Jahanparast, Suhong Moon, Minwoo Kang, and Serina Chang. 2025 · 2025
Closest in time.
Labor Force Statistics from the Current Population Survey
U.S. Bureau of Labor Statistics. 2025 · 2025
Closest in time.
National Population by Characteristics: 2020-2024
U.S. Census Bureau. 2024 · 2025
Closest in time.
Large language models that replace human participants can harmfully misportray and flatten identity groups
Angelina Wang, Jamie Morgenstern, and John P. Dickerson. 2025 · 2025
Closest in time.