Fetching the paper…
Reading the bibliography…
Language models (LMs) are increasingly used as simulacra for people, yet their ability to match the distribution of views of a specific demographic group and be \textit{distributionally aligned} remains uncertain.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 1901
Earlier work this paper cites.
Markov chains and mixing times
David A. Levin, Yuval Peres, and Elizabeth L. Wilmer. 2006 · 2006
Earlier work this paper cites.
The neglected 95%: why american psychology needs to become less american
Jeffrey Jensen Arnett. 2008 · 2008
Earlier work this paper cites.
(Mis)perceptions of Partisan Polarization in the American Public
Matthew S. Levendusky and Neil Malhotra. 2015 · 2015
Earlier work this paper cites.
Measuring public opinion with surveys
Adam Berinsky. 2017 · 2017
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul F. Christiano, Jan Leike, Tom B. Brown, Miljan Martic, Shane Legg, and Dario Amodei. 2017 · 2017
Earlier work this paper cites.
On calibration of modern neural networks
Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q. Weinberger. 2017 · 2017
Earlier work this paper cites.
The perception gap: How false impressions are pulling americans apart
Daniel A Yudkin, Stephen Hawkins, and Tim Dixon. 2019 · 2019
Earlier work this paper cites.
Beyond weird: A review of the last decade and a look ahead to the global laboratory of the future
Coren Apicella, Ara Norenzayan, and Joseph Henrich. 2020 · 2020
Earlier work this paper cites.
Designing disaggregated evaluations of ai systems: Choices, considerations, and tradeoffs
Solon Barocas, Anhong Guo, Ece Kamar, Jacquelyn Krones, Meredith Ringel Morris, Jennifer Wortman Vaughan, W. Duncan Wadsworth, and Hanna Wallach. 2021 · 2021
Earlier work this paper cites.
Simcse: Simple contrastive learning of sentence embeddings
Tianyu Gao, Xingcheng Yao, and Danqi Chen. 2021 · 2021
Earlier work this paper cites.
Consequences of asking sensitive questions in surveys
Ting Yan. 2021 · 2021
Earlier work this paper cites.
CommunityLM: Probing partisan worldviews from language models
Hang Jiang, Doug Beeferman, Brandon Roy, and Deb Roy. 2022 · 2022
Earlier work this paper cites.
Language models (mostly) know what they know
Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield-Dodds, Nova DasSarma, Eli Tran-Johnson, Scott Johnston, Sheer El-Showk, Andy Jones, Nelson Elhage, Tristan Hume, Anna Chen, Yuntao Bai, Sam Bowman, Stanislav Fort, Deep Ganguli, Danny Hernandez, Josh Jacobson, Jackson Kernion, Shauna Kravec, Liane Lovitt, Kamal Ndousse, Catherine Olsson, Sam Ringer, Dario Amodei, Tom Brown, Jack Clark, Nicholas Joseph, Ben Mann, Sam McCandlish, Chris Olah, and Jared Kaplan. 2022 · 2022
Earlier work this paper cites.
Surfacing racial stereotypes through identity portrayal
Gauri Kambhatla, Ian Stewart, and Rada Mihalcea. 2022 · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F Christiano, Jan Leike, and Ryan Lowe. 2022 · 2022
Earlier work this paper cites.
Using large language models to simulate multiple humans and replicate human subject studies
Gati V Aher, Rosa I. Arriaga, and Adam Tauman Kalai. 2023 · 2023
Earlier work this paper cites.
Out of one, many: Using language models to simulate human samples
Lisa P. Argyle, Ethan C. Busby, Nancy Fulda, Joshua R. Gubler, Christopher Rytting, and David Wingate. 2023 · 2023
Earlier work this paper cites.
CoMPosT: Characterizing and evaluating caricature in LLM simulations
Myra Cheng, Tiziano Piccardi, and Diyi Yang. 2023b · 2023
Earlier work this paper cites.
Can ai language models replace human participants?
Danica Dillion, Niket Tandon, Yuling Gu, and Kurt Gray. 2023 · 2023
Earlier work this paper cites.
Alpacafarm: A simulation framework for methods that learn from human feedback
Yann Dubois, Chen Xuechen Li, Rohan Taori, Tianyi Zhang, Ishaan Gulrajani, Jimmy Ba, Carlos Guestrin, Percy S Liang, and Tatsunori B Hashimoto. 2023 · 2023
Earlier work this paper cites.
Ai and the transformation of social science research
Igor Grossmann, Matthew Feinberg, Dawn C. Parker, Nicholas A. Christakis, Philip E. Tetlock, and William A. Cunningham. 2023 · 2023
Earlier work this paper cites.
Large language models as simulated economic agents: What can we learn from homo silicus?
John J Horton. 2023 · 2023
Earlier work this paper cites.
Prompting is not a substitute for probability measurements in large language models
Jennifer Hu and Roger P. Levy. 2023 · 2023
Earlier work this paper cites.
Aligning language models to user opinions
EunJeong Hwang, Bodhisattwa Majumder, and Niket Tandon. 2023 · 2023
Earlier work this paper cites.
Large language models as superpositions of cultural perspectives
Grgur Kovač, Masataka Sawayama, Rémy Portelas, Cédric Colas, Peter Ford Dominey, and Pierre-Yves Oudeyer. 2023 · 2023
Cited alongside, same era.
Cognitive dissonance: Why do language model outputs disagree with internal representations of truthfulness?
Kevin Liu, Stephen Casper, Dylan Hadfield-Menell, and Jacob Andreas. 2023 · 2023
Cited alongside, same era.
Discovering language model behaviors with model-written evaluations
Ethan Perez, Sam Ringer, Kamile Lukosiute, Karina Nguyen, Edwin Chen, Scott Heiner, Craig Pettit, Catherine Olsson, Sandipan Kundu, Saurav Kadavath, Andy Jones, Anna Chen, Benjamin Mann, Brian Israel, Bryan Seethor, Cameron McKinnon, Christopher Olah, Da Yan, Daniela Amodei, Dario Amodei, Dawn Drain, Dustin Li, Eli Tran-Johnson, Guro Khundadze, Jackson Kernion, James Landis, Jamie Kerr, Jared Mueller, Jeeyoon Hyun, Joshua Landau, Kamal Ndousse, Landon Goldberg, Liane Lovitt, Martin Lucas, Michael Sellitto, Miranda Zhang, Neerav Kingsland, Nelson Elhage, Nicholas Joseph, Noemi Mercado, Nova DasSarma, Oliver Rausch, Robin Larson, Sam McCandlish, Scott Johnston, Shauna Kravec, Sheer El Showk, Tamera Lanham, Timothy Telleen-Lawton, Tom Brown, Tom Henighan, Tristan Hume, Yuntao Bai, Zac Hatfield-Dodds, Jack Clark, Samuel R. Bowman, Amanda Askell, Roger Grosse, Danny Hernandez, Deep Ganguli, Evan Hubinger, Nicholas Schiefer, and Jared Kaplan. 2023 · 2023
Cited alongside, same era.
Whose opinions do language models reflect?
How random is random? evaluating the randomness and humaness of llms’ coin flips
Katherine Van Koevering and Jon Kleinberg. 2024 · 2024
Closest in time.
Benchmarking cognitive biases in large language models as evaluators
Ryan Koo, Minhwa Lee, Vipul Raheja, Jong Inn Park, Zae Myung Kim, and Dongyeop Kang. 2024 · 2024
Closest in time.
The steerability of large language models toward data-driven personas
Junyi Li, Charith Peris, Ninareh Mehrabi, Palash Goyal, Kai-Wei Chang, Aram Galstyan, Richard Zemel, and Rahul Gupta. 2024a · 2024
Closest in time.
Evaluating large language model biases in persona-steered generation
Andy Liu, Mona Diab, and Daniel Fried. 2024 · 2024
Closest in time.
Beyond probabilities: Unveiling the misalignment in evaluating large language models
Chenyang Lyu, Minghao Wu, and Alham Aji. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee, Percy Liang, and Tatsunori Hashimoto. 2023 · 2023
Cited alongside, same era.
Moral mimicry: Large language models produce moral rationalizations tailored to political identity
Gabriel Simmons. 2023 · 2023
Cited alongside, same era.
Just ask for calibration: Strategies for eliciting calibrated confidence scores from language models fine-tuned with human feedback
Katherine Tian, Eric Mitchell, Allan Zhou, Archit Sharma, Rafael Rafailov, Huaxiu Yao, Chelsea Finn, and Christopher Manning. 2023 · 2023
Cited alongside, same era.
Activation addition: Steering language models without optimization
Alexander Matt Turner, Lisa Thiergart, David Udell, Gavin Leech, Ulisse Mini, and Monte MacDiarmid. 2023 · 2023
Cited alongside, same era.
Are personalized stochastic parrots more dangerous? evaluating persona biases in dialogue systems
Yixin Wan, Jieyu Zhao, Aman Chadha, Nanyun Peng, and Kai-Wei Chang. 2023 · 2023
Cited alongside, same era.
Self-instruct: Aligning language models with self-generated instructions
Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, Daniel Khashabi, and Hannaneh Hajishirzi. 2023 · 2023
Cited alongside, same era.
Group preference optimization: Few-shot alignment of large language models
Siyan Zhao, John Dang, and Aditya Grover. 2023 · 2023
Cited alongside, same era.
Perils and opportunities in using large language models in psychological research
Suhaib Abdurahman, Mohammad Atari, Farzan Karimi-Malekabadi, Mona J Xue, Jackson Trager, Peter S Park, Preni Golazizian, Ali Omrani, and Morteza Dehghani. 2024 · 2024
Cited alongside, same era.
Is cognition and action consistent or not: Investigating large language model’s personality
Yiming Ai, Zhiwei He, Ziyin Zhang, Wenhong Zhu, Hongkun Hao, Kai Yu, Lingjun Chen, and Rui Wang. 2024 · 2024
Cited alongside, same era.
Reem I. Masoud, Ziquan Liu, Martin Ferianc, Philip Treleaven, and Miguel Rodrigues. 2024 · 2024
Closest in time.
Do ais know what the most important issue is? using language models to code open-text social survey responses at scale
Jonathan Mellon, Jack Bailey, Ralph Scott, James Breckwoldt, Marta Miori, and Phillip Schmedeman. 2024 · 2024
Closest in time.
Manuel Mondal, Ljiljana Dolamic, Gérôme Bovet, Philippe Cudré-Mauroux, and Julien Audiffren. 2024 · 2024
Closest in time.
Using llms to model the beliefs and preferences of targeted populations
Keiichi Namikoshi, Alex Filipowicz, David A. Shamma, Rumen Iliev, Candice L. Hogan, and Nikos Arechiga. 2024 · 2024
Closest in time.
Having beer after prayer? measuring cultural bias in large language models
Tarek Naous, Michael Ryan, Alan Ritter, and Wei Xu. 2024 · 2024
Closest in time.
What are the odds? language models are capable of probabilistic reasoning
Akshay Paruchuri, Jake Garrison, Shun Liao, John Hernandez, Jacob Sunshine, Tim Althoff, Xin Liu, and Daniel McDuff. 2024 · 2024
Closest in time.
Civics: Building a dataset for examining culturally-informed values in large language models
Giada Pistilli, Alina Leidinger, Yacine Jernite, Atoosa Kasirzadeh, Alexandra Sasha Luccioni, and Margaret Mitchell. 2024 · 2024
Closest in time.
Llm processes: Numerical predictive distributions conditioned on natural language
James Requeima, John Bronskill, Dami Choi, Richard E. Turner, and David Duvenaud. 2024 · 2024
Closest in time.
Political compass or spinning arrow? towards more meaningful evaluations for values and opinions in large language models
Paul Röttger, Valentin Hofmann, Valentina Pyatkin, Musashi Hinck, Hannah Kirk, Hinrich Schuetze, and Dirk Hovy. 2024 · 2024
Closest in time.
Personagym: Evaluating persona agents and llms
Vinay Samuel, Henry Peng Zou, Yue Zhou, Shreyas Chaudhari, Ashwin Kalyan, Tanmay Rajpurohit, Ameet Deshpande, Karthik Narasimhan, and Vishvak Murahari. 2024 · 2024
Closest in time.
Can llms generate novel research ideas? a large-scale human study with 100+ nlp researchers
Chenglei Si, Diyi Yang, and Tatsunori Hashimoto. 2024 · 2024
Closest in time.
A roadmap to pluralistic alignment
Taylor Sorensen, Jared Moore, Jillian Fisher, Mitchell Gordon, Niloofar Mireshghallah, Christopher Michael Rytting, Andre Ye, Liwei Jiang, Ximing Lu, Nouha Dziri, Tim Althoff, and Yejin Choi. 2024 · 2024
Closest in time.
Aligning large language models with diverse political viewpoints
Dominik Stammbach, Philine Widmer, Eunjung Cho, Caglar Gulcehre, and Elliott Ash. 2024 · 2024
Closest in time.
Top books of 2024
The New York Times. 2024 · 2024
Closest in time.
“my answer is C”: First-token probabilities do not match text answers in instruction-tuned language models
Xinpeng Wang, Bolei Ma, Chengzhi Hu, Leon Weber-Genzel, Paul Röttger, Frauke Kreuter, Dirk Hovy, and Barbara Plank. 2024c · 2024
Closest in time.
The generative AI paradox: “what it can create, it may not understand”
Peter West, Ximing Lu, Nouha Dziri, Faeze Brahman, Linjie Li, Jena D. Hwang, Liwei Jiang, Jillian Fisher, Abhilasha Ravichander, Khyathi Chandu, Benjamin Newman, Pang Wei Koh, Allyson Ettinger, and Yejin Choi. 2024 · 2024
Closest in time.
WorldValuesBench: A large-scale benchmark dataset for multi-cultural value awareness of language models
Wenlong Zhao, Debanjan Mondal, Niket Tandon, Danica Dillion, Kurt Gray, and Yuling Gu. 2024 · 2024
Closest in time.
SOTOPIA: Interactive evaluation for social intelligence in language agents
Xuhui Zhou, Hao Zhu, Leena Mathur, Ruohong Zhang, Haofei Yu, Zhengyang Qi, Louis-Philippe Morency, Yonatan Bisk, Daniel Fried, Graham Neubig, and Maarten Sap. 2024 · 2024
Closest in time.
Recovering mental representations from large language models with markov chain monte carlo
Jian-Qiao Zhu, Haijiang Yan, and Thomas L. Griffiths. 2024 · 2024
Closest in time.
Can large language models transform computational social science?
Caleb Ziems, William Held, Omar Shaikh, Jiaao Chen, Zhehao Zhang, and Diyi Yang. 2024 · 2024
Closest in time.