Fetching the paper…
Reading the bibliography…
Large language models (LLMs) appear to bias their survey answers toward certain values.
Persuasion for Good: Towards a Personalized Persuasive Dialogue System for Social Good
Xuewei Wang, Weiyan Shi, Richard Kim, Yoojung Oh, Sijia Yang, Jingwen Zhang, and Zhou Yu. 2020 · 1906
Earlier work this paper cites.
Information radius
Robin Sibson. 1969 · 1969
Earlier work this paper cites.
Intransitivity of preferences
Amos Tversky. 1969 · 1969
Earlier work this paper cites.
A Theory of Justice
John Rawls. 1971 · 1971
Earlier work this paper cites.
Universals in the Content and Structure of Values: Theoretical Advances and Empirical Tests in 20 Countries
Shalom H. Schwartz. 1992 · 1992
Earlier work this paper cites.
The international personality item pool and the future of public-domain personality measures
Lewis R. Goldberg, John A. Johnson, Herbert W. Eber, Robert Hogan, Michael C. Ashton, C. Robert Cloninger, and Harrison G. Gough. 2006 · 2006
Earlier work this paper cites.
Dimensionalizing Cultures: The Hofstede Model in Context
Geert Hofstede. 2011 · 2011
Earlier work this paper cites.
Thinking, fast and slow
Daniel Kahneman. 2011 · 2011
Earlier work this paper cites.
Transitivity of preferences
Michel Regenwetter, Jason Dana, and Clintin P. Davis-Stober. 2011 · 2011
Earlier work this paper cites.
An Overview of the Schwartz Theory of Basic Values
Shalom Schwartz. 2012 · 2012
Earlier work this paper cites.
Refining the theory of basic individual values
Shalom H. Schwartz, Jan Cieciuch, Michele Vecchione, Eldad Davidov, Ronald Fischer, Constanze Beierlein, Alice Ramos, Markku Verkasalo, Jan-Erik Lönnqvist, Kursad Demirutku, Ozlem Dirilen-Gumus, and Mark Konty. 2012 · 2012
Earlier work this paper cites.
Linguistic Harbingers of Betrayal: A Case Study on an Online Strategy Game
Vlad Niculae, Srijan Kumar, Jordan Boyd-Graber, and Cristian Danescu-Niculescu-Mizil. 2015 · 2015
Earlier work this paper cites.
Normative Uncertainty as a Voting Problem
William MacAskill. 2016 · 2016
Earlier work this paper cites.
Questionnaire Design
Jon A. Krosnick. 2018 · 2018
Earlier work this paper cites.
Let’s Make Your Request More Persuasive: Modeling Persuasive Strategies via Semi-Supervised Neural Nets on Crowdfunding Platforms
Diyi Yang, Jiaao Chen, Zichao Yang, Dan Jurafsky, and Eduard Hovy. 2019 · 2019
Earlier work this paper cites.
Climbing towards NLU: On Meaning, Form, and Understanding in the Age of Data
Emily M. Bender and Alexander Koller. 2020 · 2020
Earlier work this paper cites.
On a Generalization of the Jensen–Shannon Divergence and the Jensen–Shannon Centroid
Frank Nielsen. 2020 · 2020
Earlier work this paper cites.
It Takes Two to Lie: One to Lie, and One to Listen
Denis Peskov, Benny Cheng, Ahmed Elgohary, Joe Barrow, Cristian Danescu-Niculescu-Mizil, and Jordan Boyd-Graber. 2020 · 2020
Earlier work this paper cites.
Do Language Models Have Beliefs? Methods for Detecting, Updating, and Visualizing Model Beliefs
Peter Hase, Mona Diab, Asli Celikyilmaz, Xian Li, Zornitsa Kozareva, Veselin Stoyanov, Mohit Bansal, and Srinivasan Iyer. 2021 · 2021
Earlier work this paper cites.
Aligning AI With Shared Human Values
Dan Hendrycks, Collin Burns, Steven Basart, Andrew Critch, Jerry Li, Dawn Song, and Jacob Steinhardt. 2021 · 2021
Earlier work this paper cites.
Delphi: Towards Machine Ethics and Norms
Liwei Jiang, Jena D. Hwang, Chandra Bhagavatula, Ronan Le Bras, Maxwell Forbes, Jon Borchardt, Jenny T. Liang, Oren Etzioni, Maarten Sap, and Yejin Choi. 2021 · 2021
Earlier work this paper cites.
SCRUPLES: A Corpus of Community Ethical Judgments on 32,000 Real-Life Anecdotes
Nicholas Lourie, Ronan Le Bras, and Yejin Choi. 2021 · 2021
Earlier work this paper cites.
A Repository of Schwartz Value Scales with Instructions and an Introduction
Shalom Schwartz. 2021 · 2021
Earlier work this paper cites.
Joshua Albrecht, Ellie Kitanidis, and Abraham J. Fetterman. 2022 · 2022
Earlier work this paper cites.
Experimental Moral Philosophy
Mark Alfano, Edouard Machery, Alexandra Plakias, and Don Loeb. 2022 · 2022
Earlier work this paper cites.
Language Models as Agent Models
Jacob Andreas. 2022 · 2022
Earlier work this paper cites.
Fine-tuning language models to find agreement among humans with diverse preferences
Michiel A. Bakker, Martin J. Chadwick, Hannah R. Sheahan, Michael Henry Tessler, Lucy Campbell-Gillingham, Jan Balaguer, Nat McAleese, Amelia Glaese, John Aslanides, Matthew M. Botvinick, and Christopher Summerfield. 2022 · 2022
Earlier work this paper cites.
General Social Survey, 1972-2022 [Machine-readable data file]
Michael Davern, Rene Bautista, Jeremy Freese, Pamela Herd, and Stephen Morgan. 2022 · 2022
Earlier work this paper cites.
Jury Learning: Integrating Dissenting Voices into Machine Learning Models
Mitchell L. Gordon, Michelle S. Lam, Joon Sung Park, Kayur Patel, Jeff Hancock, Tatsunori Hashimoto, and Michael S. Bernstein. 2022 · 2022
Earlier work this paper cites.
World Values Survey Wave 7 (2017-2022) Cross-National Data-Set
Christian Haerpfer, Ronald Inglehart, Alejandro Moreno, Christian Welzel, Kseniya Kizilova, Jaime Diez-Medrano, Marta Lagos, Pippa Norris, Eduard Ponarin, and Bi Puranen. 2022a · 2022
Earlier work this paper cites.
The Ghost in the Machine has an American accent: value conflict in GPT-3
Rebecca L. Johnson, Giada Pistilli, Natalia Menédez-González, Leslye Denisse Dias Duran, Enrico Panai, Julija Kalpokiene, and Donald Jay Bertulfo. 2022 · 2022
Earlier work this paper cites.
Language Models Understand Us, Poorly
Jared Moore. 2022 · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe. 2022 · 2022
Earlier work this paper cites.
Social Simulacra: Creating Populated Prototypes for Social Computing Systems
Joon Sung Park, Lindsay Popowski, Carrie Cai, Meredith Ringel Morris, Percy Liang, and Michael S. Bernstein. 2022 · 2022
Cited alongside, same era.
ClarifyDelphi: Reinforced Clarification Questions with Defeasibility Rewards for Social and Moral Situations
Valentina Pyatkin, Jena D. Hwang, Vivek Srikumar, Ximing Lu, Liwei Jiang, Yejin Choi, and Chandra Bhagavatula. 2022 · 2022
Cited alongside, same era.
Adversarial Robustness through the Lens of Causality
Yonggang Zhang, Mingming Gong, Tongliang Liu, Gang Niu, Xinmei Tian, Bo Han, B. Schölkopf, and Kun Zhang. 2022 · 2022
Cited alongside, same era.
Probing Pre-Trained Language Models for Cross-Cultural Differences in Values
Arnav Arora, Lucie-Aimée Kaffee, and Isabelle Augenstein. 2023 · 2023
Cited alongside, same era.
Assessing LLMs for Moral Value Pluralism
Noam Benkler, Drisana Mosaphir, Scott Friedman, Andrew Smart, and Sonja Schmer-Galunder. 2023 · 2023
Cited alongside, same era.
Auditing and Mitigating Cultural Bias in LLMs
Yan Tao, Olga Viberg, Ryan S. Baker, and Rene F. Kizilcec. 2023 · 2023
Later among the works it cites.
Do LLMs exhibit human-like response biases? A case study in survey design
Lindia Tjuatja, Valerie Chen, Sherry Tongshuang Wu, Ameet Talwalkar, and Graham Neubig. 2023 · 2023
Later among the works it cites.
Llama 2: Open Foundation and Fine-Tuned Chat Models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Assessing Cross-Cultural Alignment between ChatGPT and Human Societies: An Empirical Study
Yong Cao, Li Zhou, Seolhwa Lee, Laura Cabello, Min Chen, and Daniel Hershcovich. 2023 · 2023
Cited alongside, same era.
Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Stephen Casper, Xander Davies, Claudia Shi, Thomas Krendl Gilbert, Jérémy Scheurer, Javier Rando, Rachel Freedman, Tomasz Korbak, David Lindner, Pedro Freire, Tony Wang, Samuel Marks, Charbel-Raphaël Segerie, Micah Carroll, Andi Peng, Phillip Christoffersen, Mehul Damani, Stewart Slocum, Usman Anwar, Anand Siththaranjan, Max Nadeau, Eric J. Michaud, Jacob Pfau, Dmitrii Krasheninnikov, Xin Chen, Lauro Langosco, Peter Hase, Erdem Bıyık, Anca Dragan, David Krueger, Dorsa Sadigh, and Dylan Hadfield-Menell. 2023 · 2023
Cited alongside, same era.
Do personality tests generalize to Large Language Models?
Florian E. Dorner, Tom Sühr, Samira Samadi, and Augustin Kelava. 2023 · 2023
Cited alongside, same era.
Ronald Fischer, Markus Luczak-Roesch, and Johannes A. Karl. 2023 · 2023
Cited alongside, same era.
When the Majority is Wrong: Modeling Annotator Disagreement for Subjective Tasks
Eve Fleisig, Rediet Abebe, and Dan Klein. 2023 · 2023
Cited alongside, same era.
Off The Rails: Procedural Dilemma Generation for Moral Reasoning
Jan-Philipp Fränken, Ayesha Khawaja, Kanishk Gandhi, Jared Moore, Noah D. Goodman, and Tobias Gerstenberg. 2023 · 2023
Cited alongside, same era.
Understanding Social Reasoning in Language Models with Language Models
Kanishk Gandhi, Jan-Philipp Fränken, Tobias Gerstenberg, and Noah D. Goodman. 2023 · 2023
Cited alongside, same era.
Alexander Wei, Nika Haghtalab, and Jacob Steinhardt. 2023 · 2023
Later among the works it cites.
Value FULCRA: Mapping Large Language Models to the Multidimensional Spectrum of Basic Human Values
Jing Yao, Xiaoyuan Yi, Xiting Wang, Yifan Gong, and Xing Xie. 2023 · 2023
Later among the works it cites.
Assessing Hidden Risks of LLMs: An Empirical Study on Robustness, Consistency, and Credibility
Wentao Ye, Mingfeng Ou, Tianyi Li, Yipeng chen, Xuetao Ma, Yifan Yanggong, Sai Wu, Jie Fu, Gang Chen, Haobo Wang, and Junbo Zhao. 2023 · 2023
Later among the works it cites.
GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts
Jiahao Yu, Xingwei Lin, Zheng Yu, and Xinyu Xing. 2023 · 2023
Later among the works it cites.
Measuring Value Understanding in Language Models through Discriminator-Critique Gap
Zhaowei Zhang, Fengshuo Bai, Jun Gao, and Yaodong Yang. 2023 · 2023
Later among the works it cites.
Group Preference Optimization: Few-Shot Alignment of Large Language Models
Siyan Zhao, John Dang, and Aditya Grover. 2023 · 2023
Later among the works it cites.
Navigating the Grey Area: How Expressions of Uncertainty and Overconfidence Affect Language Models
Kaitlyn Zhou, Dan Jurafsky, and Tatsunori Hashimoto. 2023 · 2023
Later among the works it cites.
Towards Measuring and Modeling “Culture”
Muhammad Farid Adilazuarda, Sagnik Mukherjee, Pradhyumna Lavania, Siddhant Singh, Ashutosh Dwivedi, Alham Fikri Aji, Jacki O’Neill, Ashutosh Modi, and Monojit Choudhury. 2024 · 2024
Closest in time.
Many-shot Jailbreaking
Cem Anil, Esin Durmus, Mrinank Sharma, Joe Benton, Sandipan Kundu, Joshua Batson, Nina Rimsky, Meg Tong, Jesse Mu, Daniel Ford, Francesco Mosconi, Rajashree Agrawal, Rylan Schaeffer, Naomi Bashkansky, Samuel Svenningsen, Mike Lambert, Ansh Radhakrishnan, Carson Denison, Evan J Hubinger, Yuntao Bai, Trenton Bricken, Timothy Maxwell, Nicholas Schiefer, Jamie Sully, Alex Tamkin, Tamera Lanham, Karina Nguyen, Tomasz Korbak, Jared Kaplan, Deep Ganguli, Samuel R Bowman, Ethan Perez, Roger Grosse, and David Duvenaud. 2024 · 2024
Closest in time.
The Echoes of Multilinguality: Tracing Cultural Value Shifts during LM Fine-tuning
Rochelle Choenni, Anne Lauscher, and Ekaterina Shutova. 2024 · 2024
Closest in time.
Bias-Augmented Consistency Training Reduces Biased Reasoning in Chain-of-Thought
James Chua, Edward Rees, Hunar Batra, Samuel R. Bowman, Julian Michael, Ethan Perez, and Miles Turpin. 2024 · 2024
Closest in time.
Towards Measuring the Representation of Subjective Global Opinions in Language Models
Esin Durmus, Karina Nguyen, Thomas I. Liao, Nicholas Schiefer, Amanda Askell, Anton Bakhtin, Carol Chen, Zac Hatfield-Dodds, Danny Hernandez, Nicholas Joseph, Liane Lovitt, Sam McCandlish, Orowa Sikder, Alex Tamkin, Janel Thamkul, Jared Kaplan, Jack Clark, and Deep Ganguli. 2024 · 2024
Closest in time.
Auxiliary task demands mask the capabilities of smaller language models
Jennifer Hu and Michael C. Frank. 2024 · 2024
Closest in time.
Speech and Language Processing , 3rd ed. draft edition
Daniel Jurafsky and James H. Martin. 2024 · 2024
Closest in time.
Debating with More Persuasive LLMs Leads to More Truthful Answers
Akbir Khan, John Hughes, Dan Valentine, Laura Ruis, Kshitij Sachan, Ansh Radhakrishnan, Edward Grefenstette, Samuel R. Bowman, Tim Rocktäschel, and Ethan Perez. 2024 · 2024
Closest in time.
What are human values, and how do we align AI to them?
Oliver Klingefjord, Ryan Lowe, and Joe Edelman. 2024 · 2024
Closest in time.
Evaluating Large Language Model Biases in Persona-Steered Generation
Andy Liu, Mona Diab, and Daniel Fried. 2024 · 2024
Closest in time.
Beyond Probabilities: Unveiling the Misalignment in Evaluating Large Language Models
Chenyang Lyu, Minghao Wu, and Alham Fikri Aji. 2024 · 2024
Closest in time.
State of What Art? A Call for Multi-Prompt LLM Evaluation
Moran Mizrahi, Guy Kaplan, Dan Malkin, Rotem Dror, Dafna Shahaf, and Gabriel Stanovsky. 2024 · 2024
Closest in time.
Having Beer after Prayer? Measuring Cultural Bias in Large Language Models
Tarek Naous, Michael J. Ryan, Alan Ritter, and Wei Xu. 2024 · 2024
Closest in time.
NORMAD: A Benchmark for Measuring the Cultural Adaptability of Large Language Models
Abhinav Rao, Akhila Yerukola, Vishwa Shah, Katharina Reinecke, and Maarten Sap. 2024 · 2024
Closest in time.
Paul Röttger, Valentin Hofmann, Valentina Pyatkin, Musashi Hinck, Hannah Rose Kirk, Hinrich Schütze, and Dirk Hovy. 2024 · 2024
Closest in time.
A Roadmap to Pluralistic Alignment
Taylor Sorensen, Jared Moore, Jillian Fisher, Mitchell Gordon, Niloofar Mireshghallah, Christopher Michael Rytting, Andre Ye, Liwei Jiang, Ximing Lu, Nouha Dziri, Tim Althoff, and Yejin Choi. 2024 · 2024
Closest in time.
Revealing Fine-Grained Values and Opinions in Large Language Models
Dustin Wright, Arnav Arora, Nadav Borenstein, Srishti Yadav, Serge Belongie, and Isabelle Augenstein. 2024 · 2024
Closest in time.
Language Models as Critical Thinking Tools: A Case Study of Philosophers
Andre Ye, Jared Moore, Rose Novick, and Amy X. Zhang. 2024 · 2024
Closest in time.
Yi: Open Foundation Models by 01.AI
Alex Young, Bei Chen, Chao Li, Chengen Huang, Ge Zhang, Guanwei Zhang, Heng Li, Jiangcheng Zhu, Jianqun Chen, Jing Chang, Kaidong Yu, Peng Liu, Qiang Liu, Shawn Yue, Senbin Yang, Shiming Yang, Tao Yu, Wen Xie, Wenhao Huang, Xiaohui Hu, Xiaoyi Ren, Xinyao Niu, Pengcheng Nie, Yuchi Xu, Yudong Liu, Yue Wang, Yuxuan Cai, Zhenyu Gu, Zhiyuan Liu, and Zonghong Dai. 2024 · 2024
Closest in time.
Yi Zeng, Hongpeng Lin, Jingwen Zhang, Diyi Yang, Ruoxi Jia, and Weiyan Shi. 2024 · 2024
Closest in time.
Wenlong Zhao, Debanjan Mondal, Niket Tandon, Danica Dillion, Kurt Gray, and Yuling Gu. 2024 · 2024
Closest in time.
Relying on the Unreliable: The Impact of Language Models’ Reluctance to Express Uncertainty
Kaitlyn Zhou, Jena D. Hwang, Xiang Ren, and Maarten Sap. 2024 · 2024
Closest in time.