Fetching the paper…
Reading the bibliography…
People increasingly rely on Large Language Models (LLMs) for moral advice, which may influence humans' decisions.
A value pluralism model of ideological reasoning
Philip E Tetlock. 1986 · 1986
Earlier work this paper cites.
Latent dirichlet allocation
David M Blei, Andrew Y Ng, and Michael I Jordan. 2003 · 2003
Earlier work this paper cites.
Is pluralism a viable model of diversity? the benefits and limits of subgroup respect
Yuen J Huo and Ludwin E Molina. 2006 · 2006
Earlier work this paper cites.
Moral foundations theory: The pragmatic validity of moral pluralism
Jesse Graham, Jonathan Haidt, Sena Koleva, Matt Motyl, Ravi Iyer, Sean P Wojcik, and Peter H Ditto. 2013 · 2013
Earlier work this paper cites.
An empirical exploration of moral foundations theory in partisan news sources
Dean Fulgoni, Jordan Carpenter, Lyle Ungar, and Daniel Preoţiuc-Pietro. 2016 · 2016
Earlier work this paper cites.
Moral foundations theory: On the advantages of moral pluralism over moral monism
Jesse Graham, Jonathan Haidt, Matt Motyl, Peter Meindl, Carol Iskiwitch, and Marlon Mooijman. 2018 · 2018
Earlier work this paper cites.
Do language models have beliefs? methods for detecting, updating, and visualizing model beliefs
Peter Hase, Mona Diab, Asli Celikyilmaz, Xian Li, Zornitsa Kozareva, Veselin Stoyanov, Mohit Bansal, and Srinivasan Iyer. 2021 · 2021
Earlier work this paper cites.
The extended moral foundations dictionary (emfd): Development and applications of a crowd-sourced approach to extracting moral intuitions from text
Frederic R Hopp, Jacob T Fisher, Devin Cornell, Richard Huskey, and René Weber. 2021 · 2021
Earlier work this paper cites.
Can machines learn morality? the delphi experiment
Liwei Jiang, Jena D Hwang, Chandra Bhagavatula, Ronan Le Bras, Jenny Liang, Jesse Dodge, Keisuke Sakaguchi, Maxwell Forbes, Jon Borchardt, Saadia Gabriel, and 1 others. 2021 · 2021
Earlier work this paper cites.
Quantifying social organization and political polarization in online platforms
Isaac Waller and Ashton Anderson. 2021 · 2021
Earlier work this paper cites.
Explainable patterns for distinction and prediction of moral judgement on reddit
Ion Stagkos Efstathiadis, Guilherme Paulino-Passos, and Francesca Toni. 2022 · 2022
Earlier work this paper cites.
Disentangling active and passive cosponsorship in the us congress
Giuseppe Russo, Christoph Gote, Laurence Brandenberger, Sophia Schlosser, and Frank Schweitzer. 2022 · 2022
Earlier work this paper cites.
Moral foundations of large language models
Marwa Abdulhai, Gregory Serapio-Garcia, Clément Crepy, Daria Valter, John Canny, and Natasha Jaques. 2023 · 2023
Earlier work this paper cites.
Towards measuring the representation of subjective global opinions in language models
Esin Durmus, Karina Nguyen, Thomas I Liao, Nicholas Schiefer, Amanda Askell, Anton Bakhtin, Carol Chen, Zac Hatfield-Dodds, Danny Hernandez, Nicholas Joseph, and 1 others. 2023 · 2023
Earlier work this paper cites.
Discovering language model behaviors with model-written evaluations
Ethan Perez, Sam Ringer, Kamile Lukosiute, Karina Nguyen, Edwin Chen, Scott Heiner, Craig Pettit, Catherine Olsson, Sandipan Kundu, Saurav Kadavath, and 1 others. 2023 · 2023
Cited alongside, same era.
Spillover of antisocial behavior from fringe platforms: The unintended consequences of community banning
Giuseppe Russo, Luca Verginer, Manoel Horta Ribeiro, and Giona Casiraghi. 2023 · 2023
Cited alongside, same era.
Whose opinions do language models reflect?
Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee, Percy Liang, and Tatsunori Hashimoto. 2023 · 2023
Cited alongside, same era.
Evaluating the moral beliefs encoded in llms
Nino Scherrer, Claudia Shi, Amir Feder, and David Blei. 2023 · 2023
Cited alongside, same era.
Evaluating gender bias of llms in making morality judgements
Divij Bajaj, Yuanyuan Lei, Jonathan Tong, and Ruihong Huang. 2024 · 2024
Cited alongside, same era.
Are large language models consistent over value-laden questions?
Jared Moore, Tanvi Deshpande, and Diyi Yang. 2024 · 2024
Later among the works it cites.
Moralbert: A fine-tuned language model for capturing moral values in social discussions
Vjosa Preniqi, Iacopo Ghinassi, Julia Ive, Charalampos Saitis, and Kyriaki Kalimeri. 2024 · 2024
Later among the works it cites.
Paul Röttger, Valentin Hofmann, Valentina Pyatkin, Musashi Hinck, Hannah Rose Kirk, Hinrich Schütze, and Dirk Hovy. 2024 · 2024
Later among the works it cites.
Localizing paragraph memorization in language models, 2024
Niklas Stoehr, Mitchell Gordon, Chiyuan Zhang, and Owen Lewis · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tim R Davidson, Viacheslav Surkov, Veniamin Veselovsky, Giuseppe Russo, Robert West, and Caglar Gulcehre. 2024 · 2024
Cited alongside, same era.
Questioning the survey responses of large language models
Ricardo Dominguez-Olmedo, Moritz Hardt, and Celestine Mendler-Dünner. 2024 · 2024
Cited alongside, same era.
Modular pluralism: Pluralistic alignment via multi-llm collaboration
Shangbin Feng, Taylor Sorensen, Yuhan Liu, Jillian Fisher, Chan Young Park, Yejin Choi, and Yulia Tsvetkov. 2024 · 2024
Cited alongside, same era.
Collective constitutional ai: Aligning a language model with public input
Saffron Huang, Divya Siddarth, Liane Lovitt, Thomas I Liao, Esin Durmus, Alex Tamkin, and Deep Ganguli. 2024 · 2024
Cited alongside, same era.
Moralbench: Moral evaluation of llms
Jianchao Ji, Yutong Chen, Mingyu Jin, Wujiang Xu, Wenyue Hua, and Yongfeng Zhang. 2024 · 2024
Cited alongside, same era.
Persona is a double-edged sword: Enhancing the zero-shot reasoning by ensembling the role-playing and neutral prompts
Junseok Kim, Nakyeong Yang, and Kyomin Jung. 2024 · 2024
Cited alongside, same era.
The ai review lottery: Widespread ai-assisted peer reviews boost paper scores and acceptance rates
Giuseppe Russo Latona, Manoel Horta Ribeiro, Tim R Davidson, Veniamin Veselovsky, and Robert West. 2024 · 2024
Cited alongside, same era.
Alex Tamkin, Miles McCain, Kunal Handa, Esin Durmus, Liane Lovitt, Ankur Rathi, Saffron Huang, Alfred Mountfield, Jerry Hong, Stuart Ritchie, and 1 others. 2024 · 2024
Later among the works it cites.
Do llms exhibit human-like response biases? a case study in survey design
Lindia Tjuatja, Valerie Chen, Tongshuang Wu, Ameet Talwalkwar, and Graham Neubig. 2024 · 2024
Later among the works it cites.
Replacing judges with juries: Evaluating llm generations with a panel of diverse models
Pat Verga, Sebastian Hofstatter, Sophia Althammer, Yixuan Su, Aleksandra Piktus, Arkady Arkhangorodsky, Minjie Xu, Naomi White, and Patrick Lewis. 2024 · 2024
Later among the works it cites.
Socialgaze: Improving the integration of human social norms in large language models
Anvesh Rao Vijjini, Rakesh R Menon, Jiayi Fu, Shashank Srivastava, and Snigdha Chaturvedi. 2024 · 2024
Later among the works it cites.
Ai language model rivals expert ethicist in perceived moral expertise
Danica Dillion, Debanjan Mondal, Niket Tandon, and Kurt Gray. 2025 · 2025
Closest in time.
Which economic tasks are performed with ai? evidence from millions of claude conversations
Kunal Handa, Alex Tamkin, Miles McCain, Saffron Huang, Esin Durmus, Sarah Heck, Jared Mueller, Jerry Hong, Stuart Ritchie, Tim Belonax, and 1 others. 2025 · 2025
Closest in time.
Values in the wild: Discovering and analyzing values in real-world language model interactions
Saffron Huang, Esin Durmus, Miles McCain, Kunal Handa, Alex Tamkin, Jerry Hong, Michael Stern, Arushi Somani, Xiuruo Zhang, and Deep Ganguli. 2025 · 2025
Closest in time.
What’s the most important value? invp: Investigating the value priorities of llms through decision-making in social scenarios
Xuelin Liu, Pengyuan Liu, and Dong Yu. 2025 · 2025
Closest in time.
Does content moderation lead users away from fringe movements? evidence from a recovery community
Giuseppe Russo, Maciej Styczen, Manoel Horta Ribeiro, and Robert West. 2025 · 2025
Closest in time.
Normative evaluation of large language models with everyday moral dilemmas
Pratik S Sachdeva and Tom van Nuenen. 2025 · 2025
Closest in time.