Fetching the paper…
Reading the bibliography…
As users increasingly seek guidance from LLMs for decision-making in daily life, many of these decisions are not clear-cut and depend significantly on the personal values and ethical standards of people.
The ethics of aristotle
James Alexander Kerr Thomson · 1956
Earlier work this paper cites.
A theory of human motivation
Abraham H Maslow · 1969
Earlier work this paper cites.
A psychoevolutionary theory of emotions, 1982
Robert Plutchik · 1982
Earlier work this paper cites.
I, robot , volume 1
Isaac Asimov · 2004
Earlier work this paper cites.
Utilitarianism and other essays
Jeremy Bentham and John Stuart Mill · 2004
Earlier work this paper cites.
Moral machines: Teaching robots right from wrong
Wendell Wallach and Colin Allen · 2008
Earlier work this paper cites.
Natural language processing with Python: analyzing text with the natural language toolkit
Steven Bird, Ewan Klein, and Edward Loper · 2009
Earlier work this paper cites.
Infinite ethics
Nick Bostrom · 2011
Earlier work this paper cites.
Moral foundations theory: The pragmatic validity of moral pluralism
Jesse Graham, Jonathan Haidt, Sena Koleva, Matt Motyl, Ravi Iyer, Sean P Wojcik, and Peter H Ditto · 2013
Earlier work this paper cites.
Prospect theory: An analysis of decision under risk
Daniel Kahneman and Amos Tversky · 2013
Earlier work this paper cites.
Child’s Conception of Space: Selected Works vol 4
Jean Piaget · 2013
Earlier work this paper cites.
The categorical imperative
Immanuel Kant · 2015
Earlier work this paper cites.
Pareto principles in infinite ethics
Amanda Askell · 2018
Earlier work this paper cites.
Virtue Ethics
Rosalind Hursthouse and Glen Pettigrove · 2018
Earlier work this paper cites.
Social chemistry 101: Learning to reason about social and moral norms
Maxwell Forbes, Jena D Hwang, Vered Shwartz, Maarten Sap, and Yejin Choi · 2020
Earlier work this paper cites.
Aligning ai with shared human values
Dan Hendrycks, Collin Burns, Steven Basart, Andrew Critch, Jerry Li, Dawn Song, and Jacob Steinhardt · 2020
Earlier work this paper cites.
A general language assistant as a laboratory for alignment
Amanda Askell, Yuntao Bai, Anna Chen, Dawn Drain, Deep Ganguli, Tom Henighan, Andy Jones, Nicholas Joseph, Ben Mann, Nova DasSarma, et al · 2021
Earlier work this paper cites.
Can machines learn morality? the delphi experiment
Liwei Jiang, Jena D Hwang, Chandra Bhagavatula, Ronan Le Bras, Jenny Liang, Jesse Dodge, Keisuke Sakaguchi, Maxwell Forbes, Jon Borchardt, Saadia Gabriel, et al · 2021
Earlier work this paper cites.
Scruples: A corpus of community ethical judgments on 32,000 real-life anecdotes
Nicholas Lourie, Ronan Le Bras, and Yejin Choi · 2021
Cited alongside, same era.
Quantifying social organization and political polarization in online platforms
Isaac Waller and Ashton Anderson · 2021
Cited alongside, same era.
Training a helpful and harmless assistant with reinforcement learning from human feedback, 2022
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, Nicholas Joseph, Saurav Kadavath, Jackson Kernion, Tom Conerly, Sheer El-Showk, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, Tristan Hume, Scott Johnston, Shauna Kravec, Liane Lovitt, Neel Nanda, Catherine Olsson, Dario Amodei, Tom Brown, Jack Clark, Sam McCandlish, Chris Olah, Ben Mann, and Jared Kaplan · 2022
Cited alongside, same era.
Understanding dataset difficulty with v-usable information
Kawin Ethayarajh, Yejin Choi, and Swabha Swayamdipta · 2022
Cited alongside, same era.
Can machines learn morality? the delphi experiment, 2022
Liwei Jiang, Jena D. Hwang, Chandra Bhagavatula, Ronan Le Bras, Jenny Liang, Jesse Dodge, Keisuke Sakaguchi, Maxwell Forbes, Jon Borchardt, Saadia Gabriel, Yulia Tsvetkov, Oren Etzioni, Maarten Sap, Regina Rini, and Yejin Choi · 2022
Reflexion: Language agents with verbal reinforcement learning
Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao · 2023
Later among the works it cites.
Voyager: An open-ended embodied agent with large language models
Guanzhi Wang, Yuqi Xie, Yunfan Jiang, Ajay Mandlekar, Chaowei Xiao, Yuke Zhu, Linxi Fan, and Anima Anandkumar · 2023
Later among the works it cites.
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al · 2023
Later among the works it cites.
Claude’s Constitution
Anthropic · 2024
Closest in time.
Vibecheck: Discover and quantify qualitative differences in large language models
Lisa Dunlap, Krishna Mandal, Trevor Darrell, Jacob Steinhardt, and Joseph E Gonzalez · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
When to make exceptions: Exploring language models as accounts of human moral judgment, 2022
Zhijing Jin, Sydney Levine, Fernando Gonzalez, Ojasv Kamal, Maarten Sap, Mrinmaya Sachan, Rada Mihalcea, Josh Tenenbaum, and Bernhard Schölkopf · 2022
Cited alongside, same era.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R Brown, Adam Santoro, Aditya Gupta, Adrià Garriga-Alonso, et al · 2022
Cited alongside, same era.
Probing pre-trained language models for cross-cultural differences in values
Arnav Arora, Lucie-aimée Kaffee, and Isabelle Augenstein · 2023
Cited alongside, same era.
Assessing cross-cultural alignment between ChatGPT and human societies: An empirical study
Yong Cao, Li Zhou, Seolhwa Lee, Laura Cabello, Min Chen, and Daniel Hershcovich · 2023
Cited alongside, same era.
Towards measuring the representation of subjective global opinions in language models
Esin Durmus, Karina Nguyen, Thomas I Liao, Nicholas Schiefer, Amanda Askell, Anton Bakhtin, Carol Chen, Zac Hatfield-Dodds, Danny Hernandez, Nicholas Joseph, et al · 2023
Cited alongside, same era.
Do the rewards justify the means? measuring trade-offs between rewards and ethical behavior in the machiavelli benchmark
Alexander Pan, Jun Shern Chan, Andy Zou, Nathaniel Li, Steven Basart, Thomas Woodside, Hanlin Zhang, Scott Emmons, and Dan Hendrycks · 2023
Cited alongside, same era.
Generative agents: Interactive simulacra of human behavior
Joon Sung Park, Joseph O’Brien, Carrie Jun Cai, Meredith Ringel Morris, Percy Liang, and Michael S Bernstein · 2023
Cited alongside, same era.
Closest in time.
Openassistant conversations-democratizing large language model alignment
Andreas Köpf, Yannic Kilcher, Dimitri von Rütte, Sotiris Anagnostidis, Zhi Rui Tam, Keith Stevens, Abdullah Barhoum, Duc Nguyen, Oliver Stanley, Richárd Nagyfi, et al · 2024
Closest in time.
Moral Dilemmas
Terrance McConnell · 2024
Closest in time.
Universal declaration of human rights
United Nations · 2024
Closest in time.
Model Spec
OpenAI · 2024
Closest in time.
2019 Subscriber Survey Data Dump!
r/AmItheAsshole · 2024
Closest in time.
Value kaleidoscope: Engaging ai with pluralistic human values, rights, and duties
Taylor Sorensen, Liwei Jiang, Jena D Hwang, Sydney Levine, Valentina Pyatkin, Peter West, Nouha Dziri, Ximing Lu, Kavel Rao, Chandra Bhagavatula, et al · 2024
Closest in time.
How to generate text: using different decoding methods for language generation with Transformers
Patrick von Platen · 2024
Closest in time.
Helpsteer2-preference: Complementing ratings with preferences
Zhilin Wang, Alexander Bukharin, Olivier Delalleau, Daniel Egert, Gerald Shen, Jiaqi Zeng, Oleksii Kuchaiev, and Yi Dong · 2024
Closest in time.
WVS Cultural Map: 2023 Version Released
WVS · 2024
Closest in time.
Helm instruct: A multidimensional instruction following evaluation framework with absolute ratings, February 2024
Yian Zhang, Yifan Mai, Josselin Somerville Roberts, Rishi Bommasani, Yann Dubois, and Percy Liang · 2024
Closest in time.
Model Spec
OpenAI · 2025
Closest in time.
Helpsteer 2: Open-source dataset for training top-performing reward models
Zhilin Wang, Yi Dong, Olivier Delalleau, Jiaqi Zeng, Gerald Shen, Daniel Egert, Jimmy Zhang, Makesh Narsimhan Sreedhar, and Oleksii Kuchaiev · 2025
Closest in time.