Fetching the paper…
Reading the bibliography…
In this study, we measure the moral reasoning ability of LLMs using the Defining Issues Test - a psychometric instrument developed for measuring the moral development stage of a person according to the Kohlberg's Cognitive Moral Development Model.
The reliability, validity, and design of the defining issues test
Richard M Martin, Michael Shafto, and William Vandeinse · 1977
Earlier work this paper cites.
Development in Judging Moral Issues
J. Rest · 1979
Earlier work this paper cites.
The philosophy of moral development: Essays on moral development
Lawrence Kohlberg · 1981
Earlier work this paper cites.
A chinese perspective on kohlberg’s theory of moral development
Dora Shu-Fang Dien · 1982
Earlier work this paper cites.
Kohlberg’s theory of moral development: Critical analysis of validation studies with the defining issues test
Stanley R Kay · 1982
Earlier work this paper cites.
Utilitarianism, moral dilemmas, and moral cost
Michael Slote · 1985
Earlier work this paper cites.
Cross-cultural universality of social-moral development: a critical review of kohlbergian research
John R Snarey · 1985
Earlier work this paper cites.
Integrating care and justice issues in professional moral education: A gender perspective
Muriel J Bebeau and Mary M Brabeck · 1987
Earlier work this paper cites.
DIT Manual: Manual for the Defining Issues Test
J.R. Rest and University of Minnesota. Center for the Study of Ethical Development · 1990
Earlier work this paper cites.
Moral development in the professions: Psychology and applied ethics
James R Rest et al · 1994
Earlier work this paper cites.
An introduction to the moral judgment test (mjt)
Georg Lind · 1998
Earlier work this paper cites.
The emotional dog and its rational tail: a social intuitionist approach to moral judgment
Jonathan Haidt · 2001
Earlier work this paper cites.
Liberal pluralism: The implications of value pluralism for political theory and practice
William A Galston and William Arthur Galston · 2002
Earlier work this paper cites.
The moral judgment of the child
Jean Piaget · 2013
Earlier work this paper cites.
Moral foundations theory: The pragmatic validity of moral pluralism
Jesse Graham, Jonathan Haidt, Sena Koleva, Matt Motyl, Ravi Iyer, Sean P Wojcik, and Peter H Ditto · 2013
Earlier work this paper cites.
Hate me, hate me not: Hate speech detection on facebook
Fabio Del Vigna12, Andrea Cimino23, Felice Dell’Orletta, Marinella Petrocchi, and Maurizio Tesconi · 2017
Earlier work this paper cites.
Fine-tuning language models from human preferences
Daniel M Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving · 2019
Earlier work this paper cites.
Social data: Biases, methodological pitfalls, and ethical boundaries
Alexandra Olteanu, Carlos Castillo, Fernando Diaz, and Emre Kıcıman · 2019
Earlier work this paper cites.
Bert has a moral compass: Improvements of ethical and moral values of machines, 2019
Patrick Schramowski, Cigdem Turan, Sophie Jentzsch, Constantin Rothkopf, and Kristian Kersting · 2019
Cited alongside, same era.
Vaccine mandates, value pluralism, and policy diversity
Mark C. Navin and Katie Attwell · 2019
Cited alongside, same era.
Toxic, hateful, offensive or abusive? what are we really classifying? an empirical analysis of hate speech datasets
Paula Fortuna, Juan Soler, and Leo Wanner · 2020
Cited alongside, same era.
Suicidal ideation detection: A review of machine learning methods and applications
Shaoxiong Ji, Shirui Pan, Xue Li, Erik Cambria, Guodong Long, and Zi Huang · 2020
Cited alongside, same era.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Cited alongside, same era.
Large pre-trained language models contain human-like biases of what is right and wrong to do, 2022
Patrick Schramowski, Cigdem Turan, Nico Andersen, Constantin A. Rothkopf, and Kristian Kersting · 2022
Later among the works it cites.
Is gpt-3 a psychopath? evaluating large language models from a psychological perspective
Xingxuan Li, Yutong Li, Linlin Liu, Lidong Bing, and Shafiq Joty · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Later among the works it cites.
Toxigen: A large-scale machine-generated dataset for adversarial and implicit hate speech detection
Thomas Hartvigsen, Saadia Gabriel, Hamid Palangi, Maarten Sap, Dipankar Ray, and Ece Kamar · 2022
Later among the works it cites.
Evaluating the factual consistency of large language models through summarization
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
The case for taking ai seriously as a threat to humanity, Oct 15, 2020
Kelsey Piper · 2020
Cited alongside, same era.
Aligning ai with shared human values
Dan Hendrycks, Collin Burns, Steven Basart, Andrew Critch, Jerry Li, Dawn Song, and Jacob Steinhardt · 2020
Cited alongside, same era.
Social chemistry 101: Learning to reason about social and moral norms
Maxwell Forbes, Jena D. Hwang, Vered Shwartz, Maarten Sap, and Yejin Choi · 2020
Cited alongside, same era.
Moral stories: Situated reasoning about norms, intents, actions, and their consequences
Denis Emelin, Ronan Le Bras, Jena D Hwang, Maxwell Forbes, and Yejin Choi · 2020
Cited alongside, same era.
Social bias frames: Reasoning about social and power implications of language
Maarten Sap, Saadia Gabriel, Lianhui Qin, Dan Jurafsky, Noah A. Smith, and Yejin Choi · 2020
Cited alongside, same era.
The moral choice machine
Patrick Schramowski, Cigdem Turan, Sophie Jentzsch, Constantin Rothkopf, and Kristian Kersting · 2020
Cited alongside, same era.
Ethical and social risks of harm from language models, 2021
Laura Weidinger, John Mellor, Maribeth Rauh, Conor Griffin, Jonathan Uesato, Po-Sen Huang, Myra Cheng, Mia Glaese, Borja Balle, Atoosa Kasirzadeh, Zac Kenton, Sasha Brown, Will Hawkins, Tom Stepleton, Courtney Biles, Abeba Birhane, Julia Haas, Laura Rimell, Lisa Anne Hendricks, William Isaac, Sean Legassick, Geoffrey Irving, and Iason Gabriel · 2021
Cited alongside, same era.
Derek Tam, Anisha Mascarenhas, Shiyue Zhang, Sarah Kwan, Mohit Bansal, and Colin Raffel · 2022
Later among the works it cites.
On the machine learning of ethical judgments from natural language
Zeerak Talat, Hagen Blix, Josef Valvoda, Maya Indira Ganesh, Ryan Cotterell, and Adina Williams · 2022
Later among the works it cites.
Propile: Probing privacy leakage in large language models
Siwon Kim, Sangdoo Yun, Hwaran Lee, Martin Gubri, Sungroh Yoon, and Seong Joon Oh · 2023
Closest in time.
Yejin Bang, Samuel Cahyawijaya, Nayeon Lee, Wenliang Dai, Dan Su, Bryan Wilie, Holy Lovenia, Ziwei Ji, Tiezheng Yu, Willy Chung, et al · 2023
Closest in time.
Siren’s song in the ai ocean: A survey on hallucination in large language models
Yue Zhang, Yafu Li, Leyang Cui, Deng Cai, Lemao Liu, Tingchen Fu, Xinting Huang, Enbo Zhao, Yu Zhang, Yulong Chen, et al · 2023
Closest in time.
Sources of hallucination by large language models on inference tasks
Nick McKenna, Tianyi Li, Liang Cheng, Mohammad Javad Hosseini, Mark Johnson, and Mark Steedman · 2023
Closest in time.
From instructions to intrinsic human values–a survey of alignment goals for big models
Jing Yao, Xiaoyuan Yi, Xiting Wang, Jindong Wang, and Xing Xie · 2023
Closest in time.
Value kaleidoscope: Engaging ai with pluralistic human values, rights, and duties
Taylor Sorensen, Liwei Jiang, Jena Hwang, Sydney Levine, Valentina Pyatkin, Peter West, Nouha Dziri, Ximing Lu, Kavel Rao, Chandra Bhagavatula, et al · 2023
Closest in time.
Rethinking machine ethics–can llms perform moral reasoning through the lens of moral theories?
Jingyan Zhou, Minda Hu, Junan Li, Xiaoying Zhang, Xixin Wu, Irwin King, and Helen Meng · 2023
Closest in time.
Lawrence Kohlberg’s stages of moral development
Cheryl E Sanders · 2023
Closest in time.
Gpt-4 technical report, 2023
OpenAI · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Closest in time.
Aligning ai with shared human values, 2023
Dan Hendrycks, Collin Burns, Steven Basart, Andrew Critch, Jerry Li, Dawn Song, and Jacob Steinhardt · 2023
Closest in time.
How is chatgpt’s behavior changing over time?
Lingjiao Chen, Matei Zaharia, and James Zou · 2023
Closest in time.