Fetching the paper…
Reading the bibliography…
Making moral judgments is an essential step toward developing ethical AI systems.
Runaround
Isaac Asimov. 1942 · 1942
Earlier work this paper cites.
Outline of a decision procedure for ethics
John Rawls. 1951 · 1951
Earlier work this paper cites.
The claim to moral adequacy of a highest stage of moral judgment
Lawrence Kohlberg. 1973 · 1973
Earlier work this paper cites.
Ethics: Inventing right and wrong
John Mackie. 1990 · 1990
Earlier work this paper cites.
Virtue ethics , volume 10
Roger Crisp and Michael Slote. 1997 · 1997
Earlier work this paper cites.
Moral dumbfounding: When intuition finds no reason
Jonathan Haidt, Fredrik Bjorklund, and Scott Murphy. 2000 · 2000
Earlier work this paper cites.
How (and where) does moral judgment work?
Joshua Greene and Jonathan Haidt. 2002 · 2002
Earlier work this paper cites.
Artificial morality: Top-down, bottom-up, and hybrid approaches
Colin Allen, Iva Smit, and Wendell Wallach. 2005 · 2005
Earlier work this paper cites.
The nature, importance, and difficulty of machine ethics
James H Moor. 2006 · 2006
Earlier work this paper cites.
Deontological ethics
Larry Alexander and Michael Moore. 2007 · 2007
Earlier work this paper cites.
Machine ethics: Creating an ethical intelligent agent
Michael Anderson and Susan Leigh Anderson. 2007 · 2007
Earlier work this paper cites.
Moral psychology, vol 2: The cognitive science of morality: Intuition and diversity
Walter Ed Sinnott-Armstrong. 2008 · 2008
Earlier work this paper cites.
Theory of mind ability in the behavioural variant of frontotemporal dementia: an analysis of the neural, cognitive, and social levels
Mauro Adenzato, Marco Cavallo, and Ivan Enrici. 2010 · 2010
Earlier work this paper cites.
The weirdest people in the world?
Joseph Henrich, Steven J Heine, and Ara Norenzayan. 2010 · 2010
Earlier work this paper cites.
Consequentialism. i stanford encyclopedia of philosophy
AS Sinnot. 2012 · 2012
Earlier work this paper cites.
Moral foundations theory: The pragmatic validity of moral pluralism
Jesse Graham, Jonathan Haidt, Sena Koleva, Matt Motyl, Ravi Iyer, Sean P Wojcik, and Peter H Ditto. 2013 · 2013
Earlier work this paper cites.
Aristotle: nicomachean ethics
Roger Crisp. 2014 · 2014
Earlier work this paper cites.
The curious tale of julie and mark: Unraveling the moral dumbfounding effect
Edward B Royzman, Kwanwoo Kim, and Robert F Leeman. 2015 · 2015
Earlier work this paper cites.
Foundations of the metaphysics of morals
Immanuel Kant. 2016 · 2016
Earlier work this paper cites.
Normative ethics
Shelly Kagan. 2018 · 2018
Earlier work this paper cites.
The theory of dyadic morality: Reinventing moral judgment by redefining harm
Chelsea Schein and Kurt Gray. 2018 · 2018
Earlier work this paper cites.
The dark side of ethical robots
Dieter Vanderelst and Alan Winfield. 2018 · 2018
Earlier work this paper cites.
Ethics beyond computation: Why we can’t (and shouldn’t) replace human moral judgment with algorithms
William Hasselberger. 2019 · 2019
Earlier work this paper cites.
Social chemistry 101: Learning to reason about social and moral norms
Maxwell Forbes, Jena D. Hwang, Vered Shwartz, Maarten Sap, and Yejin Choi. 2020 · 2020
Cited alongside, same era.
Power and moral dilemma judgments: Distinct effects of memory recall versus social roles
Bertram Gawronski and Skylar M Brannon. 2020 · 2020
Cited alongside, same era.
RealToxicityPrompts: Evaluating neural toxic degeneration in language models
Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A. Smith. 2020 · 2020
Cited alongside, same era.
Moral emotions: A review and research agenda for management scholarship
Rebecca Greenbaum, Julena Bonner, Truit Gray, and Mary Mawritz. 2020 · 2020
Cited alongside, same era.
Moral foundations twitter corpus: A collection of 35k tweets annotated for moral sentiment
Joe Hoover, Gwenyth Portillo-Wightman, Leigh Yeh, Shreya Havaldar, Aida Mostafazadeh Davani, Ying Lin, Brendan Kennedy, Mohammad Atari, Zahra Kamel, Madelyn Mendlen, et al. 2020 · 2020
Cited alongside, same era.
Large language models can self-improve
Jiaxin Huang, Shixiang Shane Gu, Le Hou, Yuexin Wu, Xuezhi Wang, Hongkun Yu, and Jiawei Han. 2022 · 2022
Later among the works it cites.
Data contamination: From memorization to exploitation
Inbal Magar and Roy Schwartz. 2022 · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022 · 2022
Later among the works it cites.
Valuenet: A new dataset for human value driven dialogue system
Liang Qiu, Yizhou Zhao, Jinchao Li, Pan Lu, Baolin Peng, Jianfeng Gao, and Song-Chun Zhu. 2022 · 2022
Later among the works it cites.
Annotators with attitudes: How annotator beliefs and identities bias toxic language detection
Maarten Sap, Swabha Swayamdipta, Laura Vianna, Xuhui Zhou, Yejin Choi, and Noah A. Smith. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A theory of justice: Revised edition
John Rawls. 2020 · 2020
Cited alongside, same era.
Social bias frames: Reasoning about social and power implications of language
Maarten Sap, Saadia Gabriel, Lianhui Qin, Dan Jurafsky, Noah A. Smith, and Yejin Choi. 2020 · 2020
Cited alongside, same era.
The importance of context in moral judgments
Chelsea Schein. 2020 · 2020
Cited alongside, same era.
A general language assistant as a laboratory for alignment
Amanda Askell, Yuntao Bai, Anna Chen, Dawn Drain, Deep Ganguli, Tom Henighan, Andy Jones, Nicholas Joseph, Ben Mann, Nova DasSarma, et al. 2021 · 2021
Cited alongside, same era.
On the dangers of stochastic parrots: Can language models be too big?
Emily M Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021 · 2021
Cited alongside, same era.
Anticipating safety issues in e2e conversational ai: Framework and tooling
Emily Dinan, Gavin Abercrombie, A Stevie Bergman, Shannon Spruit, Dirk Hovy, Y-Lan Boureau, and Verena Rieser. 2021 · 2021
Cited alongside, same era.
Moral stories: Situated reasoning about norms, intents, actions, and their consequences
Denis Emelin, Ronan Le Bras, Jena D. Hwang, Maxwell Forbes, and Yejin Choi. 2021 · 2021
Cited alongside, same era.
On the machine learning of ethical judgments from natural language
Zeerak Talat, Hagen Blix, Josef Valvoda, Maya Indira Ganesh, Ryan Cotterell, and Adina Williams. 2022 · 2022
Later among the works it cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022 · 2022
Later among the works it cites.
Towards identifying social bias in dialog systems: Framework, dataset, and benchmark
Jingyan Zhou, Jiawen Deng, Fei Mi, Yitong Li, Yasheng Wang, Minlie Huang, Xin Jiang, Qun Liu, and Helen Meng. 2022 · 2022
Later among the works it cites.
The moral integrity corpus: A benchmark for ethical dialogue systems
Caleb Ziems, Jane Yu, Yi-Chia Wang, Alon Halevy, and Diyi Yang. 2022 · 2022
Later among the works it cites.
Sparks of artificial general intelligence: Early experiments with gpt-4
Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, et al. 2023 · 2023
Closest in time.
Machine ethics: Do androids dream of being good people?
Gonzalo Génova, Valentín Moreno, and M Rosario González. 2023 · 2023
Closest in time.
Theory of mind may have spontaneously emerged in large language models
Michal Kosinski. 2023 · 2023
Closest in time.
What makes chain-of-thought prompting effective? a counterfactual study
Aman Madaan, Katherine Hermann, and Amir Yazdanbakhsh. 2023 · 2023
Closest in time.
Boosting theory-of-mind performance in large language models via prompting
Shima Rahimi Moghaddam and Christopher J Honey. 2023 · 2023
Closest in time.
OpenAI. 2023 · 2023
Closest in time.
Do the rewards justify the means? Measuring trade-offs between rewards and ethical behavior in the machiavelli benchmark
Alexander Pan, Jun Shern Chan, Andy Zou, Nathaniel Li, Steven Basart, Thomas Woodside, Hanlin Zhang, Scott Emmons, and Dan Hendrycks. 2023 · 2023
Closest in time.
Clarifydelphi: Reinforced clarification questions with defeasibility rewards for social and moral situations
Valentina Pyatkin, Jena D Hwang, Vivek Srikumar, Ximing Lu, Liwei Jiang, Yejin Choi, and Chandra Bhagavatula. 2023 · 2023
Closest in time.
Knowledge of cultural moral norms in large language models
Aida Ramezani and Yang Xu. 2023 · 2023
Closest in time.
Moral mimicry: Large language models produce moral rationalizations tailored to political identity
Gabriel Simmons. 2023 · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023 · 2023
Closest in time.
Self-guard: Empower the llm to safeguard itself
Zezhong Wang, Fangkai Yang, Lu Wang, Pu Zhao, Hongru Wang, Liang Chen, Qingwei Lin, and Kam-Fai Wong. 2023 · 2023
Closest in time.
Descriptive ethics — Wikipedia, the free encyclopedia
Wikipedia. 2023 · 2023
Closest in time.
Language models don’t always say what they think: unfaithful explanations in chain-of-thought prompting
Miles Turpin, Julian Michael, Ethan Perez, and Samuel Bowman. 2024 · 2024
Closest in time.