Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have become central to advancing automation and decision-making across various sectors, raising significant ethical questions.
Rice’s theorem
Dexter C Kozen and Dexter C Kozen · 1977
Earlier work this paper cites.
On a supposed right to lie because of philanthropic concerns
Immanuel Kant · 1993
Earlier work this paper cites.
Emotions in human and artificial intelligence
J. Martínez-Miranda and A. Aldea · 2004
Earlier work this paper cites.
Intuitive ethics : How innately prepared intuitions generate culturally variable virtues
Jonathan Haidt and Craig Joseph · 2004
Earlier work this paper cites.
Emotion and cognition : insights from studies of the human amygdala
E. Phelps · 2006
Earlier work this paper cites.
Moral judgment development across cultures : Revisiting kohlberg’s universality claims
J. Gibbs, K. Basinger, Rebecca L. Grime, and J. Snarey · 2007
Earlier work this paper cites.
Emotion and rationality : A critical review and interpretation of empirical evidence
Michel Tuan Pham · 2007
Earlier work this paper cites.
Personal values and behavior : Taking the cultural context into account
Sonia Roccas and Lilach Sagiv · 2009
Earlier work this paper cites.
Liberals and conservatives rely on different sets of moral foundations
Jesse Graham, Jonathan Haidt, and Brian A Nosek · 2009
Earlier work this paper cites.
The weirdest people in the world ?
Joseph Henrich, Steven J Heine, and Ara Norenzayan · 2010
Earlier work this paper cites.
Automation bias : a systematic review of frequency, effect mediators, and mitigators
Kate Goddard, Abdul Roudsari, and Jeremy C Wyatt · 2012
Earlier work this paper cites.
Moral intuitions and political orientation : Similarities and differences between south korea and the united states
Kisok R Kim, Je-Sang Kang, and Seongyi Yun · 2012
Earlier work this paper cites.
The moral stereotypes of liberals and conservatives : Exaggeration of differences across the political spectrum
Jesse Graham, Brian A Nosek, and Jonathan Haidt · 2012
Earlier work this paper cites.
The righteous mind : Why good people are divided by politics and religion
Jonathan Haidt · 2012
Earlier work this paper cites.
The impact of emotion on perception, attention, memory, and decision-making
T. Brosch, K. Scherer, D. Grandjean, and D. Sander · 2013
Earlier work this paper cites.
Does cultural exposure partially explain the association between personality and political orientation ?
Xiaowen Xu, Raymond A Mar, and Jordan B Peterson · 2013
Earlier work this paper cites.
The moral roots of environmental attitudes
Matthew Feinberg and Robb Willer · 2013
Earlier work this paper cites.
Cultural differences in moral judgment and behavior, across and within societies
J. Graham, P. Meindl, E. Beall, Kate M. Johnson, and Li Zhang · 2015
Cited alongside, same era.
Concrete problems in ai safety
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané · 2016
Cited alongside, same era.
The ai alignment problem : why it is hard, and where to start
Eliezer Yudkowsky · 2016
Cited alongside, same era.
Emotions and reasoning in moral decision making
V. Nadurak · 2016
Cited alongside, same era.
Binding moral foundations and the narrowing of ideological conflict to the traditional morality domain
Ariel Malka, Danny Osborne, Christopher J Soto, Lara M Greaves, Chris G Sibley, and Yphtach Lelkes · 2016
Cited alongside, same era.
Nastiness, morality and religiosity in 33 nations
Lazar Stankov and Jihyun Lee · 2016
Superintelligence cannot be contained : Lessons from computability theory
Manuel Alfonseca, Manuel Cebrian, Antonio Fernandez Anta, Lorenzo Coviello, Andrés Abeliuk, and Iyad Rahwan · 2021
Later among the works it cites.
Ai in health and medicine
Pranav Rajpurkar, Emma Chen, Oishi Banerjee, and Eric J Topol · 2022
Later among the works it cites.
Mechanistic interpretability, variables, and the importance of interpretable bases
Chris Olah · 2022
Later among the works it cites.
In-context learning and induction heads
Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Scott Johnston, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah · 2022
Later among the works it cites.
Ai alignment : A comprehensive survey
Jiaming Ji, Tianyi Qiu, Boyuan Chen, Borong Zhang, Hantao Lou, Kaile Wang, Yawen Duan, Zhonghao He, Jiayi Zhou, Zhaowei Zhang, et al · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Rational and emotional sources of moral decision-making : an evolutionary-developmental account
Kaleda K. Denton and D. Krebs · 2017
Cited alongside, same era.
Insight - amazon scraps secret ai recruiting tool that showed bias against women
Jeffrey Dastin · 2018
Cited alongside, same era.
Ai alignment problem :“human values” don’t actually exist
Alexey Turchin · 2019
Cited alongside, same era.
Religion and related morality across cultures
V. Saroglou · 2019
Cited alongside, same era.
How computers see gender : An evaluation of gender classification in commercial facial analysis services
Morgan Klaus Scheuerman, Jacob M Paul, and Jed R Brubaker · 2019
Cited alongside, same era.
Artificial intelligence, values, and alignment
Iason Gabriel · 2020
Cited alongside, same era.
Cultural differences in moral judgement
Haotong Hong · 2023
Later among the works it cites.
Probing the moral development of large language models through defining issues test
K. Tanmay, Aditi Khandelwal, Utkarsh Agarwal, and M. Choudhury · 2023
Later among the works it cites.
The moral machine experiment on large language models
Kazuhiro Takemoto · 2023
Later among the works it cites.
Potential benefits of employing large language models in research in moral education and development
Hyemin Han · 2023
Later among the works it cites.
Linear representations of sentiment in large language models
Curt Tigges, Oskar John Hollinsworth, Atticus Geiger, and Neel Nanda · 2023
Later among the works it cites.
Activation addition : Steering language models without optimization
Alex Turner, Lisa Thiergart, David Udell, Gavin Leech, Ulisse Mini, and Monte MacDiarmid · 2023
Later among the works it cites.
Towards best practices of activation patching in language models : Metrics and methods
Fred Zhang and Neel Nanda · 2023
Later among the works it cites.
Emergent linear representations in world models of self-supervised sequence models
Neel Nanda, Andrew Lee, and Martin Wattenberg · 2023
Later among the works it cites.
Machine ethics : Do androids dream of being good people ?
Gonzalo Génova, Valentín Moreno, and M Rosario González · 2023
Later among the works it cites.
Navigating and reviewing ethical dilemmas in ai development : Strategies for transparency, fairness, and accountability
Olatunji Akinrinola, Chinwe Chinazo Okoye, Onyeka Chrisanctus Ofodile, and Chinonye Esther Ugochukwu · 2024
Closest in time.
A primer on the inner workings of transformer-based language models
Javier Ferrando, Gabriele Sarti, Arianna Bisazza, and Marta R Costa-jussà · 2024
Closest in time.
How to use and interpret activation patching
Stefan Heimersheim and Neel Nanda · 2024
Closest in time.