Fetching the paper…
Reading the bibliography…
Ensuring the moral reasoning capabilities of Large Language Models (LLMs) is a growing concern as these systems are used in socially sensitive tasks.
Étude comparative de la distribution florale dans une portion des alpes et des jura
Paul Jaccard. 1901 · 1901
Earlier work this paper cites.
Predicting the type and target of offensive posts in social media
Marcos Zampieri, Shervin Malmasi, Preslav Nakov, Sara Rosenthal, Noura Farra, and Ritesh Kumar. 2019 · 1902
Earlier work this paper cites.
Eraser: A benchmark to evaluate rationalized nlp models
Jay DeYoung, Sarthak Jain, Nazneen Fatema Rajani, Eric Lehman, Caiming Xiong, Richard Socher, and Byron C Wallace. 2019 · 1911
Earlier work this paper cites.
Horizontal and vertical dimensions of individualism and collectivism: A theoretical and measurement refinement
Theodore M Singelis, Harry C Triandis, Dharm PS Bhawuk, and Michele J Gelfand. 1995 · 1995
Earlier work this paper cites.
Nltk: The natural language toolkit
Edward Loper and Steven Bird. 2002 · 2002
Earlier work this paper cites.
The kappa statistic in reliability studies: Use, interpretation, and sample size requirements
Julius Sim and Chris C Wright. 2005 · 2005
Earlier work this paper cites.
When morality opposes justice: Conservatives have moral intuitions that liberals may not recognize
Jonathan Haidt and Jesse Graham. 2007 · 2007
Earlier work this paper cites.
Modeling annotators: A generative approach to learning from annotator rationales
Omar Zaidan and Jason Eisner. 2008 · 2008
Earlier work this paper cites.
Liberals and conservatives rely on different sets of moral foundations
Jesse Graham, Jonathan Haidt, and Brian A Nosek. 2009 · 2009
Earlier work this paper cites.
The dirty dozen: a concise measure of the dark triad
Peter K Jonason and Gregory D Webster. 2010 · 2010
Earlier work this paper cites.
Differences between tight and loose cultures: A 33-nation study
Michele J Gelfand, Jana L Raver, Lisa Nishii, Lisa M Leslie, Janetta Lun, Beng Chong Lim, Lili Duan, Assaf Almaliach, Soon Ang, Jakobina Arnadottir, et al. 2011 · 2011
Earlier work this paper cites.
Mapping the moral domain
Jesse Graham, Brian A Nosek, Jonathan Haidt, Ravi Iyer, Spassena Koleva, and Peter H Ditto. 2011 · 2011
Earlier work this paper cites.
The txm platform: Building open-source textual analysis software compatible with the tei encoding scheme
Serge Heiden. 2011 · 2011
Earlier work this paper cites.
“You don’t understand, this is a new war!” Analysis of hate speech in news web sites’ comments
Karmen Erjavec and Melita Poler Kovačič. 2012 · 2012
Earlier work this paper cites.
The righteous mind: Why good people are divided by politics and religion
Jonathan Haidt. 2012 · 2012
Earlier work this paper cites.
Interrater reliability: the kappa statistic
Mary L McHugh. 2012 · 2012
Earlier work this paper cites.
A solution to the mysteries of morality
Peter DeScioli and Robert Kurzban. 2013 · 2013
Earlier work this paper cites.
The 12 item social and economic conservatism scale (secs)
Jim AC Everett. 2013 · 2013
Earlier work this paper cites.
Moral foundations theory: The pragmatic validity of moral pluralism
Jesse Graham, Jonathan Haidt, Sena Koleva, Matt Motyl, Ravi Iyer, Sean P Wojcik, and Peter H Ditto. 2013 · 2013
Earlier work this paper cites.
Virtuous violence: Hurting and killing to create, sustain, end, and honor social relationships
Alan Page Fiske and Tage Shakti Rai. 2014 · 2014
Earlier work this paper cites.
Beyond point-and-shoot morality: Why cognitive (neuro) science matters for ethics
Joshua D Greene. 2014 · 2014
Earlier work this paper cites.
Measuring moral rhetoric in text
Eyal Sagi and Morteza Dehghani. 2014 · 2014
Earlier work this paper cites.
The nature of social dominance orientation: Theorizing and measuring preferences for intergroup inequality using the new sdo 7 scale
Arnold K Ho, Jim Sidanius, Nour Kteily, Jennifer Sheehy-Skeffington, Felicia Pratto, Kristin E Henkel, Rob Foels, and Andrew L Stewart. 2015 · 2015
Earlier work this paper cites.
The language of morality
Ian Keen. 2015 · 2015
Earlier work this paper cites.
Purity homophily in social networks
Morteza Dehghani, Kate Johnson, Joe Hoover, Eyal Sagi, Justin Garten, Niki Jitendra Parmar, Stephen Vaisey, Rumen Iliev, and Jesse Graham. 2016 · 2016
Earlier work this paper cites.
The psychology of morality: A review and analysis of empirical studies published from 1940 through 2017
Naomi Ellemers, Jojanneke Van Der Toorn, Yavor Paunov, and Thed Van Leeuwen. 2019 · 2017
Earlier work this paper cites.
Short and extra-short forms of the big five inventory–2: The bfi-2-s and bfi-2-xs
Christopher J Soto and Oliver P John. 2017 · 2017
Earlier work this paper cites.
A survey on automatic detection of hate speech in text
Paula Fortuna and Sérgio Nunes. 2018 · 2018
Earlier work this paper cites.
Classification of moral foundations in microblog political discourse
Kristen Johnson and Dan Goldwasser. 2018 · 2018
Earlier work this paper cites.
The cognitive and cultural foundations of moral behavior
Benjamin Grant Purzycki, Anne C Pisor, Coren Apicella, Quentin Atkinson, Emma Cohen, Joseph Henrich, Richard McElreath, Rita A McNamara, Ara Norenzayan, Aiyana K Willard, et al. 2018 · 2018
Earlier work this paper cites.
An italian twitter corpus of hate speech against immigrants
Manuela Sanguinetti, Fabio Poletto, Cristina Bosco, Viviana Patti, and Marco Stranisci. 2018 · 2018
Cited alongside, same era.
The theory of dyadic morality: Reinventing moral judgment by redefining harm
Chelsea Schein and Kurt Gray. 2018 · 2018
Cited alongside, same era.
Racial bias in hate speech and abusive language detection datasets
Thomas Davidson, Debasmita Bhattacharya, and Ingmar Weber. 2019 · 2019
Cited alongside, same era.
Kinship, cooperation, and the evolution of moral systems
Benjamin Enke. 2019 · 2019
Cited alongside, same era.
Moral foundations dictionary 2.0
JA Frimer, R Boghrati, J Haidt, J Graham, and M Dehghani. 2019 · 2019
Cited alongside, same era.
The mad model of moral contagion: The role of motivation, attention, and design in the spread of moralized content online
William J Brady, Molly J Crockett, and Jay J Van Bavel. 2020 · 2020
HateBR: A large expert annotated corpus of Brazilian Instagram comments for offensive language and hate speech detection
Francielle Vargas, Isabelle Carvalho, Fabiana Rodrigues de Góes, Thiago Pardo, and Fabrício Benevenuto. 2022 · 2022
Later among the works it cites.
Moral foundations of large language models
Marwa Abdulhai, Gregory Serapio-Garcia, Clément Crepy, Daria Valter, John Canny, and Natasha Jaques. 2023 · 2023
Later among the works it cites.
Moral narratives around the vaccination debate on facebook
Mariano Gastón Beiró, Jacopo D’Ignazi, Victoria Perez Bustos, María Florencia Prado, and Kyriaki Kalimeri. 2023 · 2023
Later among the works it cites.
Moral narratives around the vaccination debate on facebook
Mariano Gastón Beiró, Jacopo D’Ignazi, Victoria Perez Bustos, María Florencia Prado, and Kyriaki Kalimeri. 2023 · 2023
Later among the works it cites.
Hate speech classifiers learn normative social stereotypes
Aida Mostafazadeh Davani, Mohammad Atari, Brendan Kennedy, and Morteza Dehghani. 2023 · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
ERASER: A benchmark to evaluate rationalized NLP models
Jay DeYoung, Sarthak Jain, Nazneen Fatema Rajani, Eric Lehman, Caiming Xiong, Richard Socher, and Byron C. Wallace. 2020 · 2020
Cited alongside, same era.
Toxic, hateful, offensive or abusive? what are we really classifying? an empirical analysis of hate speech datasets
Paula Fortuna, Juan Soler, and Leo Wanner. 2020 · 2020
Cited alongside, same era.
Lara Grimminger and Roman Klinger. 2021 · 2020
Cited alongside, same era.
Moral foundations twitter corpus: A collection of 35k tweets annotated for moral sentiment
Joe Hoover, Gwenyth Portillo-Wightman, Leigh Yeh, Shreya Havaldar, Aida Mostafazadeh Davani, Ying Lin, Brendan Kennedy, Mohammad Atari, Zahra Kamel, Madelyn Mendlen, Gabriela Moreno, Christina Park, Tingyee E. Chang, Jenna Chin, Christian Leong, Jun Yen Leung, Arineh Mirinjian, and Morteza Dehghani. 2020 · 2020
Cited alongside, same era.
Learning to explain: Datasets and models for identifying valid reasoning chains in multihop question-answering
Harsh Jhamtani and Peter Clark. 2020 · 2020
Cited alongside, same era.
Contextualizing hate speech classifiers with post-hoc explanation
Brendan Kennedy, Xisen Jin, Aida Mostafazadeh Davani, Morteza Dehghani, and Xiang Ren. 2020 · 2020
Cited alongside, same era.
Later among the works it cites.
ViHOS: Hate speech spans detection for Vietnamese
Phu Gia Hoang, Canh Duc Luu, Khanh Quoc Tran, Kiet Van Nguyen, and Ngan Luu-Thuy Nguyen. 2023 · 2023
Later among the works it cites.
The (moral) language of hate
Brendan Kennedy, Preni Golazizian, Jackson Trager, Mohammad Atari, Joe Hoover, Aida Mostafazadeh Davani, and Morteza Dehghani. 2023 · 2023
Later among the works it cites.
Respectful or toxic? using zero-shot learning with language models to detect hate speech
Flor Miriam Plaza-del-Arco, Debora Nozza, and Dirk Hovy. 2023 · 2023
Later among the works it cites.
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al. 2023 · 2023
Later among the works it cites.
Rethinking machine ethics–can llms perform moral reasoning through the lens of moral theories?
Jingyan Zhou, Minda Hu, Junan Li, Xiaoying Zhang, Xixin Wu, Irwin King, and Helen Meng. 2023 · 2023
Later among the works it cites.
Perils and opportunities in using large language models in psychological research
Suhaib Abdurahman, Mohammad Atari, Farzan Karimi-Malekabadi, Mona J Xue, Jackson Trager, Peter S Park, Preni Golazizian, Ali Omrani, and Morteza Dehghani. 2024 · 2024
Later among the works it cites.
Ethical reasoning and moral value alignment of LLMs depend on the language we prompt them in
Utkarsh Agarwal, Kumar Tanmay, Aditi Khandelwal, and Monojit Choudhury. 2024 · 2024
Later among the works it cites.
Lorenzo Lupo, Paul Bose, Mahyar Habibi, Dirk Hovy, and Carlo Schwarz. 2024 · 2024
Later among the works it cites.
EX-FEVER: A dataset for multi-hop explainable fact verification
Huanhuan Ma, Weizhi Xu, Yifan Wei, Liuji Chen, Liang Wang, Qiang Liu, Shu Wu, and Liang Wang. 2024 · 2024
Later among the works it cites.
Annotator in the loop: A case study of in-depth rater engagement to create a prosocial benchmark dataset
Sonja Schmer-Galunder, Ruta Wheelock, Zaria Jalan, Alyssa Chvasta, Scott Friedman, and Emily Saltz. 2024 · 2024
Later among the works it cites.
Context-aware and expert data resources for brazilian portuguese hate speech detection
Francielle Vargas, Isabelle Carvalho, Thiago A. S. Pardo, and Fabrício Benevenuto. 2024 · 2024
Later among the works it cites.
Replacing judges with juries: Evaluating llm generations with a panel of diverse models
Pat Verga, Sebastian Hofstatter, Sophia Althammer, Yixuan Su, Aleksandra Piktus, Arkady Arkhangorodsky, Minjie Xu, Naomi White, and Patrick Lewis. 2024 · 2024
Later among the works it cites.
A Conceptual Analysis of the Overlaps and Differences between Hate Speech, Misinformation and Disinformation
Claire Wardle. 2024 · 2024
Later among the works it cites.
Targeting audiences’ moral values shapes misinformation sharing
Suhaib Abdurahman, Nils K Reimer, Preni Golazizian, Elisa Baek, Yixuan Shen, Jackson Trager, Roshni Lulla, Jonas Kaplan, Carolyn Parkinson, and Morteza Dehghani. 2025 · 2025
Closest in time.
Alessio Buscemi, Cédric Lothritz, Sergio Morales, Marcos Gomez-Vazquez, Robert Clarisó, Jordi Cabot, and German Castignani. 2025 · 2025
Closest in time.
Ai language model rivals expert ethicist in perceived moral expertise
Danica Dillion, Debanjan Mondal, Niket Tandon, and Kurt Gray. 2025 · 2025
Closest in time.
Values in the wild: Discovering and analyzing values in real-world language model interactions
Saffron Huang, Esin Durmus, Miles McCain, Kunal Handa, Alex Tamkin, Jerry Hong, Michael Stern, Arushi Somani, Xiuruo Zhang, and Deep Ganguli. 2025 · 2025
Closest in time.
Can llms assist annotators in identifying morality frames? - case study on vaccination debate on social media
Tunazzina Islam and Dan Goldwasser. 2025 · 2025
Closest in time.
Enhancing large language models with neurosymbolic reasoning for multilingual tasks
Sina Bagheri Nezhad and Ameeta Agrawal. 2025 · 2025
Closest in time.
Think like a person before responding: A multi-faceted evaluation of persona-guided LLMs for countering hate speech
Mikel Ngueajio, Flor Miriam Plaza-del Arco, Yi-Ling Chung, Danda Rawat, and Amanda Cercas Curry. 2025 · 2025
Closest in time.
Robustness of large language models in moral judgements
Soyoung Oh and Vera Demberg. 2025 · 2025
Closest in time.
Towards efficient and explainable hate speech detection via model distillation
Paloma Piot and Javier Parapar. 2025 · 2025
Closest in time.
HateBRXplain: A benchmark dataset with human-annotated rationales for explainable hate speech detection in Brazilian Portuguese
Isadora Salles, Francielle Vargas, and Fabrício Benevenuto. 2025 · 2025
Closest in time.
MEEP: Is this engaging? prompting large language models for dialogue evaluation in multilingual settings
Amila Ferron, Amber Shore, Ekata Mitra, and Ameeta Agrawal. 2023 · 2078
Closest in time.