Fetching the paper…
Reading the bibliography…
In this position paper, we argue that instead of morally aligning LLMs to specific set of ethical principles, we should infuse generic ethical reasoning capabilities into them so that they can handle value pluralism at a global scale.
Freedom and Reason
R. M. Hare. 1965 · 1965
Earlier work this paper cites.
Kant: Lectures on Ethics
Immanuel Kant. 1977 · 1977
Earlier work this paper cites.
Moral dilemmas and consistency in ethics
Terrance C McConnell. 1978 · 1978
Earlier work this paper cites.
The Philosophy of Moral Development: Moral Stages and the Idea of Justice
L. Kohlberg. 1981 · 1981
Earlier work this paper cites.
Utilitarianism, moral dilemmas, and moral cost
Michael Slote. 1985 · 1985
Earlier work this paper cites.
Ethical consistency
Bernard Williams. 1988 · 1988
Earlier work this paper cites.
The metaphysics of morals
Immanuel Kant. 1996 · 1996
Earlier work this paper cites.
Filial piety. a cross-cultural comparison and its implications for the well-being of older parents
Y T Dai and M F Dimond. 1998 · 1998
Earlier work this paper cites.
The Right and the Good
David Ross and Philip Stratton-Lake. 2002 · 2002
Earlier work this paper cites.
The wvs cultural map of the world
Ronald Inglehart and Chris Welzel. 2010 · 2010
Earlier work this paper cites.
Hate me, hate me not: Hate speech detection on facebook
Fabio Del Vigna, Andrea Cimino, Felice Dell’Orletta, Marinella Petrocchi, and Maurizio Tesconi. 2017 · 2017
Earlier work this paper cites.
Evaluating the underlying gender bias in contextualized word embeddings
Christine Basta, Marta R. Costa-jussà, and Noe Casas. 2019 · 2019
Earlier work this paper cites.
Machine behaviour
Iyad Rahwan, Manuel Cebrian, Nick Obradovich, Josh Bongard, Jean-François Bonnefon, Cynthia Breazeal, Jacob W. Crandall, Nicholas A. Christakis, Iain D. Couzin, Matthew O. Jackson, Nicholas R. Jennings, Ece Kamar, Isabel M. Kloumann, Hugo Larochelle, David Lazer, Richard McElreath, Alan Mislove, David C. Parkes, Alex ‘Sandy’ Pentland, and Margaret E. Roberts. 2019 · 2019
Earlier work this paper cites.
Tweeteval: Unified benchmark and comparative evaluation for tweet classification
Francesco Barbieri, Jose Camacho-Collados, Leonardo Neves, and Luis Espinosa-Anke. 2020 · 2020
Earlier work this paper cites.
Language (technology) is power: A critical survey of “bias” in nlp
Su Lin Blodgett, Solon Barocas, Hal Daumé III, and Hanna Wallach. 2020 · 2020
Earlier work this paper cites.
Hatebert: Retraining bert for abusive language detection in english
Tommaso Caselli, Valerio Basile, Jelena Mitrović, and Michael Granitzer. 2020 · 2020
Earlier work this paper cites.
Social chemistry 101: Learning to reason about social and moral norms
Maxwell Forbes, Jena D. Hwang, Vered Shwartz, Maarten Sap, and Yejin Choi. 2020 · 2020
Earlier work this paper cites.
Toxic, hateful, offensive or abusive? what are we really classifying? an empirical analysis of hate speech datasets
Paula Fortuna, Juan Soler, and Leo Wanner. 2020 · 2020
Cited alongside, same era.
RealToxicityPrompts: Evaluating neural toxic degeneration in language models
Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A. Smith. 2020 · 2020
Cited alongside, same era.
Hatexplain: A benchmark dataset for explainable hate speech detection
Binny Mathew, Punyajoy Saha, Seid Muhie Yimam, Chris Biemann, Pawan Goyal, and Animesh Mukherjee. 2020 · 2020
Cited alongside, same era.
Social bias frames: Reasoning about social and power implications of language
Maarten Sap, Saadia Gabriel, Lianhui Qin, Dan Jurafsky, Noah A. Smith, and Yejin Choi. 2020 · 2020
Cited alongside, same era.
Deontological Ethics
Larry Alexander and Michael Moore. 2021 · 2021
Cited alongside, same era.
Virtue Ethics
Rosalind Hursthouse and Glen Pettigrove. 2022 · 2022
Later among the works it cites.
Do multilingual language models capture differing moral norms?
Katharina Hämmerl, Björn Deiseroth, Patrick Schramowski, Jindřich Libovický, Alexander Fraser, and Kristian Kersting. 2022 · 2022
Later among the works it cites.
TruthfulQA: Measuring how models mimic human falsehoods
Stephanie Lin, Jacob Hilton, and Owain Evans. 2022 · 2022
Later among the works it cites.
Aligning generative language models with human values
Ruibo Liu, Ge Zhang, Xinyu Feng, and Soroush Vosoughi. 2022 · 2022
Later among the works it cites.
Ignore previous prompt: Attack techniques for language models
Fábio Perez and Ian Ribeiro. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021 · 2021
Cited alongside, same era.
Stereotyping Norwegian salmon: An inventory of pitfalls in fairness benchmark datasets
Su Lin Blodgett, Gilsinia Lopez, Alexandra Olteanu, Robert Sim, and Hanna Wallach. 2021 · 2021
Cited alongside, same era.
How linguistically fair are multilingual pre-trained language models?
Monojit Choudhury and Amit Deshpande. 2021 · 2021
Cited alongside, same era.
Moral stories: Situated reasoning about norms, intents, actions, and their consequences
Denis Emelin, Ronan Le Bras, Jena D. Hwang, Maxwell Forbes, and Yejin Choi. 2021 · 2021
Cited alongside, same era.
Suicidal ideation detection: A review of machine learning methods and applications
Shaoxiong Ji, Shirui Pan, Xue Li, Erik Cambria, Guodong Long, and Zi Huang. 2021 · 2021
Cited alongside, same era.
Delphi: Towards machine ethics and norms
Liwei Jiang, Jena D. Hwang, Chandra Bhagavatula, Ronan Le Bras, Maxwell Forbes, Jon Borchardt, Jenny Liang, Oren Etzioni, Maarten Sap, and Yejin Choi. 2021 · 2021
Cited alongside, same era.
“that courage to encourage”: Participation and aspirations in chat-based peer support for youth living with hiv
Naveena Karusala, David Odhiambo Seeh, Cyrus Mugo, Brandon Guthrie, Megan A Moreno, Grace John-Stewart, Irene Inwani, Richard Anderson, and Keshet Ronen. 2021 · 2021
Cited alongside, same era.
John Schulman, Barret Zoph, Christina Kim, Jacob Hilton, Jacob Menick, Jiayi Weng, Juan Felipe Ceron Uribe, Liam Fedus, Luke Metz, Michael Pokorny, Rapha Gontijo Lopes, Shengjia Zhao, Arun Vijayvergiya, Eric Sigler, Adam Perelman, Chelsea Voss, Mike Heaton, Joel Parish, Dave Cummings, Rajeev Nayak, Valerie Balcom, David Schnurr, Tomer Kaftan, Chris Hallacy, Nicholas Turley, Noah Deutsch, Vik Goel, Jonathan Ward, Aris Konstantinidis, Wojciech Zaremba, Long Ouyang, Leonard Bogdonoff, Joshua Gross, David Medina, Sarah Yoo, Teddy Lee, Ryan Lowe, Dan Mossing, Joost Huizinga, Roger Jiang, Carroll Wainwright, Diogo Almeida, Steph Lin, Marvin Zhang, Kai Xiao, Katarina Slama, Steven Bills, Alex Gray, Jan Leike, Jakub Pachocki, Phil Tillet, Shantanu Jain, Greg Brockman, and Nick Ryder. 2022 · 2022
Later among the works it cites.
Consequentialism
Walter Sinnott-Armstrong. 2022 · 2022
Later among the works it cites.
On the machine learning of ethical judgments from natural language
Zeerak Talat, Hagen Blix, Josef Valvoda, Maya Indira Ganesh, Ryan Cotterell, and Adina Williams. 2022 · 2022
Later among the works it cites.
Chatgpt on whatsapp: Govt’s bhashini initiative to use ai for beneficiaries of welfare schemes
Soumyarendra Barik. 2023 · 2023
Closest in time.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. 2023 · 2023
Closest in time.
The economic potential of generative AI: The next productivity frontier
Michael Chui, Eric Hazan, Robert Rogers, Alex Singla, Kate Smaje, Alex Sukharevsky, Lareina Yee, and Rodney Zemmerl. 2023 · 2023
Closest in time.
Aligning ai with shared human values
Dan Hendrycks, Collin Burns, Steven Basart, Andrew Critch, Jerry Li, Dawn Song, and Jacob Steinhardt. 2023 · 2023
Closest in time.
Fairness in language models beyond English: Gaps and challenges
Krithika Ramesh, Sunayana Sitaram, and Monojit Choudhury. 2023 · 2023
Closest in time.
Stanford alpaca: An instruction-following llama model
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023 · 2023
Closest in time.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023 · 2023
Closest in time.
Large language models are not fair evaluators
Peiyi Wang, Lei Li, Liang Chen, Dawei Zhu, Binghuai Lin, Yunbo Cao, Qi Liu, Tianyu Liu, and Zhifang Sui. 2023 · 2023
Closest in time.