Fetching the paper…
Reading the bibliography…
The rapid adoption of large language models (LLMs) has spurred extensive research into their encoded moral norms and decision-making processes.
Aligning AI With Shared Human Values
Dan Hendrycks, Collin Burns, Steven Basart, Andrew Critch, Jerry Li, Dawn Song, and Jacob Steinhardt. 2023 · 2008
Earlier work this paper cites.
Computing Krippendorff’s alpha-reliability
Klaus Krippendorff. 2011 · 2011
Earlier work this paper cites.
Moral foundations theory: The pragmatic validity of moral pluralism
Jesse Graham, Jonathan Haidt, Sena Koleva, Matt Motyl, Ravi Iyer, Sean P Wojcik, and Peter H Ditto. 2013 · 2013
Earlier work this paper cites.
Moral foundations vignettes: A standardized stimulus database of scenarios based on moral foundations theory
Scott Clifford, Vijeth Iyengar, Roberto Cabeza, and Walter Sinnott-Armstrong. 2015 · 2015
Earlier work this paper cites.
Purity homophily in social networks
Morteza Dehghani, Kate Johnson, Joe Hoover, Eyal Sagi, Justin Garten, Niki Jitendra Parmar, Stephen Vaisey, Rumen Iliev, and Jesse Graham. 2016 · 2016
Earlier work this paper cites.
The AI alignment problem: why it is hard, and where to start
Eliezer Yudkowsky. 2016 · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei. 2017 · 2017
Earlier work this paper cites.
Fast, free, and targeted: Reddit as a source for recruiting participants online
Itamar Shatz. 2017 · 2017
Earlier work this paper cites.
The MAD model of moral contagion: The role of motivation, attention, and design in the spread of moralized content online
William J Brady, M J Crockett, and Jay J Van Bavel. 2020 · 2020
Earlier work this paper cites.
Artificial intelligence, values, and alignment
Iason Gabriel. 2020 · 2020
Earlier work this paper cites.
Artificial intelligence and communication: A Human–Machine Communication research agenda
Andrea L Guzman and Seth C Lewis. 2020 · 2020
Earlier work this paper cites.
Constitutional AI: Harmlessness from AI Feedback
Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, Carol Chen, Catherine Olsson, Christopher Olah, Danny Hernandez, Dawn Drain, Deep Ganguli, Dustin Li, Eli Tran-Johnson, Ethan Perez, Jamie Kerr, Jared Mueller, Jeffrey Ladish, Joshua Landau, Kamal Ndousse, Kamile Lukosuite, Liane Lovitt, Michael Sellitto, Nelson Elhage, Nicholas Schiefer, Noemi Mercado, Nova DasSarma, Robert Lasenby, Robin Larson, Sam Ringer, Scott Johnston, Shauna Kravec, Sheer El Showk, Stanislav Fort, Tamera Lanham, Timothy Telleen-Lawton, Tom Conerly, Tom Henighan, Tristan Hume, Samuel R. Bowman, Zac Hatfield-Dodds, Ben Mann, Dario Amodei, Nicholas Joseph, Sam McCandlish, Tom Brown, and Jared Kaplan. 2022 · 2022
Earlier work this paper cites.
Social norms on reddit: A demographic analysis. In Proceedings of the 14th ACM Web Science Conference 2022 . 139–147
Sara De Candia, Gianmarco De Francisci Morales, Corrado Monti, and Francesco Bonchi. 2022 · 2022
Earlier work this paper cites.
Does Moral Code Have a Moral Code? Probing Delphi’s Moral Philosophy
Kathleen C. Fraser, Svetlana Kiritchenko, and Esma Balkir. 2022 · 2022
Earlier work this paper cites.
The Ghost in the Machine has an American accent: value conflict in GPT-3
Rebecca L Johnson, Giada Pistilli, Natalia Menédez-González, Leslye Denisse Dias Duran, Enrico Panai, Julija Kalpokiene, and Donald Jay Bertulfo. 2022 · 2022
Earlier work this paper cites.
Mapping topics in 100,000 real-life moral dilemmas. In Proceedings of the International AAAI Conference on Web and Social Media , Vol. 16. 699–710
Tuan Dung Nguyen, Georgiana Lyall, Alasdair Tran, Minjeong Shin, Nicholas George Carroll, Colin Klein, and Lexing Xie. 2022 · 2022
Earlier work this paper cites.
The Moral Foundations Reddit Corpus
Jackson Trager, Alireza S. Ziabari, Aida Mostafazadeh Davani, Preni Golazizian, Farzan Karimi-Malekabadi, Ali Omrani, Zhihe Li, Brendan Kennedy, Nils Karl Reimer, Melissa Reyes, Kelsey Cheng, Mellow Wei, Christina Merrifield, Arta Khosravi, Evans Alvarez, and Morteza Dehghani. 2022 · 2022
Earlier work this paper cites.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Earlier work this paper cites.
PaLM 2 Technical Report
Rohan Anil, Andrew M. Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Z. Chen, Eric Chu, J. Clark, Laurent El Shafey, Yanping Huang, Kathleen S. Meier-Hellstern, Gaurav Mishra, Erica Moreira, Mark Omernick, Kevin Robinson, Sebastian Ruder, Yi Tay, Kefan Xiao, Yuanzhong Xu, Yujing Zhang, Gustavo Hernández Abrego, Junwhan Ahn, Jacob Austin, Paul Barham, Jan A. Botha, James Bradbury, Siddhartha Brahma, Kevin Michael Brooks, Michele Catasta, Yongzhou Cheng, Colin Cherry, Christopher A. Choquette-Choo, Aakanksha Chowdhery, C Crépy, Shachi Dave, Mostafa Dehghani, Sunipa Dev, Jacob Devlin, M. C. D’iaz, Nan Du, Ethan Dyer, Vladimir Feinberg, Fan Feng, Vlad Fienber, Markus Freitag, Xavier García, Sebastian Gehrmann, Lucas González, Guy Gur-Ari, Steven Hand, Hadi Hashemi, Le Hou, Joshua Howland, An Ren Hu, Jeffrey Hui, Jeremy Hurwitz, Michael Isard, Abe Ittycheriah, Matthew Jagielski, Wen Hao Jia, Kathleen Kenealy, Maxim Krikun, Sneha Kudugunta, Chang Lan, Katherine Lee, Benjamin Lee, Eric Li, Mu-Li Li, Wei Li, Yaguang Li, Jun Yu Li, Hyeontaek Lim, Han Lin, Zhong-Zhong Liu, Frederick Liu, Marcello Maggioni, Aroma Mahendru, Joshua Maynez, Vedant Misra, Maysam Moussalem, Zachary Nado, John Nham, Eric Ni, Andrew Nystrom, Alicia Parrish, Marie Pellat, Martin Polacek, Oleksandr Polozov, Reiner Pope, Siyuan Qiao, Emily Reif, Bryan Richter, Parker Riley, Alexandra Ros, Aurko Roy, Brennan Saeta, Rajkumar Samuel, Renee Marie Shelby, Ambrose Slone, Daniel Smilkov, David R. So, Daniela Sohn, Simon Tokumine, Dasha Valter, Vijay Vasudevan, Kiran Vodrahalli, Xuezhi Wang, Pidong Wang, Zirui Wang, Tao Wang, John Wieting, Yuhuai Wu, Ke Xu, Yunhan Xu, Lin Wu Xue, Pengcheng Yin, Jiahui Yu, Qiaoling Zhang, Steven Zheng, Ce Zheng, Wei Zhou, Denny Zhou, Slav Petrov, and Yonghui Wu. 2023 · 2023
Earlier work this paper cites.
Assessing LLMs for Moral Value Pluralism
Noam Benkler, Drisana Mosaphir, Scott Friedman, Andrew Smart, and Sonja Schmer-Galunder. 2023 · 2023
Earlier work this paper cites.
Assessing Cross-Cultural Alignment between ChatGPT and Human Societies: An Empirical Study
Yong Cao, Li Zhou, Seolhwa Lee, Laura Cabello, Min Chen, and Daniel Hershcovich. 2023 · 2023
Earlier work this paper cites.
Psychology’s Weird Problems
Guilherme Sanches de Oliveira and Edward Baggs. 2023 · 2023
Cited alongside, same era.
Ronald Fischer, Markus Luczak-Roesch, and Johannes A. Karl. 2023 · 2023
Cited alongside, same era.
Revisiting the political biases of ChatGPT
Sasuke Fujimoto and Kazuhiro Takemoto. 2023 · 2023
Cited alongside, same era.
Author as character and narrator: Deconstructing personal narratives from the r/amitheasshole reddit community. In Proceedings of the International AAAI Conference on Web and Social Media , Vol. 17. 233–244
Salvatore Giorgi, Ke Zhao, Alexander H Feng, and Lara J Martin. 2023 · 2023
Cited alongside, same era.
The political ideology of conversational AI: Converging evidence on ChatGPT’s pro-environmental, left-libertarian orientation
Jochen Hartmann, Jasper Schwenzow, and Maximilian Witte. 2023 · 2023
DailyDilemmas: Revealing Value Preferences of LLMs with Quandaries of Daily Life
Yu Ying Chiu, Liwei Jiang, and Yejin Choi. 2024 · 2024
Later among the works it cites.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al · 2024
Later among the works it cites.
The Moral Turing Test: Evaluating Human-LLM Alignment in Moral Decision-Making
Basile Garcia, Crystal Qian, and Stefano Palminteri. 2024 · 2024
Later among the works it cites.
Whose Emotions and Moral Sentiments Do Language Models Reflect?
Zihao He, Siyi Guo, Ashwin Rao, and Kristina Lerman. 2024 · 2024
Later among the works it cites.
Large language models in mental health care: a scoping review
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al · 2023
Cited alongside, same era.
Understanding the benefits and challenges of using large language model-based conversational agents for mental well-being support. In AMIA Annual Symposium Proceedings , Vol. 2023. 1105
Zilin Ma, Yiyang Mei, and Zhaoyuan Su. 2024a · 2023
Cited alongside, same era.
Are Large Language Models Moral Hypocrites? A Study Based on Moral Foundations
José Luiz Nunes, Guilherme F. C. F. Almeida, Marcelo de Araujo, and Simone D. J. Barbosa. 2023 · 2023
Cited alongside, same era.
Knowledge of cultural moral norms in large language models. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki (Eds.). Association for Computational Linguistics, Toronto, Canada, 428–446
Aida Ramezani and Yang Xu. 2023 · 2023
Cited alongside, same era.
The Self-Perception and Political Biases of ChatGPT
Jérôme Rutinowski, Sven Franke, Jan Endendyk, Ina Dormuth, and Markus Pauly. 2023 · 2023
Cited alongside, same era.
Whose Opinions Do Language Models Reflect?
Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee, Percy Liang, and Tatsunori Hashimoto. 2023 · 2023
Cited alongside, same era.
Evaluating the Moral Beliefs Encoded in LLMs
Nino Scherrer, Claudia Shi, Amir Feder, and David M. Blei. 2023 · 2023
Cited alongside, same era.
Yining Hua, Fenglin Liu, Kailai Yang, Zehan Li, Hongbin Na, Yi-han Sheu, Peilin Zhou, Lauren V Moran, Sophia Ananiadou, Andrew Beam, et al · 2024
Later among the works it cites.
MoralBench: Moral Evaluation of LLMs
Jianchao Ji, Yutong Chen, Mingyu Jin, Wujiang Xu, Wenyue Hua, and Yongfeng Zhang. 2024 · 2024
Later among the works it cites.
A systematic survey and critical review on evaluating large language models: Challenges, limitations, and recommendations. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing . 13785–13816
Md Tahmid Rahman Laskar, Sawsan Alqahtani, M Saiful Bari, Mizanur Rahman, Mohammad Abdullah Matin Khan, Haidar Khan, Israt Jahan, Amran Bhuiyan, Chee Wei Tan, Md Rizwan Parvez, et al · 2024
Later among the works it cites.
The opportunities and risks of large language models in mental health
Hannah R Lawrence, Renee A Schneider, Susan B Rubin, Maja J Matarić, Daniel J McDuff, and Megan Jones Bell. 2024 · 2024
Later among the works it cites.
The Potential and Challenges of Evaluating Attitudes, Opinions, and Values in Large Language Models. In Findings of the Association for Computational Linguistics: EMNLP 2024 , Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (Eds.). Association for Computational Linguistics, Miami, Florida, USA, 8783–8805
Bolei Ma, Xinpeng Wang, Tiancheng Hu, Anna-Carolina Haensch, Michael A. Hedderich, Barbara Plank, and Frauke Kreuter. 2024b · 2024
Later among the works it cites.
More human than human: measuring ChatGPT political bias
Fabio Motoki, Valdemar Pinho Neto, and Victor Rodrigues. 2024 · 2024
Later among the works it cites.
A Comprehensive Survey of Bias in LLMs: Current Landscape and Future Directions
Rajesh Ranjan, Shailja Gupta, and Surya Narayan Singh. 2024 · 2024
Later among the works it cites.
Yuanyi Ren, Haoran Ye, Hanjun Fang, Xin Zhang, and Guojie Song. 2024 · 2024
Later among the works it cites.
The Political Preferences of LLMs
David Rozado. 2024 · 2024
Later among the works it cites.
Paul Röttger, Valentin Hofmann, Valentina Pyatkin, Musashi Hinck, Hannah Rose Kirk, Hinrich Schütze, and Dirk Hovy. 2024 · 2024
Later among the works it cites.
Melanie Sclar, Yejin Choi, Yulia Tsvetkov, and Alane Suhr. 2024 · 2024
Later among the works it cites.
Gemma: Open models based on gemini research and technology
Gemma Team, Thomas Mesnard, Cassidy Hardin, Robert Dadashi, Surya Bhupatiraju, Shreya Pathak, Laurent Sifre, Morgane Rivière, Mihir Sanjay Kale, Juliette Love, et al · 2024
Later among the works it cites.
A comprehensive survey of LLM alignment techniques: RLHF, RLAIF, PPO, DPO and more
Zhichao Wang, Bin Bi, Shiva Kumar Pentyala, Kiran Ramnath, Sougata Chaudhuri, Shubham Mehrotra, Xiang-Bo Mao, Sitaram Asur, et al · 2024
Later among the works it cites.
Hallucination is inevitable: An innate limitation of large language models
Ziwei Xu, Sanjay Jain, and Mohan Kankanhalli. 2024 · 2024
Later among the works it cites.
Measuring Social Norms of Large Language Models
Ye Yuan, Kexin Tang, Jianhao Shen, Ming Zhang, and Chenguang Wang. 2024 · 2024
Later among the works it cites.
Large Language Models Are Not Robust Multiple Choice Selectors
Chujie Zheng, Hao Zhou, Fandong Meng, Jie Zhou, and Minlie Huang. 2024 · 2024
Later among the works it cites.
praw-dev/praw
Bryce Boe, Joel Payne, PythonCoderAS, Joey Rees-Hill, Andreas Damgaard Pedersen, Ethan Dalool, Tim, Levi Roth, 13steinj, nmtake, Pyprohly, Julian Berman, Watchful1, kungming2, Alexander Putilin, MaybeNetwork, Liudvikam, Nemec, CrackedP0t, Jamie Magee, C.A.M. Gerlach, A. Baisero, Chad Birch, vallard192, bakonydraco, D0cR3d, Steve C, Aurélien Bombo, Zhifu Ge, and Michael Lazar. 2025 · 2025
Closest in time.