Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have made unprecedented breakthroughs, yet their increasing integration into everyday life might raise societal risks due to generated unethical content.
The psychology of social norms
Muzafer Sherif · 1936
Earlier work this paper cites.
Runaround. i, robot
Isaac Asimov · 1950
Earlier work this paper cites.
Stages of moral development
Lawrence Kohlberg · 1971
Earlier work this paper cites.
The cognitive-developmental approach to moral education
Lawrence Kohlberg · 1975
Earlier work this paper cites.
Social learning theory , volume 1
Albert Bandura and Richard H Walters · 1977
Earlier work this paper cites.
Perplexity—a measure of the difficulty of speech recognition tasks
Fred Jelinek, Robert L Mercer, Lalit R Bahl, and James K Baker · 1977
Earlier work this paper cites.
Moral development: A review of the theory
Lawrence Kohlberg and Richard H Hersh · 1977
Earlier work this paper cites.
Rule utilitarianism, rights, obligations and the theory of rational behavior
John C Harsanyi and John C Harsanyi · 1982
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
Simulated annealing
Dimitris Bertsimas and John Tsitsiklis · 1993
Earlier work this paper cites.
A view of the em algorithm that justifies incremental, sparse, and other variants
Radford M Neal and Geoffrey E Hinton · 1998
Earlier work this paper cites.
A theory of cultural values and some implications for work
Shalom H Schwartz et al · 1999
Earlier work this paper cites.
The construction of moral dilemmas in everyday life
Gillian R Wark and Dennis L Krebs · 2000
Earlier work this paper cites.
Common morality: Deciding what to do
Bernard Gert · 2004
Earlier work this paper cites.
Intuitive ethics: How innately prepared intuitions generate culturally variable virtues
Jonathan Haidt and Craig Joseph · 2004
Earlier work this paper cites.
The nature, importance, and difficulty of machine ethics
James H Moor · 2006
Earlier work this paper cites.
Basic human values: Theory, measurement, and applications
Shalom H Schwartz · 2007
Earlier work this paper cites.
Moral foundations questionnaire
Jesse Graham, Brian A Nosek, Jonathan Haidt, Ravi Iyer, Koleva Spassena, and Peter H Ditto · 2008
Earlier work this paper cites.
Should i save or should i not kill? how people solve moral dilemmas depends on which rule is most accessible
Ron Broeders, Kees Van Den Bos, Patrick A Müller, and Jaap Ham · 2011
Earlier work this paper cites.
Dimensionalizing cultures: The hofstede model in context
Geert Hofstede · 2011
Earlier work this paper cites.
An overview of the schwartz theory of basic values
Shalom H Schwartz · 2012
Earlier work this paper cites.
Moral foundations theory: The pragmatic validity of moral pluralism
Jesse Graham, Jonathan Haidt, Sena Koleva, Matt Motyl, Ravi Iyer, Sean P Wojcik, and Peter H Ditto · 2013
Earlier work this paper cites.
Modern psychometrics: The science of psychological assessment
John Rust and Susan Golombok · 2014
Earlier work this paper cites.
Impure or just weird? scenario sampling bias raises questions about the foundation of morality
Kurt Gray and Jonathan E Keeney · 2015
Earlier work this paper cites.
A handbook of test construction (psychology revivals): introduction to psychometric design
Paul Kline · 2015
Earlier work this paper cites.
Scalable bayesian optimization using deep neural networks
Jasper Snoek, Oren Rippel, Kevin Swersky, Ryan Kiros, Nadathur Satish, Narayanan Sundaram, Mostofa Patwary, Mr Prabhat, and Ryan Adams · 2015
Earlier work this paper cites.
A diversity-promoting objective function for neural conversation models
Jiwei Li, Michel Galley, Chris Brockett, Jianfeng Gao, and Bill Dolan · 2016
Earlier work this paper cites.
Validation of the moral foundations questionnaire in turkey and its relation to cultural schemas of individualism and collectivism
Onurcan Yilmaz, Mehmet Harma, Hasan G Bahçekapili, and Sevim Cesur · 2016
Earlier work this paper cites.
Guided open vocabulary image captioning with constrained beam search
Peter Anderson, Basura Fernando, Mark Johnson, and Stephen Gould · 2017
Earlier work this paper cites.
Applying moral foundations theory to understanding public views of sexual offending
Craig A Harper and Andrew J Harris · 2017
Earlier work this paper cites.
Intuitive ethics and political orientations: Testing moral foundations as a theory of political ideology
Kevin B Smith, John R Alford, John R Hibbing, Nicholas G Martin, and Peter K Hatemi · 2017
Earlier work this paper cites.
A tutorial on bayesian optimization
Peter I Frazier · 2018
Earlier work this paper cites.
Rankme: Reliable human ratings for natural language generation
Jekaterina Novikova, Ondřej Dušek, and Verena Rieser · 2018
Earlier work this paper cites.
Fast lexically constrained decoding with dynamic beam allocation for neural machine translation
Matt Post and David Vilar · 2018
Earlier work this paper cites.
Expanding the scope and content of morality policy research: lessons from moral foundations theory
Raymond Tatalovich and Dane G Wendell · 2018
Earlier work this paper cites.
Texygen: A benchmarking platform for text generation models
Yaoming Zhu, Sidi Lu, Lei Zheng, Jiaxian Guo, Weinan Zhang, Jun Wang, and Yong Yu · 2018
Cited alongside, same era.
The curious case of neural text degeneration
Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi · 2019
Cited alongside, same era.
Sex differences in moral judgements across 67 countries
Mohammad Atari, Mark HC Lai, and Morteza Dehghani · 2020
Cited alongside, same era.
Social chemistry 101: Learning to reason about social and moral norms
Maxwell Forbes, Jena D Hwang, Vered Shwartz, Maarten Sap, and Yejin Choi · 2020
Cited alongside, same era.
Artificial intelligence, values, and alignment
Iason Gabriel · 2020
Cited alongside, same era.
Realtoxicityprompts: Evaluating neural toxic degeneration in language models
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Later among the works it cites.
Cold decoding: Energy-based constrained text generation with langevin dynamics
Lianhui Qin, Sean Welleck, Daniel Khashabi, and Yejin Choi · 2022
Later among the works it cites.
Self-critiquing models for assisting human evaluators
William Saunders, Catherine Yeh, Jeff Wu, Steven Bills, Long Ouyang, Jonathan Ward, and Jan Leike · 2022
Later among the works it cites.
Moral mimicry: Large language models produce moral rationalizations tailored to political identity
Gabriel Simmons · 2022
Later among the works it cites.
Self-instruct: Aligning language model with self generated instructions
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A Smith · 2020
Cited alongside, same era.
Aligning ai with shared human values
Dan Hendrycks, Collin Burns, Steven Basart, Andrew Critch, Jerry Li, Dawn Song, and Jacob Steinhardt · 2020
Cited alongside, same era.
A cross-cultural examination on global orientations and moral foundations
Xiaomeng Hu, Yijie Zhu, Feng Yu, David A Wilder, Li Zhang, Sylvia Xiaohua Chen, and Kaiping Peng · 2020
Cited alongside, same era.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Cited alongside, same era.
Model inversion networks for model-based optimization
Aviral Kumar and Sergey Levine · 2020
Cited alongside, same era.
Towards controllable biases in language generation
Emily Sheng, Kai-Wei Chang, Prem Natarajan, and Nanyun Peng · 2020
Cited alongside, same era.
Sliced score matching: A scalable approach to density and score estimation
Yang Song, Sahaj Garg, Jiaxin Shi, and Stefano Ermon · 2020
Cited alongside, same era.
Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A Smith, Daniel Khashabi, and Hannaneh Hajishirzi · 2022
Later among the works it cites.
Large language models are human-level prompt engineers
Yongchao Zhou, Andrei Ioan Muresanu, Ziwen Han, Keiran Paster, Silviu Pitis, Harris Chan, and Jimmy Ba · 2022
Later among the works it cites.
The moral integrity corpus: A benchmark for ethical dialogue systems
Caleb Ziems, Jane A Yu, Yi-Chia Wang, Alon Halevy, and Diyi Yang · 2022
Later among the works it cites.
Probing pre-trained language models for cross-cultural differences in values
Arnav Arora, Lucie-Aimée Kaffee, and Isabelle Augenstein · 2023
Closest in time.
Enabling classifiers to make judgements explicitly aligned with human values
Yejin Bang, Tiezheng Yu, Andrea Madotto, Zhaojiang Lin, Mona Diab, and Pascale Fung · 2023
Closest in time.
Sparks of artificial general intelligence: Early experiments with gpt-4
Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, et al · 2023
Closest in time.
Assessing cross-cultural alignment between chatgpt and human societies: An empirical study
Yong Cao, Li Zhou, Seolhwa Lee, Laura Cabello, Min Chen, and Daniel Hershcovich · 2023
Closest in time.
Toxicity in chatgpt: Analyzing persona-assigned language models
Ameet Deshpande, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan, and Karthik Narasimhan · 2023
Closest in time.
The capacity for moral self-correction in large language models
Deep Ganguli, Amanda Askell, Nicholas Schiefer, Thomas Liao, Kamilė Lukošiūtė, Anna Chen, Anna Goldie, Azalia Mirhoseini, Catherine Olsson, Danny Hernandez, et al · 2023
Closest in time.
Time travel in llms: Tracing data contamination in large language models
Shahriar Golchin and Mihai Surdeanu · 2023
Closest in time.
Critic: Large language models can self-correct with tool-interactive critiquing
Zhibin Gou, Zhihong Shao, Yeyun Gong, Yelong Shen, Yujiu Yang, Nan Duan, and Weizhu Chen · 2023
Closest in time.
Evaluating the robustness of discrete prompts
Yoichi Ishibashi, Danushka Bollegala, Katsuhito Sudoh, and Satoshi Nakamura · 2023
Closest in time.
Beavertails: Towards improved safety alignment of llm via a human-preference dataset
Jiaming Ji, Mickel Liu, Juntao Dai, Xuehai Pan, Chi Zhang, Ce Bian, Ruiyang Sun, Yizhou Wang, and Yaodong Yang · 2023
Closest in time.
Hannah Rose Kirk, Bertie Vidgen, Paul Röttger, and Scott A Hale · 2023
Closest in time.
Biastestgpt: Using chatgpt for social bias testing of language models
Rafal Kocielnik, Shrimai Prabhumoye, Vivian Zhang, Roy Jiang, R. Michael Alvarez, and Anima Anandkumar · 2023
Closest in time.
Chatgpt: Jack of all trades, master of none
Jan Kocoń, Igor Cichecki, Oliwier Kaszyca, Mateusz Kochanek, Dominika Szydło, Joanna Baran, Julita Bielaniewicz, Marcin Gruza, Arkadiusz Janz, Kamil Kanclerz, et al · 2023
Closest in time.
A systematic study and comprehensive evaluation of ChatGPT on benchmark datasets
Md Tahmid Rahman Laskar, M Saiful Bari, Mizanur Rahman, Md Amran Hossen Bhuiyan, Shafiq Joty, and Jimmy Huang · 2023
Closest in time.
Multi-step jailbreaking privacy attacks on chatgpt
Haoran Li, Dadi Guo, Wei Fan, Mingshi Xu, and Yangqiu Song · 2023
Closest in time.
Training socially aligned language models in simulated human society
Ruibo Liu, Ruixin Yang, Chenyan Jia, Ge Zhang, Denny Zhou, Andrew M Dai, Diyi Yang, and Soroush Vosoughi · 2023
Closest in time.
The decontaminated evaluation of gpt-4, 2023
Benjamin Marie · 2023
Closest in time.
Inverse scaling: When bigger isn’t better
Ian R McKenzie, Alexander Lyzhov, Michael Pieler, Alicia Parrish, Aaron Mueller, Ameya Prabhu, Euan McLean, Aaron Kirtland, Alexis Ross, Alisa Liu, et al · 2023
Closest in time.
OpenAI · 2023
Closest in time.
Discovering language model behaviors with model-written evaluations
Ethan Perez, Sam Ringer, Kamile Lukosiute, Karina Nguyen, Edwin Chen, Scott Heiner, Craig Pettit, Catherine Olsson, Sandipan Kundu, Saurav Kadavath, Andy Jones, Anna Chen, Benjamin Mann, Brian Israel, Bryan Seethor, Cameron McKinnon, Christopher Olah, Da Yan, Daniela Amodei, Dario Amodei, Dawn Drain, Dustin Li, Eli Tran-Johnson, Guro Khundadze, Jackson Kernion, James Landis, Jamie Kerr, Jared Mueller, Jeeyoon Hyun, Joshua Landau, Kamal Ndousse, Landon Goldberg, Liane Lovitt, Martin Lucas, Michael Sellitto, Miranda Zhang, Neerav Kingsland, Nelson Elhage, Nicholas Joseph, Noemi Mercado, Nova DasSarma, Oliver Rausch, Robin Larson, Sam McCandlish, Scott Johnston, Shauna Kravec, Sheer El Showk, Tamera Lanham, Timothy Telleen-Lawton, Tom Brown, Tom Henighan, Tristan Hume, Yuntao Bai, Zac Hatfield-Dodds, Jack Clark, Samuel R. Bowman, Amanda Askell, Roger Grosse, Danny Hernandez, Deep Ganguli, Evan Hubinger, Nicholas Schiefer, and Jared Kaplan · 2023
Closest in time.
Evaluating the moral beliefs encoded in llms
Nino Scherrer, Claudia Shi, Amir Feder, and David M Blei · 2023
Closest in time.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Closest in time.
Miles Turpin, Julian Michael, Ethan Perez, and Samuel R Bowman · 2023
Closest in time.
Fake alignment: Are llms really aligned well?
Yixu Wang, Yan Teng, Kexin Huang, Chengqi Lyu, Songyang Zhang, Wenwei Zhang, Xingjun Ma, and Yingchun Wang · 2023
Closest in time.
Simple synthetic data reduces sycophancy in large language models
Jerry Wei, Da Huang, Yifeng Lu, Denny Zhou, and Quoc V Le · 2023
Closest in time.
Rrhf: Rank responses to align language models with human feedback without tears
Zheng Yuan, Hongyi Yuan, Chuanqi Tan, Wei Wang, Songfang Huang, and Fei Huang · 2023
Closest in time.
Exploring ai ethics of chatgpt: A diagnostic analysis
Terry Yue Zhuo, Yujin Huang, Chunyang Chen, and Zhenchang Xing · 2023
Closest in time.