Fetching the paper…
Reading the bibliography…
As Large Language Models (LLMs) achieve remarkable breakthroughs, aligning their values with humans has become imperative for their responsible development and customized applications.
A value pluralism model of ideological reasoning
Philip E Tetlock. 1986 · 1986
Earlier work this paper cites.
Slinging arrows at democracy: Social choice theory, value pluralism, and democratic politics
Richard H Pildes and Elizabeth S Anderson. 1990 · 1990
Earlier work this paper cites.
Sus-a quick and dirty usability scale
John Brooke et al. 1996 · 1996
Earlier work this paper cites.
The big five personality factors: the psycholexical approach to personality
Boele De Raad. 2000 · 2000
Earlier work this paper cites.
The information bottleneck method
Naftali Tishby, Fernando C Pereira, and William Bialek. 2000 · 2000
Earlier work this paper cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020 · 2001
Earlier work this paper cites.
Value pluralism
Elinor Mason. 2006 · 2006
Earlier work this paper cites.
Incommensurable values
Nien-hê Hsieh and Henrik Andersson. 2007 · 2007
Earlier work this paper cites.
A suggested change in terminology and emphasis regarding validity and education
Robert W Lissitz and Karen Samuelsen. 2007 · 2007
Earlier work this paper cites.
Aligning ai with shared human values
Dan Hendrycks, Collin Burns, Steven Basart, Andrew Critch, Jerry Li, Dawn Song, and Jacob Steinhardt. 2020 · 2008
Earlier work this paper cites.
Realtoxicityprompts: Evaluating neural toxic degeneration in language models
Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A Smith. 2020 · 2009
Earlier work this paper cites.
Crows-pairs: A challenge dataset for measuring social biases in masked language models
Nikita Nangia, Clara Vania, Rasika Bhalerao, and Samuel R Bowman. 2020 · 2010
Earlier work this paper cites.
Social choice and individual values , volume 12
Kenneth J Arrow. 2012 · 2012
Earlier work this paper cites.
An overview of the schwartz theory of basic values
Shalom H Schwartz. 2012 · 2012
Earlier work this paper cites.
Moral foundations theory: The pragmatic validity of moral pluralism
Jesse Graham, Jonathan Haidt, Sena Koleva, Matt Motyl, Ravi Iyer, Sean P Wojcik, and Peter H Ditto. 2013 · 2013
Earlier work this paper cites.
The developmental origins of fairness: The knowledge–behavior gap
Peter R Blake, Katherine McAuliffe, and Felix Warneken. 2014 · 2014
Earlier work this paper cites.
Quantifying the user experience: Practical statistics for user research
Jeff Sauro and James R Lewis. 2016 · 2016
Earlier work this paper cites.
Predicting ideological prejudice
Mark J Brandt. 2017 · 2017
Earlier work this paper cites.
Personal values in human life
Lilach Sagiv, Sonia Roccas, Jan Cieciuch, and Shalom H Schwartz. 2017 · 2017
Earlier work this paper cites.
Welfare as equity equivalents
Loïc Berger and Johannes Emmerling. 2020 · 2020
Earlier work this paper cites.
Disentangling perceptions of offensiveness: Cultural and moral correlates
Aida Davani, Mark Díaz, Dylan Baker, and Vinodkumar Prabhakaran. 2024 · 2021
Earlier work this paper cites.
Can machines learn morality? the delphi experiment
Liwei Jiang, Jena D Hwang, Chandra Bhagavatula, Ronan Le Bras, Jenny Liang, Jesse Dodge, Keisuke Sakaguchi, Maxwell Forbes, Jon Borchardt, Saadia Gabriel, et al. 2021 · 2021
Earlier work this paper cites.
Human-versus artificial intelligence
JE (Hans) Korteling, Geertje C van de Boer-Visschedijk, Romy AM Blankendaal, Rob C Boonekamp, and A Roos Eikelboom. 2021 · 2021
Earlier work this paper cites.
Bbq: A hand-built bias benchmark for question answering
Alicia Parrish, Angelica Chen, Nikita Nangia, Vishakh Padmakumar, Jason Phang, Jana Thompson, Phu Mon Htut, and Samuel R Bowman. 2021 · 2021
Earlier work this paper cites.
Ethical and social risks of harm from language models
Laura Weidinger, John Mellor, Maribeth Rauh, Conor Griffin, Jonathan Uesato, Po-Sen Huang, Myra Cheng, Mia Glaese, Borja Balle, Atoosa Kasirzadeh, et al. 2021 · 2021
Earlier work this paper cites.
Does moral code have a moral code? probing delphi’s moral philosophy
Kathleen C Fraser, Svetlana Kiritchenko, and Esma Balkir. 2022 · 2022
Earlier work this paper cites.
Holistic evaluation of language models
Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, et al. 2022 · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022 · 2022
Cited alongside, same era.
Moral mimicry: Large language models produce moral rationalizations tailored to political identity
Gabriel Simmons. 2022 · 2022
Cited alongside, same era.
Self-instruct: Aligning language model with self generated instructions
Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A Smith, Daniel Khashabi, and Hannaneh Hajishirzi. 2022 · 2022
Cited alongside, same era.
A survey on evaluation of large language models
Yupeng Chang, Xu Wang, Jindong Wang, Yuan Wu, Linyi Yang, Kaijie Zhu, Hao Chen, Xiaoyuan Yi, Cunxiang Wang, Yidong Wang, et al. 2024 · 2024
Later among the works it cites.
Chatbot arena: An open platform for evaluating llms by human preference
Wei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos, Tianle Li, Dacheng Li, Banghua Zhu, Hao Zhang, Michael Jordan, Joseph E Gonzalez, et al. 2024 · 2024
Later among the works it cites.
Why do you answer like that? psychological analysis on underlying connections between llm’s values and safety risks
Sooyung Choi, Xiaoyuan Yi, Jing Yao, Xing Xie, and JinYeong Bak. 2024 · 2024
Later among the works it cites.
Yihong Dong, Xue Jiang, Huanyu Liu, Zhi Jin, Bin Gu, Mengfei Yang, and Ge Li. 2024 · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, et al. 2022 · 2022
Cited alongside, same era.
The moral integrity corpus: A benchmark for ethical dialogue systems
Caleb Ziems, Jane A Yu, Yi-Chia Wang, Alon Halevy, and Diyi Yang. 2022 · 2022
Cited alongside, same era.
Moral foundations of large language models
Marwa Abdulhai, Gregory Serapio-Garcia, Clément Crepy, Daria Valter, John Canny, and Natasha Jaques. 2023 · 2023
Cited alongside, same era.
Sparks of artificial general intelligence: Early experiments with gpt-4
Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, et al. 2023 · 2023
Cited alongside, same era.
Shitong Duan, Xiaoyuan Yi, Peng Zhang, Tun Lu, Xing Xie, and Ning Gu. 2023 · 2023
Cited alongside, same era.
Nphardeval: Dynamic benchmark on reasoning ability of large language models via complexity classes
Lizhou Fan, Wenyue Hua, Lingyao Li, Haoyang Ling, and Yongfeng Zhang. 2023 · 2023
Cited alongside, same era.
Trustgpt: A benchmark for trustworthy and responsible large language models
Yue Huang, Qihui Zhang, Lichao Sun, et al. 2023 · 2023
Cited alongside, same era.
Reem I Masoud, Ziquan Liu, Martin Ferianc, Philip Treleaven, and Miguel Rodrigues. 2023 · 2023
Cited alongside, same era.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. 2024 · 2024
Later among the works it cites.
Flames: Benchmarking value alignment of LLMs in Chinese
Kexin Huang, Xiangyang Liu, Qianyu Guo, Tianxiang Sun, Jiawei Sun, Yaru Wang, Zeyang Zhou, Yixu Wang, Yan Teng, Xipeng Qiu, Yingchun Wang, and Dahua Lin. 2024a · 2024
Later among the works it cites.
Flames: Benchmarking value alignment of llms in chinese
Kexin Huang, Xiangyang Liu, Qianyu Guo, Tianxiang Sun, Jiawei Sun, Yaru Wang, Zeyang Zhou, Yixu Wang, Yan Teng, Xipeng Qiu, et al. 2024b · 2024
Later among the works it cites.
Aaron Hurst, Adam Lerer, Adam P Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, et al. 2024 · 2024
Later among the works it cites.
Moralbench: Moral evaluation of llms
Jianchao Ji, Yutong Chen, Mingyu Jin, Wujiang Xu, Wenyue Hua, and Yongfeng Zhang. 2024 · 2024
Later among the works it cites.
Raising the bar: Investigating the values of large language models via generative evolving testing
Han Jiang, Xiaoyuan Yi, Zhihua Wei, Shu Wang, and Xing Xie. 2024 · 2024
Later among the works it cites.
Opencompass
Shanghai AI Lab. 2024 · 2024
Later among the works it cites.
Salad-bench: A hierarchical and comprehensive safety benchmark for large language models
Lijun Li, Bowen Dong, Ruohui Wang, Xuhao Hu, Wangmeng Zuo, Dahua Lin, Yu Qiao, and Jing Shao. 2024 · 2024
Later among the works it cites.
Inadequacies of large language model benchmarks in the era of generative artificial intelligence
Timothy R McIntosh, Teo Susnjak, Tong Liu, Paul Watters, and Malka N Halgamuge. 2024 · 2024
Later among the works it cites.
Gwenyth Isobel Meadows, Nicholas Wai Long Lau, Eva Adelina Susanto, Chi Lok Yu, and Aditya Paul. 2024 · 2024
Later among the works it cites.
Sg-bench: Evaluating llm safety generalization across diverse tasks and prompt types
Yutao Mou, Shikun Zhang, and Wei Ye. 2024 · 2024
Later among the works it cites.
Yuanyi Ren, Haoran Ye, Hanjun Fang, Xin Zhang, and Guojie Song. 2024 · 2024
Later among the works it cites.
Towards understanding sycophancy in language models
Mrinank Sharma, Meg Tong, Tomasz Korbak, David Duvenaud, Amanda Askell, Samuel R Bowman, Esin DURMUS, Zac Hatfield-Dodds, Scott R Johnston, Shauna M Kravec, et al. 2024 · 2024
Later among the works it cites.
Chatbot arena
UC Berkeley SkyLab and LMArena. 2024 · 2024
Later among the works it cites.
Trustllm: Trustworthiness in large language models
Lichao Sun, Yue Huang, Haoran Wang, Siyuan Wu, Qihui Zhang, Chujie Gao, Yixin Huang, Wenhan Lyu, Yixuan Zhang, Xiner Li, et al. 2024 · 2024
Later among the works it cites.
Clave: An adaptive framework for evaluating values of llm generated responses
Jing Yao, Xiaoyuan Yi, and Xing Xie. 2024 · 2024
Later among the works it cites.
Yifan Zeng. 2024 · 2024
Later among the works it cites.
Safetybench: Evaluating the safety of large language models
Zhexin Zhang, Leqi Lei, Lindong Wu, Rui Sun, Yongkang Huang, Chong Long, Xiao Liu, Xuanyu Lei, Jie Tang, and Minlie Huang. 2024 · 2024
Later among the works it cites.
Wenlong Zhao, Debanjan Mondal, Niket Tandon, Danica Dillion, Kurt Gray, and Yuling Gu. 2024 · 2024
Later among the works it cites.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et al. 2025 · 2025
Closest in time.
Cultural alignment in large language models: An explanatory analysis based on hofstede’s cultural dimensions
Reem Masoud, Ziquan Liu, Martin Ferianc, Philip C Treleaven, and Miguel Rodrigues Rodrigues. 2025 · 2025
Closest in time.
Aligning llms with individual preferences via interaction
Shujin Wu, Yi R Fung, Cheng Qian, Jeonghwan Kim, Dilek Hakkani-Tur, and Heng Ji. 2025 · 2025
Closest in time.