Fetching the paper…
Reading the bibliography…
Existing research primarily evaluates the values of LLMs by examining their stated inclinations towards specific values.
Eraser: A benchmark to evaluate rationalized nlp models
Jay DeYoung, Sarthak Jain, Nazneen Fatema Rajani, Eric Lehman, Caiming Xiong, Richard Socher, and Byron C Wallace. 2019 · 1911
Earlier work this paper cites.
Understanding attitudes and predictiing social behavior
Icek Ajzen. 1980 · 1980
Earlier work this paper cites.
Predicting and understanding consumer behavior: Attitude-behavior correspondence
Martin Fishbein and Icek Ajzen. 1980 · 1980
Earlier work this paper cites.
Universals in the content and structure of values: Theoretical advances and empirical tests in 20 countries
Shalom H Schwartz. 1992 · 1992
Earlier work this paper cites.
Are there universal aspects in the structure and contents of human values?
Shalom H Schwartz. 1994 · 1994
Earlier work this paper cites.
Overcoming the ‘value-action gap’in environmental policy: Tensions between national policy and local experience
James Blake. 1999 · 1999
Earlier work this paper cites.
Environmental attitude and ecological behaviour
Florian G Kaiser, Sybille Wölfing, and Urs Fuhrer. 1999 · 1999
Earlier work this paper cites.
Bridging the intention–behaviour gap: The role of moral norm
Gaston Godin, Mark Conner, and Paschal Sheeran. 2005 · 2005
Earlier work this paper cites.
Robustness and fruitfulness of a theory of universals in individual values
Shalom H Schwartz. 2005 · 2005
Earlier work this paper cites.
Impact of values, involvement and perceptions on consumer attitudes and intentions towards sustainable consumption
Iris Vermeir and Wim Verbeke. 2006 · 2006
Earlier work this paper cites.
The value-action gap in waste recycling: The case of undergraduates in hong kong
Shan-Shan Chung and Monica Miu-Yin Leung. 2007 · 2007
Earlier work this paper cites.
An overview of the schwartz theory of basic values
Shalom H Schwartz. 2012 · 2012
Earlier work this paper cites.
Heterogeneous value alignment evaluation for large language models
Zhaowei Zhang, Ceyao Zhang, Nian Liu, Siyuan Qi, Ziqi Rong, Song-Chun Zhu, and Yaodong Yang. 2025 · 2012
Earlier work this paper cites.
Prolific
Prolific. 2024 · 2014
Earlier work this paper cites.
General social survey
Public-Use Microdata File. 2017 · 2017
Earlier work this paper cites.
Artificial intelligence, values, and alignment
Iason Gabriel. 2020 · 2020
Earlier work this paper cites.
World values survey: Round seven-country-pooled datafile. madrid, spain & vienna, austria: Jd systems institute & wvsa secretariat
Christian Haerpfer, Ronald Inglehart, Alejandro Moreno, Christian Welzel, K Kizilova, Jaime Diez-Medrano, Marta Lagos, Pippa Norris, Eduard Ponarin, Bi Puranen, and 1 others. 2020 · 2020
Cited alongside, same era.
Interpretable deep learning under fire
Xinyang Zhang, Ningfei Wang, Hua Shen, Shouling Ji, Xiapu Luo, and Ting Wang. 2020 · 2020
Cited alongside, same era.
Human-ai interaction in human resource management: Understanding why employees resist algorithmic evaluation at workplaces and how to mitigate burdens
Hyanghee Park, Daehwan Ahn, Kartik Hosanagar, and Joonhwan Lee. 2021 · 2021
Cited alongside, same era.
A framework of severity for harmful content online
Morgan Klaus Scheuerman, Jialun Aaron Jiang, Casey Fiesler, and Jed R Brubaker. 2021 · 2021
Cited alongside, same era.
Constitutional ai: Harmlessness from ai feedback
Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, and 1 others. 2022 · 2022
Faithfulness vs. plausibility: On the (un) reliability of explanations from large language models
Chirag Agarwal, Sree Harsha Tanneru, and Himabindu Lakkaraju. 2024 · 2024
Later among the works it cites.
High-dimension human value representation in large language models
Samuel Cahyawijaya, Delong Chen, Yejin Bang, Leila Khalatbari, Bryan Wilie, Ziwei Ji, Etsuko Ishii, and Pascale Fung. 2024 · 2024
Later among the works it cites.
" they are uncultured": Unveiling covert harms and social threats in llm generated conversations
Preetam Prabhu Srikar Dammu, Hayoung Jung, Anjali Singh, Monojit Choudhury, and Tanushree Mitra. 2024 · 2024
Later among the works it cites.
Risk and response in large language models: Evaluating key threat categories
Bahareh Harandizadeh, Abel Salinas, and Fred Morstatter. 2024 · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned
Deep Ganguli, Liane Lovitt, Jackson Kernion, Amanda Askell, Yuntao Bai, Saurav Kadavath, Ben Mann, Ethan Perez, Nicholas Schiefer, Kamal Ndousse, and 1 others. 2022 · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, and 1 others. 2022 · 2022
Cited alongside, same era.
Improving fairness in speaker verification via group-adapted fusion network
Hua Shen, Yuguang Yang, Guoli Sun, Ryan Langman, Eunjung Han, Jasha Droppo, and Andreas Stolcke. 2022 · 2022
Cited alongside, same era.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, and 1 others. 2023 · 2023
Cited alongside, same era.
How (not) to use sociodemographic information for subjective nlp tasks
Tilman Beck, Hendrik Schuff, Anne Lauscher, and Iryna Gurevych. 2023 · 2023
Cited alongside, same era.
Study and analysis of chat gpt and its impact on different fields of study
Dinesh Kalla, Nathan Smith, Fnu Samaah, and Sivaraju Kuraku. 2023 · 2023
Cited alongside, same era.
What is the impact of chatgpt on education? a rapid review of the literature
Chung Kwan Lo. 2023 · 2023
Cited alongside, same era.
Aaron Hurst, Adam Lerer, Adam P Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, and 1 others. 2024 · 2024
Later among the works it cites.
Can language models reason about individualistic human values and preferences?
Liwei Jiang, Sydney Levine, and Yejin Choi. 2024b · 2024
Later among the works it cites.
Llm-globe: A benchmark evaluating the cultural values embedded in llm output
Elise Karinshak, Amanda Hu, Kewen Kong, Vishwanatha Rao, Jingren Wang, Jindong Wang, and Yi Zeng. 2024 · 2024
Later among the works it cites.
Julia Kharchenko, Tanya Roosta, Aman Chadha, and Chirag Shah. 2024 · 2024
Later among the works it cites.
The benefits, risks and bounds of personalizing the alignment of large language models to individuals
Hannah Rose Kirk, Bertie Vidgen, Paul Röttger, and Scott A Hale. 2024 · 2024
Later among the works it cites.
The generation gap: Exploring age bias in the value systems of large language models
Siyang Liu, Trisha Maturi, Bowen Yi, Siqi Shen, and Rada Mihalcea. 2024 · 2024
Later among the works it cites.
Yuanyi Ren, Haoran Ye, Hanjun Fang, Xin Zhang, and Guojie Song. 2024 · 2024
Later among the works it cites.
Paul Röttger, Valentin Hofmann, Valentina Pyatkin, Musashi Hinck, Hannah Rose Kirk, Hinrich Schütze, and Dirk Hovy. 2024 · 2024
Later among the works it cites.
Evaluating large language models with fmeval
Pola Schwöbel, Luca Franceschi, Muhammad Bilal Zafar, Keerthan Vasist, Aman Malhotra, Tomer Shenhar, Pinal Tailor, Pinar Yilmaz, Michael Diamond, and Michele Donini. 2024 · 2024
Later among the works it cites.
A roadmap to pluralistic alignment
Taylor Sorensen, Jared Moore, Jillian Fisher, Mitchell Gordon, Niloofar Mireshghallah, Christopher Michael Rytting, Andre Ye, Liwei Jiang, Ximing Lu, Nouha Dziri, and 1 others. 2024 · 2024
Later among the works it cites.
Gender, race, and intersectional bias in resume screening via language model retrieval
Kyra Wilson and Aylin Caliskan. 2024 · 2024
Later among the works it cites.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
DeepSeek-AI. 2025 · 2025
Closest in time.