Fetching the paper…
Reading the bibliography…
The awareness of multi-cultural human values is critical to the ability of language models (LMs) to generate safe and personalized responses.
Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned
Deep Ganguli, Liane Lovitt, Jackson Kernion, Amanda Askell, Yuntao Bai, Saurav Kadavath, Ben Mann, Ethan Perez, Nicholas Schiefer, Kamal Ndousse, et al. 2022 · 2022
Earlier work this paper cites.
Demographic-aware language model fine-tuning as a bias mitigation technique
Aparna Garimella, Rada Mihalcea, and Akhash Amarnath. 2022 · 2022
Earlier work this paper cites.
World values survey wave 7 (2017-2022) cross-national data-set, version 4.0.0
C. Haerpfer, R. Inglehart, A. Moreno, C. Welzel, K. Kizilova, J. Diez-Medrano, M. Lagos, P. Norris, E. Ponarin, and Puranen B. 2022 · 2022
Earlier work this paper cites.
World values survey: All rounds - country-pooled datafile, version 3.0.0
R. Inglehart, C. Haerpfer, A. Moreno, C. Welzel, J. Diez-Medrano K. Kizilova, M. Lagos, P. Norris, E. Ponarin, and B. Puranen. 2022 · 2022
Earlier work this paper cites.
The ghost in the machine has an american accent: value conflict in gpt-3
Rebecca L Johnson, Giada Pistilli, Natalia Menédez-González, Leslye Denisse Dias Duran, Enrico Panai, Julija Kalpokiene, and Donald Jay Bertulfo. 2022 · 2022
Earlier work this paper cites.
A cross-national multilevel analysis of fear of crime: Exploring the roles of institutional confidence and institutional performance
Kai Lin. 2022 · 2022
Earlier work this paper cites.
Aligning generative language models with human values
Ruibo Liu, Ge Zhang, Xinyu Feng, and Soroush Vosoughi. 2022 · 2022
Cited alongside, same era.
When large language models meet personalization: Perspectives of challenges and opportunities
Jin Chen, Zheng Liu, Xu Huang, Chenwang Wu, Qi Liu, Gangwei Jiang, Yuanhao Pu, Yuxuan Lei, Xiaolong Chen, Xingmei Wang, et al. 2023 · 2023
Cited alongside, same era.
Towards measuring the representation of subjective global opinions in language models
Esin Durmus, Karina Nyugen, Thomas Liao, Nicholas Schiefer, Amanda Askell, Anton Bakhtin, Carol Chen, Zac Hatfield-Dodds, Danny Hernandez, Nicholas Joseph, Liane Lovitt, Sam McCandlish, Orowa Sikder, Alex Tamkin, Janel Thamkul, Jared Kaplan, Jack Clark, and Deep Ganguli. 2023 · 2023
Cited alongside, same era.
Prioritizing the environment or economic growth: insights from the world values survey
Nuria López de Calle Bastida. 2023 · 2023
Cited alongside, same era.
Whose opinions do language models reflect?
Fine-grained human feedback gives better rewards for language model training
Zeqiu Wu, Yushi Hu, Weijia Shi, Nouha Dziri, Alane Suhr, Prithviraj Ammanabrolu, Noah A. Smith, Mari Ostendorf, and Hanna Hajishirzi. 2023 · 2023
Later among the works it cites.
Constructive large language models alignment with diverse feedback
Tianshu Yu, Ting-En Lin, Yuchuan Wu, Min Yang, Fei Huang, and Yongbin Li. 2023 · 2023
Later among the works it cites.
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Haotong Zhang, Joseph Gonzalez, and Ion Stoica. 2023 · 2023
Later among the works it cites.
Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de Las Casas, Emma Bou Hanna, Florian Bressand, Gianna Lengyel, Guillaume Bour, Guillaume Lample, L’elio Renard Lavaud, Lucile Saulnier, Marie-Anne Lachaux, Pierre Stock, Sandeep Subramanian, Sophia Yang, Szymon Antoniak, Teven Le Scao, Théophile Gervet, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. 2024 · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee, Percy Liang, and Tatsunori Hashimoto. 2023 · 2023
Cited alongside, same era.
Stanford alpaca: An instruction-following llama model
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023 · 2023
Cited alongside, same era.
Probing pre-trained language models for cross-cultural differences in values
Arnav Arora, Lucie-Aimée Kaffee, and Isabelle Augenstein. 2022a
Cited in the paper.
Probing pre-trained language models for cross-cultural differences in values
Arnav Arora, Lucie-Aimée Kaffee, and Isabelle Augenstein. 2022b
Cited in the paper.
Closest in time.
Culturellm: Incorporating cultural differences into large language models
Cheng Li, Mengzhou Chen, Jindong Wang, Sunayana Sitaram, and Xing Xie. 2024 · 2024
Closest in time.
GPT-3.5 Turbo fine-tuning and API updates
Andrew Peng, Michael Wu, John Allard, Logan Kilpatrick, and Steven Heidel. 2023 · 2024
Closest in time.