Fetching the paper…
Reading the bibliography…
We present a survey of more than 90 recent papers that aim to study cultural representation and inclusion in large language models (LLMs).
The concept of culture*
Leslie A. White. 1959 · 1959
Earlier work this paper cites.
CrowS-pairs: A challenge dataset for measuring social biases in masked language models
Nikita Nangia, Clara Vania, Rasika Bhalerao, and Samuel R. Bowman. 2020 · 1967
Earlier work this paper cites.
Outline of a Theory of Practice
P. Bourdieu. 1972 · 1972
Earlier work this paper cites.
Culture and social system revisited
Talcott Parsons. 1972 · 1972
Earlier work this paper cites.
The Interpretation Of Cultures
C. Geertz. 1973 · 1973
Earlier work this paper cites.
Development in Judging Moral Issues
J.R. Rest and L. Kohlberg. 1979 · 1979
Earlier work this paper cites.
Theory of Culture
R. Münch, N.J. Smelser, American Sociological Association. Theory Section, Deutsche Gesellschaft fur Soziologie. Sektion Soziologische Theorien, and Deutsche Gesellschaft für Soziologie. Sektion Soziologische Theorien. 1992 · 1992
Earlier work this paper cites.
Workflow from within and without: Technology and cooperative work on the print industry shopfloor
John Bowers, Graham Button, and Wes Sharrock. 1995 · 1995
Earlier work this paper cites.
At home with the technology: an ethnographic study of a set-top-box trial
Jon O’Brien, Tom Rodden, Mark Rouncefield, and John A. Hughes. 1999 · 1999
Earlier work this paper cites.
On defining the cultural heritage
Janet Blake. 2000 · 2000
Earlier work this paper cites.
The Cultural Psychology of Development: One Mind, Many Mentalities , volume 1
Richard Shweder, Jacqueline Goodnow, Giyoo Hatano, Robert LeVine, Hazel Markus, and Peggy Miller. 2007 · 2007
Earlier work this paper cites.
The number of distinct basic values and their structure assessed by pvq–40
Jan Cieciuch and Shalom Schwartz. 2012 · 2012
Earlier work this paper cites.
A Cultural Approach to Interpersonal Communication: Essential Readings
L. Monaghan, J.E. Goodman, and J. Robinson. 2012 · 2012
Earlier work this paper cites.
Cultures and organizations: Software of the mind intercultural cooperation and its importance for survival
Cristina Mora. 2013 · 2013
Earlier work this paper cites.
Peer-to-peer in the workplace: A view from the road
Syed Ishtiaque Ahmed, Nicola J. Bidwell, Himanshu Zade, Srihari H. Muralidhar, Anupama Dhareshwar, Baneen Karachiwala, Cedrick N. Tandong, and Jacki O’Neill. 2016 · 2016
Earlier work this paper cites.
Semantics derived automatically from language corpora contain human-like biases
Aylin Caliskan, Joanna J. Bryson, and Arvind Narayanan. 2017 · 2017
Earlier work this paper cites.
How should my chatbot interact? a survey on social characteristics in human–chatbot interaction design
Ana Paula Chaves and Marco Aurélio Gerosa. 2019 · 2019
Earlier work this paper cites.
Cross-cultural transfer learning for text classification
Dor Ringel, Gal Lavee, Ido Guy, and Kira Radinsky. 2019 · 2019
Earlier work this paper cites.
Language (technology) is power: A critical survey of “bias” in NLP
Su Lin Blodgett, Solon Barocas, Hal Daumé III, and Hanna Wallach. 2020 · 2020
Earlier work this paper cites.
Social chemistry 101: Learning to reason about social and moral norms
Maxwell Forbes, Jena D. Hwang, Vered Shwartz, Maarten Sap, and Yejin Choi. 2020 · 2020
Earlier work this paper cites.
Making chat at home in the hospital: Exploring chat use by nurses
Naveena Karusala, Ding Wang, and Jacki O’Neill. 2020 · 2020
Earlier work this paper cites.
RiSAWOZ: A large-scale multi-domain Wizard-of-Oz dataset with rich semantic annotations for task-oriented dialogue modeling
Jun Quan, Shian Zhang, Qian Cao, Zizhong Li, and Deyi Xiong. 2020 · 2020
Earlier work this paper cites.
Cultural influences on word meanings revealed through large-scale semantic alignment
Bill Thompson, Se’an G. Roberts, and Gary Lupyan. 2020 · 2020
Earlier work this paper cites.
On the dangers of stochastic parrots: Can language models be too big?
Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021 · 2021
Earlier work this paper cites.
StereoSet: Measuring stereotypical bias in pretrained language models
Moin Nadeem, Anna Bethke, and Siva Reddy. 2021 · 2021
Earlier work this paper cites.
Re-imagining algorithmic fairness in india and beyond
Nithya Sambasivan, Erin Arnesen, Ben Hutchinson, Tulsee Doshi, and Vinodkumar Prabhakaran. 2021 · 2021
Earlier work this paper cites.
Large pre-trained language models contain human-like biases of what is right and wrong to do
Patrick Schramowski, Cigdem Turan, Nico Andersen, Constantin A. Rothkopf, and Kristian Kersting. 2021 · 2021
Earlier work this paper cites.
A word on machine ethics: A response to jiang et al. (2021)
Zeerak Talat, Hagen Blix, Josef Valvoda, Maya Indira Ganesh, Ryan Cotterell, and Adina Williams. 2021 · 2021
Earlier work this paper cites.
Restoring and attributing ancient texts using deep neural networks
Yannis Assael, Thea Sommerschield, Brendan Shillingford, Mahyar Bordbar, John Pavlopoulos, Maria Chatzipanagiotou, Ion Androutsopoulos, Jonathan Prag, and Nando Freitas. 2022 · 2022
Earlier work this paper cites.
Sapir’s thought-grooves and whorf’s tensors: Reconciling transformer architectures with cultural anthropology
Michael Castelle. 2022 · 2022
Earlier work this paper cites.
Ai ethics and the future of where large language models are heading
Lance Eliot. 2022 · 2022
Earlier work this paper cites.
Joint evs/wvs 2017-2022 dataset (joint evs/wvs)
EVS/WVS. 2022 · 2022
Earlier work this paper cites.
Improving alignment of dialogue agents via targeted human judgements
Amelia Glaese, Nat McAleese, Maja Trębacz, John Aslanides, Vlad Firoiu, Timo Ewalds, Maribeth Rauh, Laura Weidinger, Martin Chadwick, Phoebe Thacker, Lucy Campbell-Gillingham, Jonathan Uesato, Po-Sen Huang, Ramona Comanescu, Fan Yang, Abigail See, Sumanth Dathathri, Rory Greig, Charlie Chen, Doug Fritz, Jaume Sanchez Elias, Richard Green, Soňa Mokrá, Nicholas Fernando, Boxi Wu, Rachel Foley, Susannah Young, Iason Gabriel, William Isaac, John Mellor, Demis Hassabis, Koray Kavukcuoglu, Lisa Anne Hendricks, and Geoffrey Irving. 2022 · 2022
Earlier work this paper cites.
Challenges and strategies in cross-cultural NLP
Daniel Hershcovich, Stella Frank, Heather Lent, Miryam de Lhoneux, Mostafa Abdou, Stephanie Brandl, Emanuele Bugliarello, Laura Cabello Piqueras, Ilias Chalkidis, Ruixiang Cui, Constanza Fierro, Katerina Margatina, Phillip Rust, and Anders Søgaard. 2022 · 2022
Earlier work this paper cites.
Can machines learn morality? the delphi experiment
Liwei Jiang, Jena D. Hwang, Chandra Bhagavatula, Ronan Le Bras, Jenny Liang, Jesse Dodge, Keisuke Sakaguchi, Maxwell Forbes, Jon Borchardt, Saadia Gabriel, Yulia Tsvetkov, Oren Etzioni, Maarten Sap, Regina Rini, and Yejin Choi. 2022 · 2022
Earlier work this paper cites.
The ghost in the machine has an american accent: value conflict in gpt-3
Rebecca L Johnson, Giada Pistilli, Natalia Menédez-González, Leslye Denisse Dias Duran, Enrico Panai, Julija Kalpokiene, and Donald Jay Bertulfo. 2022 · 2022
Earlier work this paper cites.
ArtELingo: A million emotion annotations of WikiArt with emphasis on diversity over language and culture
Youssef Mohamed, Mohamed Abdelfattah, Shyma Alhuwaider, Feifan Li, Xiangliang Zhang, Kenneth Church, and Mohamed Elhoseiny. 2022 · 2022
Earlier work this paper cites.
Annotators with attitudes: How annotator beliefs and identities bias toxic language detection
Maarten Sap, Swabha Swayamdipta, Laura Vianna, Xuhui Zhou, Yejin Choi, and Noah A. Smith. 2022 · 2022
Earlier work this paper cites.
Using large language models to simulate multiple humans and replicate human subject studies
Gati Aher, Rosa I. Arriaga, and Adam Tauman Kalai. 2023 · 2023
Earlier work this paper cites.
SODAPOP: Open-ended discovery of social biases in social commonsense reasoning models
Haozhe An, Zongxia Li, Jieyu Zhao, and Rachel Rudinger. 2023 · 2023
Earlier work this paper cites.
Social commonsense for explanation and cultural bias discovery
Lisa Bauer, Hanna Tischer, and Mohit Bansal. 2023 · 2023
Earlier work this paper cites.
GD-COMET: A geo-diverse commonsense inference model
Mehar Bhatia and Vered Shwartz. 2023 · 2023
Cited alongside, same era.
Assessing cross-cultural alignment between ChatGPT and human societies: An empirical study
Yong Cao, Li Zhou, Seolhwa Lee, Laura Cabello, Min Chen, and Daniel Hershcovich. 2023 · 2023
Cited alongside, same era.
Sociocultural norm similarities and differences via situational alignment and explainable textual entailment
Sky CH-Wang, Arkadiy Saakyan, Oliver Li, Zhou Yu, and Smaranda Muresan. 2023 · 2023
Cited alongside, same era.
Toward cultural bias evaluation datasets: The case of Bengali gender, religious, and national identity
Dipto Das, Shion Guha, and Bryan Semaan. 2023 · 2023
Cited alongside, same era.
Building socio-culturally inclusive stereotype resources with community engagement
Sunipa Dev, Jaya Goyal, Dinesh Tewari, Shachi Dave, and Vinodkumar Prabhakaran. 2023 · 2023
Cited alongside, same era.
Navigating cultural chasms: Exploring and unlocking the cultural pov of text-to-image models
Mor Ventura, Eyal Ben-David, Anna Korhonen, and Roi Reichart. 2023 · 2023
Later among the works it cites.
Are personalized stochastic parrots more dangerous? evaluating persona biases in dialogue systems
Yixin Wan, Jieyu Zhao, Aman Chadha, Nanyun Peng, and Kai-Wei Chang. 2023 · 2023
Later among the works it cites.
Seaeval for multilingual foundation models: From cross-lingual alignment to cultural reasoning
Bin Wang, Zhengyuan Liu, Xin Huang, Fangkai Jiao, Yang Ding, Ai Ti Aw, and Nancy F. Chen. 2023 · 2023
Later among the works it cites.
Copal-id: Indonesian language reasoning with local culture and nuances
Haryo Akbarianto Wibowo, Erland Hilman Fuadi, Made Nindyatama Nityasya, Radityo Eko Prasojo, and Alham Fikri Aji. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Towards measuring the representation of subjective global opinions in language models
Esin Durmus, Karina Nyugen, Thomas I. Liao, Nicholas Schiefer, Amanda Askell, Anton Bakhtin, Carol Chen, Zac Hatfield-Dodds, Danny Hernandez, Nicholas Joseph, Liane Lovitt, Sam McCandlish, Orowa Sikder, Alex Tamkin, Janel Thamkul, Jared Kaplan, Jack Clark, and Deep Ganguli. 2023 · 2023
Cited alongside, same era.
EtiCor: Corpus for analyzing LLMs for etiquettes
Ashutosh Dwivedi, Pradhyumna Lavania, and Ashutosh Modi. 2023 · 2023
Cited alongside, same era.
From pretraining data to language models to downstream tasks: Tracking the trails of political biases leading to unfair NLP models
Shangbin Feng, Chan Young Park, Yuhan Liu, and Yulia Tsvetkov. 2023 · 2023
Cited alongside, same era.
EPIC: Multi-perspective annotation of a corpus of irony
Simona Frenda, Alessandro Pedrani, Valerio Basile, Soda Marem Lo, Alessandra Teresa Cignarella, Raffaella Panizzon, Cristina Marco, Bianca Scarlini, Viviana Patti, Cristina Bosco, and Davide Bernardi. 2023 · 2023
Cited alongside, same era.
Revision Transformers: Instructing Language Models to Change Their Values
Felix Friedrich, Wolfgang Stammer, Patrick Schramowski, and Kristian Kersting. 2023 · 2023
Cited alongside, same era.
NORMSAGE: Multi-lingual multi-cultural norm discovery from conversations on-the-fly
Yi Fung, Tuhin Chakrabarty, Hao Guo, Owen Rambow, Smaranda Muresan, and Heng Ji. 2023 · 2023
Cited alongside, same era.
Chatgpt and the global south: how are journalists in sub-saharan africa engaging with generative ai?
Greg Gondwe. 2023 · 2023
Cited alongside, same era.
Winston Wu, Lu Wang, and Rada Mihalcea. 2023 · 2023
Later among the works it cites.
From instructions to intrinsic human values – a survey of alignment goals for big models
Jing Yao, Xiaoyuan Yi, Xiting Wang, Jindong Wang, and Xing Xie. 2023 · 2023
Later among the works it cites.
Socialdial: A benchmark for socially-aware dialogue systems
Haolan Zhan, Zhuang Li, Yufei Wang, Linhao Luo, Tao Feng, Xiaoxi Kang, Yuncheng Hua, Lizhen Qu, Lay-Ki Soon, Suraj Sharma, Ingrid Zukerman, Zhaleh Semnani-Azad, and Gholamreza Haffari. 2023 · 2023
Later among the works it cites.
The skipped beat: A study of sociopragmatic understanding in LLMs for 64 languages
Chiyu Zhang, Khai Doan, Qisheng Liao, and Muhammad Abdul-Mageed. 2023 · 2023
Later among the works it cites.
Cultural compass: Predicting transfer learning success in offensive language detection with cultural features
Li Zhou, Antonia Karamolegkou, Wenyu Chen, and Daniel Hershcovich. 2023b · 2023
Later among the works it cites.
NormBank: A knowledge bank of situational social norms
Caleb Ziems, Jane Dwivedi-Yu, Yi-Chia Wang, Alon Halevy, and Diyi Yang. 2023 · 2023
Later among the works it cites.
Investigating cultural alignment of large language models
Badr AlKhamissi, Muhammad ElNokrashy, Mai AlKhamissi, and Mona Diab. 2024 · 2024
Closest in time.
Studying large language models as compression algorithms for human culture
Nicholas Buttrick. 2024 · 2024
Closest in time.
Bridging cultural nuances in dialogue agents through cultural value surveys
Yong Cao, Min Chen, and Daniel Hershcovich. 2024a · 2024
Closest in time.
Benchmarking large language models in retrieval-augmented generation
Jiawei Chen, Hongyu Lin, Xianpei Han, and Le Sun. 2024 · 2024
Closest in time.
Yu Ying Chiu, Liwei Jiang, Maria Antoniak, Chan Young Park, Shuyue Stella Li, Mehar Bhatia, Sahithya Ravi, Yulia Tsvetkov, Vered Shwartz, and Yejin Choi. 2024 · 2024
Closest in time.
The echoes of multilinguality: Tracing cultural value shifts during lm fine-tuning
Rochelle Choenni, Anne Lauscher, and Ekaterina Shutova. 2024 · 2024
Closest in time.
”it’s how you do things that matters”: Attending to process to better serve indigenous communities with language technologies
Ned Cooper, Courtney Heldreth, and Ben Hutchinson. 2024 · 2024
Closest in time.
" they are uncultured": Unveiling covert harms and social threats in llm generated conversations
Preetam Prabhu Srikar Dammu, Hayoung Jung, Anjali Singh, Monojit Choudhury, and Tanushree Mitra. 2024 · 2024
Closest in time.
Towards measuring the representation of subjective global opinions in language models
Esin Durmus, Karina Nguyen, Thomas I. Liao, Nicholas Schiefer, Amanda Askell, Anton Bakhtin, Carol Chen, Zac Hatfield-Dodds, Danny Hernandez, Nicholas Joseph, Liane Lovitt, Sam McCandlish, Orowa Sikder, Alex Tamkin, Janel Thamkul, Jared Kaplan, Jack Clark, and Deep Ganguli. 2024 · 2024
Closest in time.
Massively multi-cultural knowledge acquisition and lm benchmarking
Yi Fung, Ruining Zhao, Jae Doo, Chenkai Sun, and Heng Ji. 2024 · 2024
Closest in time.
Self-assessment tests are unreliable measures of llm personality
Akshat Gupta, Xiaoyang Song, and Gopala Anumanchipalli. 2024 · 2024
Closest in time.
Factors influencing intention to engage in human–chatbot interaction: examining user perceptions and context culture orientation
Luna Luan Haoyue and Hichang Cho. 2024 · 2024
Closest in time.
Kobbq: Korean bias benchmark for question answering
Jiho Jin, Jiseon Kim, Nayeon Lee, Haneul Yoo, Alice Oh, and Hwaran Lee. 2024 · 2024
Closest in time.
CLIcK: A benchmark dataset of cultural and linguistic intelligence in Korean
Eunsu Kim, Juyoung Suk, Philhoon Oh, Haneul Yoo, James Thorne, and Alice Oh. 2024 · 2024
Closest in time.
Hannah Rose Kirk, Alexander Whitefield, Paul Röttger, Andrew Bean, Katerina Margatina, Juan Ciro, Rafael Mosquera, Max Bartolo, Adina Williams, He He, Bertie Vidgen, and Scott A. Hale. 2024 · 2024
Closest in time.
Fajri Koto, Rahmad Mahendra, Nurul Aisyah, and Timothy Baldwin. 2024 · 2024
Closest in time.
A "perspectival" mirror of the elephant: Investigating language bias on google, chatgpt, youtube, and wikipedia
Queenie Luo, Michael J. Puett, and Michael D. Smith. 2024 · 2024
Closest in time.
Reem I. Masoud, Ziquan Liu, Martin Ferianc, Philip Treleaven, and Miguel Rodrigues. 2024 · 2024
Closest in time.
In-between visuals and visible: The impacts of text-to-image generative ai tools on digital image-making practices in the global south
Nusrat Jahan Mim, Dipannita Nandi, Sadaf Sumyia Khan, Arundhuti Dey, and Syed Ishtiaque Ahmed. 2024 · 2024
Closest in time.
Multi-cultural commonsense knowledge distillation
Tuan-Phong Nguyen, Simon Razniewski, and Gerhard Weikum. 2024 · 2024
Closest in time.
Millicent Ochieng, Varun Gumma, Sunayana Sitaram, Jindong Wang, Vishrav Chaudhary, Keshet Ronen, Kalika Bali, and Jacki O’Neill. 2024 · 2024
Closest in time.
Can llm generate culturally relevant commonsense qa data? case study in indonesian and sundanese
Rifki Afina Putri, Faiz Ghifari Haznitrama, Dea Adhista, and Alice Oh. 2024 · 2024
Closest in time.
A cross-cultural analysis of social norms in bollywood and hollywood movies
Sunny Rai, Khushang Jilesh Zaveri, Shreya Havaldar, Soumna Nema, Lyle Ungar, and Sharath Chandra Guntuku. 2024 · 2024
Closest in time.
Weiyan Shi, Ryan Li, Yutong Zhang, Caleb Ziems, Chunhua yu, Raya Horesh, Rogério Abreu de Paula, and Diyi Yang. 2024 · 2024
Closest in time.
Kmmlu: Measuring massive multitask language understanding in korean
Guijin Son, Hanwool Lee, Sungdong Kim, Seungone Kim, Niklas Muennighoff, Taekyoon Choi, Cheonbok Park, Kang Min Yoo, and Stella Biderman. 2024 · 2024
Closest in time.
Gradient-based language model red teaming
Nevan Wichers, Carson Denison, and Ahmad Beirami. 2024 · 2024
Closest in time.
Benchmarking llm-based machine translation on cultural awareness
Binwei Yao, Ming Jiang, Diyi Yang, and Junjie Hu. 2024 · 2024
Closest in time.
Renovi: A benchmark towards remediating norm violations in socio-cultural conversations
Haolan Zhan, Zhuang Li, Xiaoxi Kang, Tao Feng, Yuncheng Hua, Lizhen Qu, Yi Ying, Mei Rianto Chandra, Kelly Rosalin, Jureynolds Jureynolds, Suraj Sharma, Shilin Qu, Linhao Luo, Lay-Ki Soon, Zhaleh Semnani Azad, Ingrid Zukerman, and Gholamreza Haffari. 2024 · 2024
Closest in time.
WorldValuesBench: A large-scale benchmark dataset for multi-cultural value awareness of language models
Wenlong Zhao, Debanjan Mondal, Niket Tandon, Danica Dillion, Kurt Gray, and Yuling Gu. 2024 · 2024
Closest in time.
Does mapo tofu contain coffee? probing llms for food-related cultural knowledge
Li Zhou, Taelin Karidi, Nicolas Garneau, Yong Cao, Wanlong Liu, Wenyu Chen, and Daniel Hershcovich. 2024 · 2024
Closest in time.