Fetching the paper…
Reading the bibliography…
Cultural bias is pervasive in many large language models (LLMs), largely due to the deficiency of data representative of different cultures.
Extracting cultural commonsense knowledge at scale
Tuan-Phong Nguyen, Simon Razniewski, Aparna Varde, and Gerhard Weikum · 1917
Earlier work this paper cites.
Cognitive conflict and goal conflict effects on task performance
Richard A Cosier and Gerald L Rose · 1977
Earlier work this paper cites.
Social cognition
Susan T Fiske and Shelley E Taylor · 1991
Earlier work this paper cites.
Situated learning: Legitimate peripheral participation
Jean Lave and Etienne Wenger · 1991
Earlier work this paper cites.
Situated learning and education
John R Anderson, Lynne M Reder, and Herbert A Simon · 1996
Earlier work this paper cites.
Modernity at large: Cultural dimensions of globalization
Arjun Appadurai · 1996
Earlier work this paper cites.
On the cognitive conflict as an instructional strategy for conceptual change: A critical appraisal
Margarita Limón · 2001
Earlier work this paper cites.
Cultures and Organizations: Software of the Mind, Third Edition
Michael Minkov Geert Hofstede, Gert Jan Hofstede · 2010
Earlier work this paper cites.
What is culture
Helen Spencer-Oatey and Peter Franklin · 2012
Earlier work this paper cites.
The 2012 stein rokkan lecture: Three decades of popu list radical right parties in western europe: so what?
Cas Mudde · 2016
Earlier work this paper cites.
Measuring the Reliability of Hate Speech Annotations: The Case of the European Refugee Crisis
Björn Ross, Michael Rist, Guillermo Carbonell, Benjamin Cabrera, Nils Kurowsky, and Michael Wojatzki · 2016
Earlier work this paper cites.
Hateful symbols or hateful people? predictive features for hate speech detection on twitter
Zeerak Waseem and Dirk Hovy · 2016
Earlier work this paper cites.
Offensive comments in the brazilian web: a dataset and baseline results
Rogers P. de Pelle and Viviane P. Moreira · 2017
Earlier work this paper cites.
Bangla-abusive-comment-dataset
aimansnigdha · 2018
Earlier work this paper cites.
Overview of mex-a3t at ibereval 2018: Authorship and aggressiveness analysis in mexican spanish tweets
Miguel Á Álvarez-Carmona, Estefanıa Guzmán-Falcón, Manuel Montes-y Gómez, Hugo Jair Escalante, Luis Villasenor-Pineda, Verónica Reyes-Meza, and Antonio Rico-Sulayes · 2018
Earlier work this paper cites.
Overview of the task on automatic misogyny identification at ibereval 2018
Elisabetta Fersini, Paolo Rosso, Maria Anzovino, et al · 2018
Earlier work this paper cites.
Overview of the germeval 2018 shared task on the identification of offensive language
Michael Wiegand, Melanie Siegel, and Josef Ruppenhofer · 2018
Earlier work this paper cites.
Semeval-2019 task 5: Multilingual detection of hate speech against immigrants and women in twitter
Valerio Basile, Cristina Bosco, Elisabetta Fersini, Debora Nozza, Viviana Patti, Francisco Manuel Rangel Pardo, Paolo Rosso, and Manuela Sanguinetti · 2019
Earlier work this paper cites.
Detect camouflaged spam content via stoneskipping: Graph and text joint embedding for chinese character variation representation
Zhuoren Jiang, Zhe Gao, Guoxiu He, Yangyang Kang, Changlong Sun, Qiong Zhang, Luo Si, and Xiaozhong Liu · 2019
Earlier work this paper cites.
Detecting and monitoring hate speech in twitter
Juan Carlos Pereira-Kohatsu, Lara Quijano-Sánchez, Federico Liberatore, and Miguel Camacho-Collados · 2019
Earlier work this paper cites.
Turkish Spam V01
TurkishSpamV01 · 2019
Earlier work this paper cites.
Developing a multilingual annotated corpus of misogyny and aggression
Shiladitya Bhattacharya, Siddharth Singh, Ritesh Kumar, Akanksha Bansal, Akash Bhagat, Yogesh Dawer, Bornini Lahiri, and Atul Kr. Ojha · 2020
Earlier work this paper cites.
I feel offended, don’t be abusive! implicit/explicit messages in offensive and abusive language
Tommaso Caselli, Valerio Basile, Jelena Mitrović, Inga Kartoziya, and Michael Granitzer · 2020
Earlier work this paper cites.
A corpus of turkish offensive language on social media
Çağrı Çöltekin · 2020
Earlier work this paper cites.
A multi-platform arabic news comment dataset for offensive language detection
Shammur Absar Chowdhury, Hamdy Mubarak, Ahmed Abdelali, Soon-gyo Jung, Bernard J Jansen, and Joni Salminen · 2020
Earlier work this paper cites.
Korean hatespeech dataset
daanVeer · 2020
Cited alongside, same era.
What is culture
Werner Delanoy · 2020
Cited alongside, same era.
Hasoc2020
HASOC · 2020
Cited alongside, same era.
F Husain · 2020
Cited alongside, same era.
Joao A Leite, Diego F Silva, Kalina Bontcheva, and Carolina Scarton · 2020
Cited alongside, same era.
BEEP! Korean corpus of online news comments for toxic speech detection
Jihyung Moon, Won Ik Cho, and Junbum Lee · 2020
Cited alongside, same era.
Harmonizing global voices: Culturally-aware models for enhanced content moderation
Alex J Chan, José Luis Redondo García, Fabrizio Silvestri, Colm O’Donnel, and Konstantina Palla · 2023
Later among the works it cites.
Recommender systems in the era of large language models (llms)
Wenqi Fan, Zihuai Zhao, Jiatong Li, Yunqing Liu, Xiaowei Mei, Yiqi Wang, Jiliang Tang, and Qing Li · 2023
Later among the works it cites.
Large language models as superpositions of cultural perspectives
Grgur Kovač, Masataka Sawayama, Rémy Portelas, Cédric Colas, Peter Ford Dominey, and Pierre-Yves Oudeyer · 2023
Later among the works it cites.
Prompt distillation for efficient llm-based recommendation
Lei Li, Yongfeng Zhang, and Li Chen · 2023
Later among the works it cites.
Taiwan llm: Bridging the linguistic divide with a culturally aligned language model
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Marcos Zampieri, Preslav Nakov, Sara Rosenthal, Pepa Atanasova, Georgi Karadzhov, Hamdy Mubarak, Leon Derczynski, Zeses Pitenis, and Çağrı Çöltekin · 2020
Cited alongside, same era.
Angel Felipe Magnossao de Paula and Ipek Baris Schlicht · 2021
Cited alongside, same era.
5k turkish tweets with incivil content
Kaggle · 2021
Cited alongside, same era.
Detecting abusive instagram comments in turkish using convolutional neural network and machine learning methods
Habibe Karayiğit, Çiğdem İnan Acı, and Ali Akdağlı · 2021
Cited alongside, same era.
Offendes: A new corpus in spanish for offensive language research
Flor Miriam Plaza-del Arco, Arturo Montejo-Ráez, L Alfonso Urena Lopez, and María-Teresa Martín-Valdivia · 2021
Cited alongside, same era.
Hate speech detection in the bengali language: A dataset and its baseline evaluation
Nauros Romim, Mosahed Ahmed, Hriteshwar Talukder, and Md Saiful Islam · 2021
Cited alongside, same era.
Yen-Ting Lin and Yun-Nung Chen · 2023
Later among the works it cites.
Chen Cecilia Liu, Fajri Koto, Timothy Baldwin, and Iryna Gurevych · 2023
Later among the works it cites.
Reem I Masoud, Ziquan Liu, Martin Ferianc, Philip Treleaven, and Miguel Rodrigues · 2023
Later among the works it cites.
Having beer after prayer? measuring cultural bias in large language models
Tarek Naous, Michael J Ryan, and Wei Xu · 2023
Later among the works it cites.
Typhoon: Thai large language models
Kunat Pipatanakul, Phatrasek Jirabovonvisut, Potsawee Manakul, Sittipong Sripaisarnmongkol, Ruangsak Patomwong, Pathomporn Chokchainant, and Kasima Tharnpipitchai · 2023
Later among the works it cites.
Sabiá: Portuguese large language models
Ramon Pires, Hugo Abonizio, Thales Sales Almeida, and Rodrigo Nogueira · 2023
Later among the works it cites.
Ethical reasoning over moral alignment: A case and framework for in-context ethical policies in llms
Abhinav Rao, Aditi Khandelwal, Kumar Tanmay, Utkarsh Agarwal, and Monojit Choudhury · 2023
Later among the works it cites.
Rehearsal: Simulating conflict to teach conflict resolution
Omar Shaikh, Valentino Chai, Michele J Gelfand, Diyi Yang, and Michael S Bernstein · 2023
Later among the works it cites.
Not all countries celebrate thanksgiving: On the cultural dominance in large language models
Wenxuan Wang, Wenxiang Jiao, Jingyuan Huang, Ruyi Dai, Jen-tse Huang, Zhaopeng Tu, and Michael R Lyu · 2023
Later among the works it cites.
Cvalues: Measuring the values of chinese large language models from safety to responsibility
Guohai Xu, Jiayi Liu, Ming Yan, Haotian Xu, Jinghui Si, Zhuoran Zhou, Peng Yi, Xing Gao, Jitao Sang, Rong Zhang, Ji Zhang, Chao Peng, Fei Huang, and Jingren Zhou · 2023
Later among the works it cites.
Massively multi-cultural knowledge acquisition & lm benchmarking
Yi Fung, Ruining Zhao, Jae Doo, Chenkai Sun, and Heng Ji · 2024
Closest in time.
Dataset of arabic spam and ham tweets
Sanaa Kaddoura and Safaa Henno · 2024
Closest in time.
Culturellm: Incorporating cultural differences into large language models
Cheng Li, Mengzhou Chen, Jindong Wang, Sunayana Sitaram, and Xing Xie · 2024
Closest in time.
Mala-500: Massive language adaptation of large language models
Peiqin Lin, Shaoxiong Ji, Jörg Tiedemann, André FT Martins, and Hinrich Schütze · 2024
Closest in time.
Blend: A benchmark for llms on everyday knowledge in diverse cultures and languages
Junho Myung, Nayeon Lee, Yi Zhou, Jiho Jin, Rifki Afina Putri, Dimosthenis Antypas, Hsuvas Borkakoty, Eunsu Kim, Carla Perez-Almendros, Abinew Ali Ayele, et al · 2024
Closest in time.
text-embedding-3-small
OpenAI · 2024
Closest in time.
Normad: A benchmark for measuring the cultural adaptability of large language models
Abhinav Rao, Akhila Yerukola, Vishwa Shah, Katharina Reinecke, and Maarten Sap · 2024
Closest in time.
Unintended impacts of llm alignment on global representation
Michael J Ryan, William Held, and Diyi Yang · 2024
Closest in time.
Weiyan Shi, Ryan Li, Yutong Zhang, Caleb Ziems, Raya Horesh, Rogério Abreu de Paula, Diyi Yang, et al · 2024
Closest in time.
Social skill training with large language models
Diyi Yang, Caleb Ziems, William Held, Omar Shaikh, Michael S Bernstein, and John Mitchell · 2024
Closest in time.
Wenlong Zhao, Debanjan Mondal, Niket Tandon, Danica Dillion, Kurt Gray, and Yuling Gu · 2024
Closest in time.