Fetching the paper…
Reading the bibliography…
As the reach of large language models (LMs) expands globally, their ability to cater to diverse cultural contexts becomes crucial.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020 · 1901
Earlier work this paper cites.
Extracting cultural commonsense knowledge at scale
Tuan-Phong Nguyen, Simon Razniewski, Aparna Varde, and Gerhard Weikum. 2023 · 1917
Earlier work this paper cites.
CrowS-pairs: A challenge dataset for measuring social biases in masked language models
Nikita Nangia, Clara Vania, Rasika Bhalerao, and Samuel R. Bowman. 2020a · 1967
Earlier work this paper cites.
A cluster separation measure
David L. Davies and Donald W. Bouldin. 1979 · 1979
Earlier work this paper cites.
Culture’s consequences: International differences in work-related values , volume 5
Geert Hofstede. 1984 · 1984
Earlier work this paper cites.
Food is culture
Massimo Montanari. 2006 · 2006
Earlier work this paper cites.
Anersys: An arabic named entity recognition system based on maximum entropy
Yassine Benajiba, Paolo Rosso, and José Miguel Benedíruiz. 2007 · 2007
Earlier work this paper cites.
Visualizing data using t-sne
Laurens Van der Maaten and Geoffrey Hinton. 2008 · 2008
Earlier work this paper cites.
The community, autonomy, and divinity scale (cads): A new tool for the cross-cultural study of morality
Valeschka M Guerra and Roger Giner-Sorolla. 2010 · 2010
Earlier work this paper cites.
Crows-pairs: A challenge dataset for measuring social biases in masked language models
Nikita Nangia, Clara Vania, Rasika Bhalerao, and Samuel R Bowman. 2020b · 2010
Earlier work this paper cites.
Explaining odds ratios
Magdalena Szumilas. 2010 · 2010
Earlier work this paper cites.
Mapping the moral domain
Jesse Graham, Brian A Nosek, Jonathan Haidt, Ravi Iyer, Spassena Koleva, and Peter H Ditto. 2011 · 2011
Earlier work this paper cites.
The opengrm open-source finite-state grammar software libraries
Brian Roark, Richard Sproat, Cyril Allauzen, Michael Riley, Jeffrey Sorensen, and Terry Tai. 2012 · 2012
Earlier work this paper cites.
Rule-based information extraction is dead! long live rule-based information extraction systems!
Laura Chiticariu, Yunyao Li, and Frederick R. Reiss. 2013 · 2013
Earlier work this paper cites.
Wikidata: a free collaborative knowledgebase
Denny Vrandečić and Markus Krötzsch. 2014 · 2014
Earlier work this paper cites.
Farasa: A fast and furious segmenter for arabic
Ahmed Abdelali, Kareem Darwish, Nadir Durrani, and Hamdy Mubarak. 2016 · 2016
Earlier work this paper cites.
1.5 billion words Arabic corpus
Ibrahim Abu El-Khair. 2016 · 2016
Earlier work this paper cites.
The abc of stereotypes about groups: Agency/socioeconomic success, conservative–progressive beliefs, and communion
Alex Koch, Roland Imhoff, Ron Dotsch, Christian Unkelbach, and Hans Alves. 2016 · 2016
Earlier work this paper cites.
Semantics derived automatically from language corpora contain human-like biases
Aylin Caliskan, Joanna J Bryson, and Arvind Narayanan. 2017 · 2017
Earlier work this paper cites.
Hotel arabic-reviews dataset construction for sentiment analysis applications
Ashraf Elnagar, Yasmin S Khalifa, and Anas Einea. 2018 · 2018
Earlier work this paper cites.
Football clubs as symbols of regional identities
Adriano Gómez-Bantel. 2018 · 2018
Earlier work this paper cites.
Cultural and social identity in clothing matters “different cultures, different meanings”
Fatjri Nur Tajuddin. 2018 · 2018
Earlier work this paper cites.
Gender bias in coreference resolution: Evaluation and debiasing methods
Jieyu Zhao, Tianlu Wang, Mark Yatskar, Vicente Ordonez, and Kai-Wei Chang. 2018 · 2018
Earlier work this paper cites.
Good secretaries, bad truck drivers? occupational gender stereotypes in sentiment analysis
Jayadev Bhaskaran and Isha Bhallamudi. 2019 · 2019
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
Are we consistently biased? multidimensional analysis of biases in distributional word vectors
Anne Lauscher and Goran Glavaš. 2019 · 2019
Earlier work this paper cites.
On measuring social biases in sentence encoders
Chandler May, Alex Wang, Shikha Bordia, Samuel Bowman, and Rachel Rudinger. 2019 · 2019
Earlier work this paper cites.
Asynchronous pipeline for processing huge corpora on medium to low resource infrastructures
Pedro Javier Ortiz Suárez, Benoît Sagot, and Laurent Romary. 2019 · 2019
Earlier work this paper cites.
Assessing social and intersectional biases in contextualized word representations
Yi Chern Tan and L Elisa Celis. 2019 · 2019
Earlier work this paper cites.
OSIAN: Open source international Arabic news corpus-preparation and integration into the CLARIN-infrastructure
Imad Zeroual, Dirk Goldhahn, Thomas Eckart, and Abdelhak Lakhouaja. 2019 · 2019
Earlier work this paper cites.
AraBERT: Transformer-based model for arabic language understanding
Wissam Antoun, Fady Baly, and Hazem Hajj. 2020 · 2020
Earlier work this paper cites.
Unsupervised cross-lingual representation learning at scale
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Édouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020 · 2020
Earlier work this paper cites.
Cultural differences in bias? origin and gender bias in pre-trained german and french word embeddings
Mascha Kurpicz-Briki. 2020 · 2020
Earlier work this paper cites.
An empirical study of pre-trained transformers for arabic information extraction
Wuwei Lan, Yang Chen, Wei Xu, and Alan Ritter. 2020 · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020 · 2020
Earlier work this paper cites.
Masked language model scoring
Julian Salazar, Davis Liang, Toan Q Nguyen, and Katrin Kirchhoff. 2020 · 2020
Earlier work this paper cites.
ARBERT & MARBERT: Deep bidirectional transformers for arabic
Muhammad Abdul-Mageed, AbdelRahim Elmadany, et al. 2021 · 2021
Cited alongside, same era.
Persistent anti-muslim bias in large language models
Abubakar Abid, Maheen Farooqi, and James Zou. 2021b · 2021
Cited alongside, same era.
Mitigating language-dependent ethnic bias in BERT
Jaimeen Ahn and Alice Oh. 2021 · 2021
Cited alongside, same era.
AraGPT2: Pre-trained transformer for arabic language generation
Wissam Antoun, Fady Baly, and Hazem Hajj. 2021 · 2021
Cited alongside, same era.
Quantifying social biases in nlp: A generalization and empirical comparison of extrinsic fairness metrics
Paula Czarnowska, Yogarshi Vyas, and Kashif Shah. 2021 · 2021
Cited alongside, same era.
OSCaR: Orthogonal subspace correction and rectification of biases in word embeddings
Sunipa Dev, Tao Li, Jeff M Phillips, and Vivek Srikumar. 2021 · 2021
Bloom+1: Adding language support to bloom for zero-shot prompting
Zheng-Xin Yong, Hailey Schoelkopf, Niklas Muennighoff, Alham Fikri Aji, David Ifeoluwa Adelani, Khalid Almubarak, M Saiful Bari, Lintang Sutawika, Jungo Kasai, Ahmed Baruwa, et al. 2022 · 2022
Later among the works it cites.
Moral foundations of large language models
Marwa Abdulhai, Gregory Serapio-Garcia, Clément Crepy, Daria Valter, John Canny, and Natasha Jaques. 2023 · 2023
Closest in time.
Mega: Multilingual evaluation of generative ai
Kabir Ahuja, Rishav Hada, Millicent Ochieng, Prachi Jain, Harshita Diddee, Samuel Maina, Tanuja Ganu, Sameer Segal, Maxamed Axmed, Kalika Bali, et al. 2023 · 2023
Closest in time.
SODAPOP: Open-ended discovery of social biases in social commonsense reasoning models
Haozhe An, Zongxia Li, Jieyu Zhao, and Rachel Rudinger. 2023 · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Detecting emergent intersectional biases: Contextualized word embeddings contain a distribution of human-like biases
Wei Guo and Aylin Caliskan. 2021 · 2021
Cited alongside, same era.
World values survey: Round seven
Christian Haerpfer, Ronald Inglehart, Alejandro Moreno, Christian Welzel, Kseniya Kizilova, Jaime Diez-Medrano, Marta Lagos, Pippa Norris, E Ponarin, and B Puranen. 2021 · 2021
Cited alongside, same era.
LoRA: Low-rank adaptation of large language models
Edward J Hu, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen, et al. 2021 · 2021
Cited alongside, same era.
The interplay of variant, size, and task type in arabic pre-trained language models
Go Inoue, Bashar Alhafni, Nurpeiis Baimukan, Houda Bouamor, and Nizar Habash. 2021 · 2021
Cited alongside, same era.
HONEST: Measuring hurtful sentence completion in language models
Debora Nozza, Federico Bianchi, and Dirk Hovy. 2021 · 2021
Cited alongside, same era.
Measuring social biases in grounded vision and language embeddings
Candace Ross, Boris Katz, and Andrei Barbu. 2021 · 2021
Cited alongside, same era.
Arnav Arora, Lucie-Aimée Kaffee, and Isabelle Augenstein. 2023 · 2023
Closest in time.
Assessing cross-cultural alignment between ChatGPT and human societies: An empirical study
Yong Cao, Li Zhou, Seolhwa Lee, Laura Cabello, Min Chen, and Daniel Hershcovich. 2023 · 2023
Closest in time.
Toward cultural bias evaluation datasets: The case of bengali gender, religious, and national identity
Dipto Das, Shion Guha, and Bryan Semaan. 2023 · 2023
Closest in time.
Towards measuring the representation of subjective global opinions in language models
Esin Durmus, Karina Nyugen, Thomas I Liao, Nicholas Schiefer, Amanda Askell, Anton Bakhtin, Carol Chen, Zac Hatfield-Dodds, Danny Hernandez, Nicholas Joseph, et al. 2023 · 2023
Closest in time.
ORCA: A challenging benchmark for Arabic language understanding
AbdelRahim Elmadany, ElMoatez Billah Nagoudi, and Muhammad Abdul-Mageed. 2023 · 2023
Closest in time.
Shangbin Feng, Chan Young Park, Yuhan Liu, and Yulia Tsvetkov. 2023 · 2023
Closest in time.
Bias and fairness in large language models: A survey
Isabel O Gallegos, Ryan A Rossi, Joe Barrow, Md Mehrab Tanjim, Sungchul Kim, Franck Dernoncourt, Tong Yu, Ruiyi Zhang, and Nesreen K Ahmed. 2023 · 2023
Closest in time.
Jochen Hartmann, Jasper Schwenzow, and Maximilian Witte. 2023 · 2023
Closest in time.
Muslim-violence bias persists in debiased gpt models
Babak Hemmatian, Razan Baltaji, and Lav R Varshney. 2023 · 2023
Closest in time.
AceGPT, localizing large language models in Arabic
Huang Huang, Fei Yu, Jianqing Zhu, Xuening Sun, Hao Cheng, Dingjie Song, Zhihong Chen, Abdulmohsen Alharthi, Bang An, Ziche Liu, et al. 2023 · 2023
Closest in time.
Culturally aware natural language inference
Jing Huang and Diyi Yang. 2023 · 2023
Closest in time.
Amr Keleg and Walid Magdy. 2023 · 2023
Closest in time.
Hate speech classifiers are culturally insensitive
Nayeon Lee, Chani Jung, and Alice Oh. 2023 · 2023
Closest in time.
Comparing biases and the impact of multilingual training across multiple languages
Sharon Levy, Neha Anna John, Ling Liu, Yogarshi Vyas, Jie Ma, Yoshinari Fujinuma, Miguel Ballesteros, Vittorio Castelli, and Dan Roth. 2023 · 2023
Closest in time.
A survey on fairness in large language models
Yingji Li, Mengnan Du, Rui Song, Xin Wang, and Ying Wang. 2023 · 2023
Closest in time.
Intersectional stereotypes in large language models: Dataset and analysis
Weicheng Ma, Brian Chiang, Tong Wu, Lili Wang, and Soroush Vosoughi. 2023a · 2023
Closest in time.
Deciphering stereotypes in pre-trained language models
Weicheng Ma, Henry Scheible, Brian Wang, Goutham Veeramachaneni, Pratim Chowdhary, Alan Sun, Andrew Koulogeorge, Lili Wang, Diyi Yang, and Soroush Vosoughi. 2023b · 2023
Closest in time.
Social bias probing: Fairness benchmarking for language models
Marta Marchiori Manerba, Karolina Stańczak, Riccardo Guidotti, and Isabelle Augenstein. 2023 · 2023
Closest in time.
Reem I Masoud, Ziquan Liu, Martin Ferianc, Philip Treleaven, and Miguel Rodrigues. 2023 · 2023
Closest in time.
FORK: A bite-sized test set for probing culinary cultural biases in commonsense reasoning models
Shramay Palta and Rachel Rudinger. 2023 · 2023
Closest in time.
Knowledge of cultural moral norms in large language models
Aida Ramezani and Yang Xu. 2023 · 2023
Closest in time.
Whose opinions do language models reflect?
Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee, Percy Liang, and Tatsunori Hashimoto. 2023 · 2023
Closest in time.
Neha Sengupta, Sunil Kumar Sahu, Bokang Jia, Satheesh Katipomu, Haonan Li, Fajri Koto, Osama Mohammed Afzal, Samta Kamboj, Onkar Pandit, Rahul Pal, et al. 2023 · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023 · 2023
Closest in time.
On evaluating and mitigating gender biases in multilingual settings
Aniket Vashishtha, Kabir Ahuja, and Sunayana Sitaram. 2023 · 2023
Closest in time.
Nationality bias in text generation
Pranav Narayanan Venkit, Sanjana Gautam, Ruchi Panchanadikar, Ting-Hao Huang, and Shomir Wilson. 2023 · 2023
Closest in time.
“kelly is a warm person, joseph is a role model”: Gender biases in llm-generated reference letters
Yixin Wan, George Pu, Jiao Sun, Aparna Garimella, Kai-Wei Chang, and Nanyun Peng. 2023 · 2023
Closest in time.
Align on the fly: Adapting chatbot behavior to established norms
Chunpu Xu, Steffi Chern, Ethan Chern, Ge Zhang, Zekun Wang, Ruibo Liu, Jing Li, Jie Fu, and Pengfei Liu. 2023 · 2023
Closest in time.
A shocking amount of the web is machine translated: Insights from multi-way parallelism
Brian Thompson, Mehak Preet Dhaliwal, Peter Frisch, Tobias Domhan, and Marcello Federico. 2024 · 2024
Closest in time.
Aya model: An instruction finetuned open-access multilingual language model
Ahmet Üstün, Viraat Aryabumi, Zheng-Xin Yong, Wei-Yin Ko, Daniel D’souza, Gbemileke Onilude, Neel Bhandari, Shivalika Singh, Hui-Lee Ooi, Amr Kayid, et al. 2024 · 2024
Closest in time.
GeoMLAMA: Geo-diverse commonsense probing on multilingual pre-trained language models
Da Yin, Hritik Bansal, Masoud Monajatipoor, Liunian Harold Li, and Kai-Wei Chang. 2022 · 2055
Closest in time.