Fetching the paper…
Reading the bibliography…
Based on the foundation of Large Language Models (LLMs), Multilingual LLMs (MLLMs) have been developed to address the challenges faced in multilingual natural language processing, hoping to achieve knowledge transfer from high-resource languages to low-resource languages.
Language models are few-shot learners
Brown T, Mann B, Ryder N, Subbiah M, Kaplan J D, Dhariwal P, Neelakantan A, Shyam P, Sastry G, Askell A et al · 1901
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
Mnih V, Badia A P, Mirza M, Graves A, Lillicrap T, Harley T, Silver D, Kavukcuoglu K · 1937
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Srivastava N, Hinton G, Krizhevsky A, Sutskever I, Salakhutdinov R · 1958
Earlier work this paper cites.
Cross-lingual name tagging and linking for 282 languages
Pan X, Zhang B, May J, Nothman J, Knight K, Ji H · 1958
Earlier work this paper cites.
CrowS-pairs: A challenge dataset for measuring social biases in masked language models
Nangia N, Vania C, Bhalerao R, Bowman S R · 1967
Earlier work this paper cites.
A new algorithm for data compression
Gage P · 1994
Earlier work this paper cites.
Catastrophic forgetting in connectionist networks
French R M · 1999
Earlier work this paper cites.
Visually grounded reasoning across languages and cultures
Liu F, Bugliarello E, Ponti E M, Reddy S, Collier N, Elliott D · 2009
Earlier work this paper cites.
Japanese and korean voice search
Schuster M, Nakajima K · 2012
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
Mikolov T, Sutskever I, Chen K, Corrado G, Dean J · 2013
Earlier work this paper cites.
Glove: Global vectors for word representation
Pennington J, Socher R, Manning C D · 2014
Earlier work this paper cites.
OpenSubtitles2016: Extracting large parallel corpora from movie and TV subtitles
Lison P, Tiedemann J · 2016
Earlier work this paper cites.
Attention is all you need
Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez A N, Kaiser L, Polosukhin I · 2017
Earlier work this paper cites.
Transfer learning across low-resource, related languages for neural machine translation
Nguyen T Q, Chiang D · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman J, Wolski F, Dhariwal P, Radford A, Klimov O · 2017
Earlier work this paper cites.
Enriching word vectors with subword information
Bojanowski P, Grave E, Joulin A, Mikolov T · 2017
Earlier work this paper cites.
Semantics derived automatically from language corpora contain human-like biases
Caliskan A, Bryson J J, Narayanan A · 2017
Earlier work this paper cites.
Improving language understanding by generative pre-training
Radford A, Narasimhan K, Salimans T, Sutskever I · 2018
Earlier work this paper cites.
Word translation without parallel data
Lample G, Conneau A, Ranzato M, Denoyer L, Jégou H · 2018
Earlier work this paper cites.
Wikipedia monolingual corpora, 2018
2018
Earlier work this paper cites.
A robust self-learning method for fully unsupervised cross-lingual mappings of word embeddings
Artetxe M, Labaka G, Agirre E · 2018
Earlier work this paper cites.
On the limitations of unsupervised bilingual dictionary induction
Søgaard A, Ruder S, Vulić I · 2018
Earlier work this paper cites.
Norma: Neighborhood sensitive maps for multilingual word embeddings
Nakashole N · 2018
Earlier work this paper cites.
Gromov-wasserstein alignment of word embedding spaces
Alvarez-Melis D, Jaakkola T · 2018
Earlier work this paper cites.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Wang A, Singh A, Michael J, Hill F, Levy O, Bowman S · 2018
Earlier work this paper cites.
Gender bias in coreference resolution
Rudinger R, Naradowsky J, Leonard B, Van Durme B · 2018
Earlier work this paper cites.
Gender bias in coreference resolution: Evaluation and debiasing methods
Zhao J, Wang T, Yatskar M, Ordonez V, Chang K W · 2018
Earlier work this paper cites.
Examining gender and race bias in two hundred sentiment analysis systems
Kiritchenko S, Mohammad S · 2018
Earlier work this paper cites.
XNLI: Evaluating cross-lingual sentence representations
Conneau A, Rinott R, Lample G, Williams A, Bowman S, Schwenk H, Stoyanov V · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin J, Chang M W, Lee K, Toutanova K · 2019
Earlier work this paper cites.
Cross-lingual language model pretraining
Conneau A, Lample G · 2019
Earlier work this paper cites.
Vries d W, Cranenburgh v A, Bisazza A, Caselli T, Noord v G, Nissim M · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford A, Wu J, Child R, Luan D, Amodei D, Sutskever I · 2019
Earlier work this paper cites.
How multilingual is multilingual bert?
Pires T, Schlinger E, Garrette D · 2019
Earlier work this paper cites.
Is multilingual bert fluent in language generation?
Rönnqvist S, Kanerva J, Salakoski T, Ginter F · 2019
Earlier work this paper cites.
Cross-lingual alignment of contextual word embeddings, with applications to zero-shot dependency parsing
Schuster T, Ram O, Barzilay R, Globerson A · 2019
Earlier work this paper cites.
Revisiting adversarial autoencoder for unsupervised word translation with cycle consistency and improved training
Mohiuddin T, Joty S · 2019
Earlier work this paper cites.
Bert is not an interlingua and the bias of tokenization
Singh J, McCann B, Socher R, Xiong C · 2019
Earlier work this paper cites.
Multilingual word translation using auxiliary languages
Taitelbaum H, Chechik G, Goldberger J · 2019
Earlier work this paper cites.
Cross-lingual ability of multilingual bert: An empirical study
Karthikeyan K, Wang Z, Mayhew S, Roth D · 2019
Earlier work this paper cites.
On measuring social biases in sentence encoders
May C, Wang A, Bordia S, Bowman S R, Rudinger R · 2019
Earlier work this paper cites.
Evaluating gender bias in machine translation
Stanovsky G, Smith N A, Zettlemoyer L · 2019
Earlier work this paper cites.
Bias in bios: A case study of semantic representation bias in a high-stakes setting
De-Arteaga M, Romanov A, Wallach H M, Chayes J T, Borgs C, Chouldechova A, Geyik S C, Kenthapadi K, Kalai A T · 2019
Earlier work this paper cites.
Are we consistently biased? multidimensional analysis of biases in distributional word vectors
Lauscher A, Glavaš G · 2019
Earlier work this paper cites.
Gender bias in contextualized word embeddings
Zhao J, Wang T, Yatskar M, Cotterell R, Ordonez V, Chang K W · 2019
Earlier work this paper cites.
Unsupervised cross-lingual representation learning at scale
Conneau A, Khandelwal K, Goyal N, Chaudhary V, Wenzek G, Guzmán F, Grave É, Ott M, Zettlemoyer L, Stoyanov V · 2020
Earlier work this paper cites.
Multilingual alignment of contextual word representations
Cao S, Kitaev N, Klein D · 2020
Earlier work this paper cites.
Social biases in nlp models as barriers for persons with disabilities
Hutchinson B, Prabhakaran V, Denton E, Webster K, Zhong Y, Denuyl S · 2020
Earlier work this paper cites.
Flaubert: Unsupervised language model pre-training for french
Le H, Vial L, Frej J, Segonne V, Coavoux M, Lecouteux B, Allauzen A, Crabbé B, Besacier L, Schwab D · 2020
Earlier work this paper cites.
Arabert: Transformer-based model for arabic language understanding
Antoun W, Baly F, Hajj H · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel C, Shazeer N, Roberts A, Lee K, Narang S, Matena M, Zhou Y, Li W, Liu P J · 2020
Earlier work this paper cites.
Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension
Lewis M, Liu Y, Goyal N, Ghazvininejad M, Mohamed A, Levy O, Stoyanov V, Zettlemoyer L · 2020
Earlier work this paper cites.
Multilingual denoising pre-training for neural machine translation
Liu Y, Gu J, Goyal N, Li X, Edunov S, Ghazvininejad M, Lewis M, Zettlemoyer L · 2020
Earlier work this paper cites.
On the cross-lingual transferability of monolingual representations
Artetxe M, Ruder S, Yogatama D · 2020
Earlier work this paper cites.
Pre-trained models for natural language processing: A survey
Qiu X, Sun T, Xu Y, Shao Y, Dai N, Huang X · 2020
Earlier work this paper cites.
Learning to summarize with human feedback
Stiennon N, Ouyang L, Wu J, Ziegler D, Lowe R, Voss C, Radford A, Amodei D, Christiano P F · 2020
Earlier work this paper cites.
Extending multilingual bert to low-resource languages
Wang Z, Karthikeyan K, Mayhew S, Roth D · 2020
Earlier work this paper cites.
XED: A multilingual dataset for sentiment analysis and emotion detection
Öhman E, Pàmies M, Kajava K, Tiedemann J · 2020
Earlier work this paper cites.
CLIRMatrix: A massively large collection of bilingual and multilingual datasets for cross-lingual information retrieval
Sun S, Duh K · 2020
Earlier work this paper cites.
The multilingual Amazon reviews corpus
Keung P, Lu Y, Szarvas G, Smith N A · 2020
Earlier work this paper cites.
Probing pretrained language models for lexical semantics
Vulić I, Ponti E M, Litschko R, Glavaš G, Korhonen A · 2020
Earlier work this paper cites.
A graph-based coarse-to-fine method for unsupervised bilingual lexicon induction
Ren S, Liu S, Zhou M, Ma S · 2020
Earlier work this paper cites.
LNMap: Departures from isomorphic assumption in bilingual lexicon induction through non-linear mapping in latent space
Mohiuddin T, Bari M S, Joty S · 2020
Earlier work this paper cites.
Non-linear instance-based cross-lingual mapping for non-isomorphic embedding spaces
Glavaš G, Vulić I · 2020
Earlier work this paper cites.
A study of cross-lingual ability and language-specific information in multilingual bert
Liu C L, Hsu T Y, Chuang Y S, Lee H Y · 2020
Earlier work this paper cites.
Gender bias in multilingual embeddings and cross-lingual transfer
Zhao J, Mukherjee S, Hosseini S, Chang K, Awadallah A H · 2020
Earlier work this paper cites.
Are all languages created equal in multilingual bert?
Wu S, Dredze M · 2020
Cited alongside, same era.
Bertscore: Evaluating text generation with BERT
Zhang T, Kishore V, Wu F, Weinberger K Q, Artzi Y · 2020
Cited alongside, same era.
BLEURT: Learning robust metrics for text generation
Sellam T, Das D, Parikh A · 2020
Cited alongside, same era.
Towards debiasing sentence representations
Liang P P, Li I M, Zheng E, Lim Y C, Salakhutdinov R, Morency L P · 2020
Cited alongside, same era.
Null it out: Guarding protected attributes by iterative nullspace projection
Ravfogel S, Elazar Y, Gonen H, Twiton M, Goldberg Y · 2020
Cited alongside, same era.
Measuring and reducing gendered correlations in pre-trained models
Webster K, Wang X, Tenney I, Beutel A, Pitler E, Pavlick E, Chen J, Chi E, Petrov S · 2020
Bertscore is unfair: On social bias in language model-based metrics for text generation
Sun T, He J, Qiu X, Huang X · 2022
Later among the works it cites.
Measuring fairness with biased rulers: A comparative study on bias metrics for pre-trained language models
Delobelle P, Tokpo E, Calders T, Berendt B · 2022
Later among the works it cites.
French CrowS-pairs: Extending a challenge dataset for measuring social bias in masked language models to a language other than English
Névéol A, Dupont Y, Bezançon J, Fort K · 2022
Later among the works it cites.
Auto-debias: Debiasing masked language models with automated biased prompts
Guo Y, Yang Y, Abbasi A · 2022
Later among the works it cites.
Understanding stereotypes in language models: Towards robust measurement and zero-shot debiasing
Chen Y, Raghuram V C, Mattern J, Mihalcea R, Jin Z · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Xtreme: A massively multilingual multi-task benchmark for evaluating cross-lingual generalisation
Hu J, Ruder S, Siddhant A, Neubig G, Firat O, Johnson M · 2020
Cited alongside, same era.
Identifying elements essential for BERT’s multilinguality
Dufter P, Schütze H · 2020
Cited alongside, same era.
mt5: A massively multilingual pre-trained text-to-text transformer
Xue L, Constant N, Roberts A, Kale M, Al-Rfou R, Siddhant A, Barua A, Raffel C · 2021
Cited alongside, same era.
On the dangers of stochastic parrots: Can language models be too big?��
Bender E M, Gebru T, McMillan-Major A, Shmitchell S · 2021
Cited alongside, same era.
Stereoset: Measuring stereotypical bias in pretrained language models
Nadeem M, Bethke A, Reddy S · 2021
Cited alongside, same era.
How good is your tokenizer? on the monolingual performance of multilingual language models
Rust P, Pfeiffer J, Vulić I, Ruder S, Gurevych I · 2021
Cited alongside, same era.
The bigscience roots corpus: A 1.6tb composite multilingual dataset
Laurençon H, Saulnier L, Wang T, Akiki C, Moral d A V, Scao T L, Werra L V, Mou C, Ponferrada E G, Nguyen H et al · 2022
Later among the works it cites.
Quality at a glance: An audit of web-crawled multilingual datasets
Kreutzer J, Caswell I, Wang L, Wahab A, Esch v D, Ulzii-Orshikh N, Tapo A, Subramani N, Sokolov A, Sikasote C et al · 2022
Later among the works it cites.
Counterfactually augmented data and unintended bias: The case of sexism and hate speech detection
Sen I, Samory M, Wagner C, Augenstein I · 2022
Later among the works it cites.
An investigation of the (in)effective-ness of counterfactually augmented data
Joshi N, He H · 2022
Later among the works it cites.
A survey of multilingual models for automatic speech recognition
Yadav H, Sitaram S · 2022
Later among the works it cites.
KinyaBERT: a morphology-aware Kinyarwanda language model
Nzeyimana A, Niyongabo Rubungo A · 2022
Later among the works it cites.
Llama: Open and efficient foundation language models
Touvron H, Lavril T, Izacard G, Martinet X, Lachaux M A, Lacroix T, Rozière B, Goyal N, Hambro E, Azhar F et al · 2023
Later among the works it cites.
OpenAI , Achiam J, Adler S, Agarwal S, Ahmad L, Akkaya I, Aleman F L, Almeida D, Altenschmidt J, Altman S et al · 2023
Later among the works it cites.
Palm: Scaling language modeling with pathways
Chowdhery A, Narang S, Devlin J, Bosma M, Mishra G, Roberts A, Barham P, Chung H W, Sutton C, Gehrmann S et al · 2023
Later among the works it cites.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
Chiang W L, Li Z, Lin Z, Sheng Y, Wu Z, Zhang H, Zheng L, Zhuang S, Zhuang Y, Gonzalez J E et al · 2023
Later among the works it cites.
Gemini: a family of highly capable multimodal models
Team G, Anil R, Borgeaud S, Alayrac J B, Yu J, Soricut R, Schalkwyk J, Dai A M, Hauth A, Millican K et al · 2023
Later among the works it cites.
Palm-e: An embodied multimodal language model
Driess D, Xia F, Sajjadi M S M, Lynch C, Chowdhery A, Ichter B, Wahid A, Tompson J, Vuong Q, Yu T et al · 2023
Later among the works it cites.
Stanford alpaca: An instruction-following llama model
Taori R, Gulrajani I, Zhang T, Dubois Y, Li X, Guestrin C, Liang P, Hashimoto T B · 2023
Later among the works it cites.
Ren X, Zhou P, Meng X, Huang X, Wang Y, Wang W, Li P, Zhang X, Podolskiy A, Arshinov G et al · 2023
Later among the works it cites.
Pythia: A suite for analyzing large language models across training and scaling
Biderman S, Schoelkopf H, Anthony Q G, Bradley H, O’Brien K, Hallahan E, Khan M A, Purohit S, Prashanth U S, Raff E et al · 2023
Later among the works it cites.
Anil R, Dai A M, Firat O, Johnson M, Lepikhin D, Passos A, Shakeri S, Taropa E, Bailey P, Chen Z et al · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Touvron H, Martin L, Stone K, Albert P, Almahairi A, Babaei Y, Bashlykov N, Batra S, Bhargava P, Bhosale S et al · 2023
Later among the works it cites.
An overview of bard: an early experiment with generative ai
Manyika J, Hsiao S · 2023
Later among the works it cites.
Baichuan 2: Open large-scale language models
Yang A, Xiao B, Wang B, Zhang B, Bian C, Yin C, Lv C, Pan D, Wang D, Yan D et al · 2023
Later among the works it cites.
Phi-2: The surprising power of small language models
MICROSOFT · 2023
Later among the works it cites.
A survey of large language models
Zhao W X, Zhou K, Li J, Tang T, Wang X, Hou Y, Min Y, Zhang B, Zhang J, Dong Z et al · 2023
Later among the works it cites.
Large language model alignment: A survey
Shen T, Jin R, Huang Y, Liu C, Dong W, Guo Z, Wu X, Liu Y, Xiong D · 2023
Later among the works it cites.
Improving language models with advantage-based offline policy gradients
Baheti A, Lu X, Brahman F, Bras R L, Sap M, Riedl M · 2023
Later among the works it cites.
Aligning language models with preferences through f-divergence minimization
Go D, Korbak T, Kruszewski G, Rozen J, Ryu N, Dymetman M · 2023
Later among the works it cites.
Named entity recognition for low-resource languages-profiting from language families
Torge S, Politov A, Lehmann C, Saffar B, Tao Z · 2023
Later among the works it cites.
How do languages influence each other? studying cross-lingual data sharing during LM fine-tuning
Choenni R, Garrette D, Shutova E · 2023
Later among the works it cites.
Exploring vision-language models for imbalanced learning
Wang Y, Yu Z, Wang J, Heng Q, Chen H, Ye W, Xie R, Xie X, Zhang S · 2023
Later among the works it cites.
Balanced and explainable social media analysis for public health with large language models
Jiang Y, Qiu R, Zhang Y, Zhang P F · 2023
Later among the works it cites.
Mini but mighty: Efficient multilingual pretraining with linguistically-informed data selection
Ogunremi T, Jurafsky D, Manning C D · 2023
Later among the works it cites.
Towards a common understanding of contributing factors for cross-lingual transfer in multilingual language models: A review
Philippy F, Guo S, Haddadan S · 2023
Later among the works it cites.
Yayi 2: Multilingual open-source large language models
Luo Y, Kong Q, Xu N, Cao J, Hao B, Qu B, Chen B, Zhu C, Zhao C, Zhang D et al · 2023
Later among the works it cites.
NollySenti: Leveraging transfer learning and machine translation for Nigerian movie sentiment classification
Shode I, Adelani D I, Peng J, Feldman A · 2023
Later among the works it cites.
Taxi1500: A multilingual dataset for text classification in 1500 languages
Ma C, ImaniGooghari A, Ye H, Asgari E, Schütze H · 2023
Later among the works it cites.
Should chatgpt be biased? challenges and risks of bias in large language models
Ferrara E · 2023
Later among the works it cites.
Comparing biases and the impact of multilingual training across multiple languages
Levy S, John N A, Liu L, Vyas Y, Ma J, Fujinuma Y, Ballesteros M, Castelli V, Roth D · 2023
Later among the works it cites.
Towards explainable evaluation metrics for machine translation
Leiter C, Lertvittayakumjorn P, Fomicheva M, Zhao W, Gao Y, Eger S · 2023
Later among the works it cites.
“kelly is a warm person, joseph is a role model”: Gender biases in LLM-generated reference letters
Wan Y, Pu G, Sun J, Garimella A, Chang K W, Peng N · 2023
Later among the works it cites.
Causal-debias: Unifying debiasing in pretrained language models and fine-tuning via causal invariant learning
Zhou F, Mao Y, Yu L, Yang Y, Zhong T · 2023
Later among the works it cites.
Overwriting pretrained bias with finetuning data
Wang A, Russakovsky O · 2023
Later among the works it cites.
Queer people are people first: Deconstructing sexual identity stereotypes in large language models
Dhingra H, Jayashanker P, Moghe S, Strubell E · 2023
Later among the works it cites.
People make better edits: Measuring the efficacy of LLM-generated counterfactually augmented data for harmful language detection
Sen I, Assenmacher D, Samory M, Augenstein I, Aalst W, Wagner C · 2023
Later among the works it cites.
Bias beyond English: Counterfactual tests for bias in sentiment analysis in four languages
Goldfarb-Tarrant S, Lopez A, Blanco R, Marcheggiani D · 2023
Later among the works it cites.
A comprehensive overview of large language models
Naveed H, Khan A U, Qiu S, Saqib M, Anwar S, Usman M, Akhtar N, Barnes N, Mian A · 2023
Later among the works it cites.
MM-LLMs: Recent advances in MultiModal large language models
Zhang D, Yu Y, Dong J, Li C, Su D, Chu C, Yu D · 2024
Closest in time.
Chatglm: A family of large language models from glm-130b to glm-4 all tools
GLM T, Zeng A, Xu B, Wang B, Zhang C, Yin D, Zhang D, Rojas D, Feng G, Zhao H et al · 2024
Closest in time.
The claude 3 model family: Opus, sonnet, haiku
Anthropic · 2024
Closest in time.
Dubey A, Jauhri A, Pandey A, Kadian A, Al-Dahle A, Letman A, Mathur A, Schelten A, Yang A, Fan A et al · 2024
Closest in time.
The refinedweb dataset for falcon llm: outperforming curated corpora with web data only
Penedo G, Malartic Q, Hesslow D, Cojocaru R, Alobeidli H, Cappelli A, Pannier B, Almazrouei E, Launay J · 2024
Closest in time.
FuxiTranyu: A multilingual large language model trained with balanced data
Sun H, Jin R, Xu S, Pan L, Supryadi , Cui M, Du J, Lei Y, Yang L, Shi L et al · 2024
Closest in time.
Multilingual machine translation with large language models: Empirical results and analysis
Zhu W, Liu H, Dong Q, Xu J, Huang S, Kong L, Chen J, Li L · 2024
Closest in time.
A comprehensive survey of bias in llms: Current landscape and future directions
Ranjan R, Gupta S, Singh S N · 2024
Closest in time.
Agr: Age group fairness reward for bias mitigation in llms
Cao S, Cheng R, Wang Z · 2024
Closest in time.
Having beer after prayer? measuring cultural bias in large language models
Naous T, Ryan M J, Ritter A, Xu W · 2024
Closest in time.
Benchmarking cognitive biases in large language models as evaluators
Koo R, Lee M, Raheja V, Park J I, Kim Z M, Kang D · 2024
Closest in time.
A trip towards fairness: Bias and de-biasing in large language models
Ranaldi L, Ruzzetti E S, Venditti D, Onorati D, Zanzotto F M · 2024
Closest in time.
CulturaX: A cleaned, enormous, and multilingual dataset for large language models in 167 languages
Nguyen T, Nguyen C V, Lai V D, Man H, Ngo N T, Dernoncourt F, Rossi R A, Nguyen T H · 2024
Closest in time.
Exploring accuracy-fairness trade-off in large language models
Zhang Q, Duan Q, Yuan B, Shi Y, Liu J · 2024
Closest in time.
Towards trustworthy llms: a review on debiasing and dehallucinating in large language models
Lin Z, Guan S, Zhang W, Zhang H, Li Y, Zhang H · 2024
Closest in time.
Mitigating biases for instruction-following language models via bias neurons elimination
Yang N, Kang T, Choi S J, Lee H, Jung K · 2024
Closest in time.
Multilingual open text release 1: Public domain news in 44 languages
Palen-Michel C, Kim J, Lignos C · 2089
Closest in time.