Fetching the paper…
Reading the bibliography…
In this work, we introduce EMMA-500, a large-scale multilingual language model continue-trained on texts across 546 languages designed for enhanced multilingual performance, focusing on improving language coverage for low-resource languages.
Bertscore: Evaluating text generation with bert
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. 2019 · 1904
Earlier work this paper cites.
The md5 message-digest algorithm
Ronald Rivest. 1992 · 1992
Earlier work this paper cites.
On the resemblance and containment of documents
Andrei Z Broder. 1997 · 1997
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Named entity recognition for south and south East Asian languages: Taking stock
Anil Kumar Singh. 2008 · 2008
Earlier work this paper cites.
W2C – web to corpus – corpora
Martin Majliš. 2011 · 2011
Earlier work this paper cites.
The Kyoto free translation task
Graham Neubig. 2011 · 2011
Earlier work this paper cites.
Mining of massive datasets
Anand Rajaraman and Jeffrey D Ullman. 2011 · 2011
Earlier work this paper cites.
Building large monolingual dictionaries at the Leipzig corpora collection: From 100 to 200 languages
Dirk Goldhahn, Thomas Eckart, and Uwe Quasthoff. 2012 · 2012
Earlier work this paper cites.
The Winograd schema challenge
Hector Levesque, Ernest Davis, and Leora Morgenstern. 2012 · 2012
Earlier work this paper cites.
Parallel data, tools and interfaces in OPUS
Jörg Tiedemann. 2012 · 2012
Earlier work this paper cites.
Evenki life newspaper
Evenki Life. 2014 · 2014
Earlier work this paper cites.
Creating a massively parallel bible corpus
Thomas Mayer and Michael Cysouw. 2014 · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba. 2015 · 2015
Earlier work this paper cites.
chrF: character n-gram F-score for automatic MT evaluation
Maja Popović. 2015 · 2015
Earlier work this paper cites.
A large-scale multilingual disambiguation of glosses
José Camacho-Collados, Claudio Delli Bovi, Alessandro Raganato, and Roberto Navigli. 2016 · 2016
Earlier work this paper cites.
Lsdsem 2017 shared task: The story cloze test
Nasrin Mostafazadeh, Michael Roth, Annie Louis, Nathanael Chambers, and James Allen. 2017 · 2017
Earlier work this paper cites.
Tilde MODEL - multilingual open data for EU languages
Roberts Rozis and Raivis Skadiņš. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Parallel corpora for bi-lingual English-Ethiopian languages statistical machine translation
Solomon Teferra Abate, Michael Melese, Martha Yifiru Tachbelie, Million Meshesha, Solomon Atinafu, Wondwossen Mulugeta, Yaregal Assabie, Hafte Abera, Binyam Ephrem, Tewodros Abebe, Wondimagegnhue Tsegaye, Amanuel Lemma, Tsegaye Andargie, and Seifedin Shifaw. 2018 · 2018
Earlier work this paper cites.
Shami: A corpus of Levantine Arabic dialects
Kathrein Abu Kwaik, Motaz Saad, Stergios Chatzikyriakidis, and Simon Dobnik. 2018 · 2018
Earlier work this paper cites.
Developing new linguistic resources and tools for the Galician language
Rodrigo Agerri, Xavier Gómez Guinovart, German Rigau, and Miguel Anxo Solla Portela. 2018 · 2018
Earlier work this paper cites.
DART: A large dataset of dialectal Arabic tweets
Israa Alsarsour, Esraa Mohamed, Reem Suwaileh, and Tamer Elsayed. 2018 · 2018
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. 2018 · 2018
Earlier work this paper cites.
Xnli: Evaluating cross-lingual sentence representations
Alexis Conneau, Ruty Rinott, Guillaume Lample, Adina Williams, Samuel R. Bowman, Holger Schwenk, and Veselin Stoyanov. 2018 · 2018
Earlier work this paper cites.
Arabic dialect identification in the context of bivalency and code-switching
Mahmoud El-Haj, Paul Rayson, and Mariam Aboelezz. 2018 · 2018
Earlier work this paper cites.
The IIT Bombay English-Hindi parallel corpus
Anoop Kunchukuttan, Pratik Mehta, and Pushpak Bhattacharyya. 2018 · 2018
Earlier work this paper cites.
A call for clarity in reporting BLEU scores
Matt Post. 2018 · 2018
Earlier work this paper cites.
Texygen: A benchmarking platform for text generation models
Yaoming Zhu, Sidi Lu, Lei Zheng, Jiaxian Guo, Weinan Zhang, Jun Wang, and Yong Yu. 2018 · 2018
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
Asynchronous pipeline for processing huge corpora on medium to low resource infrastructures
Pedro Javier Ortiz Suárez, Benoît Sagot, and Laurent Romary. 2019 · 2019
Earlier work this paper cites.
Huggingface’s transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, et al. 2019 · 2019
Earlier work this paper cites.
TICO-19: the translation initiative for COvid-19
Antonios Anastasopoulos, Alessandro Cattelan, Zi-Yi Dou, Marcello Federico, Christian Federmann, Dmitriy Genzel, Franscisco Guzmán, Junjie Hu, Macduff Hughes, Philipp Koehn, Rosie Lazar, Will Lewis, Graham Neubig, Mengmeng Niu, Alp Öktem, Eric Paquin, Grace Tang, and Sylwia Tur. 2020 · 2020
Earlier work this paper cites.
ParaCrawl: Web-scale acquisition of parallel corpora
Marta Bañón, Pinzhen Chen, Barry Haddow, Kenneth Heafield, Hieu Hoang, Miquel Esplà-Gomis, Mikel L. Forcada, Amir Kamran, Faheem Kirefu, Philipp Koehn, Sergio Ortiz Rojas, Leopoldo Pla Sempere, Gema Ramírez-Sánchez, Elsa Sarrías, Marek Strelec, Brian Thompson, William Waites, Dion Wiggins, and Jaume Zaragoza. 2020 · 2020
Earlier work this paper cites.
Unsupervised cross-lingual representation learning at scale
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Édouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020 · 2020
Earlier work this paper cites.
Mapping languages: the corpus of global language use
Jonathan Dunn. 2020 · 2020
Earlier work this paper cites.
Habibi - a multi dialect multi national Arabic song lyrics corpus
Mahmoud El-Haj. 2020 · 2020
Earlier work this paper cites.
IndicNLPSuite: Monolingual corpora, evaluation benchmarks and pre-trained multilingual language models for Indian languages
Divyanshu Kakwani, Anoop Kunchukuttan, Satish Golla, Gokul N.C., Avik Bhattacharyya, Mitesh M. Khapra, and Pratyush Kumar. 2020 · 2020
Earlier work this paper cites.
Towards computational linguistics in Minangkabau language: Studies on sentiment analysis and machine translation
Fajri Koto and Ikhwan Koto. 2020 · 2020
Earlier work this paper cites.
S2ORC: The semantic scholar open research corpus
Kyle Lo, Lucy Lu Wang, Mark Neumann, Rodney Kinney, and Daniel Weld. 2020 · 2020
Earlier work this paper cites.
JParaCrawl: A large scale web-based English-Japanese parallel corpus
Makoto Morishita, Jun Suzuki, and Masaaki Nagata. 2020 · 2020
Earlier work this paper cites.
XCOPA: A multilingual dataset for causal commonsense reasoning
Edoardo Maria Ponti, Goran Glavaš, Olga Majewska, Qianchu Liu, Ivan Vulić, and Anna Korhonen. 2020 · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020 · 2020
Earlier work this paper cites.
AraBench: Benchmarking dialectal Arabic-English machine translation
Hassan Sajjad, Ahmed Abdelali, Nadir Durrani, and Fahim Dalvi. 2020 · 2020
Cited alongside, same era.
The Tatoeba Translation Challenge – Realistic data sets for low resource and multilingual MT
Jörg Tiedemann. 2020 · 2020
Cited alongside, same era.
CCNet: Extracting high quality monolingual datasets from web crawl data
Guillaume Wenzek, Marie-Anne Lachaux, Alexis Conneau, Vishrav Chaudhary, Francisco Guzmán, Armand Joulin, and Edouard Grave. 2020 · 2020
Cited alongside, same era.
Indonlu: Benchmark and resources for evaluating indonesian natural language understanding
Bryan Wilie, Karissa Vincentio, Genta Indra Winata, Samuel Cahyawijaya, X. Li, Zhi Yuan Lim, S. Soleman, R. Mahendra, Pascale Fung, Syafri Bahar, and A. Purwarianti. 2020 · 2020
Cited alongside, same era.
QADI: Arabic dialect identification in the wild
Ahmed Abdelali, Hamdy Mubarak, Younes Samih, Sabit Hassan, and Kareem Darwish. 2021 · 2021
Cited alongside, same era.
Dataset card for ”project gutenberg”
Manuel Faysse. 2023 · 2023
Later among the works it cites.
A framework for few-shot language model evaluation
Leo Gao, Jonathan Tow, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Kyle McDonell, Niklas Muennighoff, et al. 2023 · 2023
Later among the works it cites.
Glot500: Scaling multilingual corpora and language models to 500 languages
Ayyoob Imani, Peiqin Lin, Amir Hossein Kargaran, Silvia Severini, Masoud Jalili Sabet, Nora Kassner, Chunlan Ma, Helmut Schmid, André F. T. Martins, François Yvon, and Hinrich Schütze. 2023 · 2023
Later among the works it cites.
GlotLID: Language identification for low-resource languages
Amir Hossein Kargaran, Ayyoob Imani, François Yvon, and Hinrich Schütze. 2023 · 2023
Later among the works it cites.
The stack: 3 TB of permissively licensed source code
Denis Kocetkov, Raymond Li, Loubna Ben Allal, Jia Li, Chenghao Mou, Yacine Jernite, Margaret Mitchell, Carlos Muñoz Ferrandis, Sean Hughes, Thomas Wolf, Dzmitry Bahdanau, Leandro von Werra, and Harm de Vries. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
The effect of domain and diacritics in Yoruba–English neural machine translation
David Adelani, Dana Ruiter, Jesujoba Alabi, Damilola Adebonojo, Adesina Ayeni, Mofe Adeyemi, Ayodele Esther Awokoya, and Cristina España-Bonet. 2021 · 2021
Cited alongside, same era.
Evaluating large language models trained on code
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde De Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al. 2021 · 2021
Cited alongside, same era.
Training verifiers to solve math word problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. 2021 · 2021
Cited alongside, same era.
Tlmd: Tigrinya language modeling dataset (1.0.0)
Fitsum Gaim, Wonsuk Yang, and Jong C. Park. 2021 · 2021
Cited alongside, same era.
Experiments on a Guarani corpus of news and social media
Santiago Góngora, Nicolás Giossa, and Luis Chiruzzo. 2021 · 2021
Cited alongside, same era.
Many-to-English machine translation tools, data, and pretrained models
Thamme Gowda, Zhao Zhang, Chris Mattmann, and Jonathan May. 2021 · 2021
Cited alongside, same era.
XL-sum: Large-scale multilingual abstractive summarization for 44 languages
Tahmid Hasan, Abhik Bhattacharjee, Md. Saiful Islam, Kazi Mubasshir, Yuan-Fang Li, Yong-Bin Kang, M. Sohel Rahman, and Rifat Shahriyar. 2021 · 2021
Cited alongside, same era.
Efficient memory management for large language model serving with pagedattention
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, and Ion Stoica. 2023 · 2023
Later among the works it cites.
Okapi: Instruction-tuned large language models in multiple languages with reinforcement learning from human feedback
Viet Dac Lai, Chien Van Nguyen, Nghia Trung Ngo, Thuat Nguyen, Franck Dernoncourt, Ryan A. Rossi, and Thien Huu Nguyen. 2023 · 2023
Later among the works it cites.
Colossal-ai: A unified deep learning system for large-scale parallel training
Shenggui Li, Hongxin Liu, Zhengda Bian, Jiarui Fang, Haichen Huang, Yuliang Liu, Boxiang Wang, and Yang You. 2023 · 2023
Later among the works it cites.
An empirical study of catastrophic forgetting in large language models during continual fine-tuning
Yun Luo, Zhen Yang, Fandong Meng, Yafu Li, Jie Zhou, and Yue Zhang. 2023 · 2023
Later among the works it cites.
Taxi1500: A multilingual dataset for text classification in 1500 languages
Chunlan Ma, Ayyoob ImaniGooghari, Haotian Ye, Ehsaneddin Asgari, and Hinrich Schütze. 2023 · 2023
Later among the works it cites.
Lacuna project
Masakhane. 2023 · 2023
Later among the works it cites.
Chenghaomou/text-dedup: Reference snapshot
Chenghao Mou, Chris Ha, Kenneth Enevoldsen, and Peiyuan Liu. 2023 · 2023
Later among the works it cites.
Culturax: A cleaned, enormous, and multilingual dataset for large language models in 167 languages
Thuat Nguyen, Chien Van Nguyen, Viet Dac Lai, Hieu Man, Nghia Trung Ngo, Franck Dernoncourt, Ryan A. Rossi, and Thien Huu Nguyen. 2023 · 2023
Later among the works it cites.
Oscar (open super-large crawled aggregated corpus) 2301
OSCAR. 2023 · 2023
Later among the works it cites.
Code llama: Open foundation models for code
Baptiste Roziere, Jonas Gehring, Fabian Gloeckle, Sten Sootla, Itai Gat, Xiaoqing Ellen Tan, Yossi Adi, Jingyu Liu, Tal Remez, Jérémy Rapin, et al. 2023 · 2023
Later among the works it cites.
peS2o (Pretraining Efficiently on S2ORC) Dataset
Luca Soldaini and Kyle Lo. 2023 · 2023
Later among the works it cites.
Code translation with compiler representations
Marc Szafraniec, Baptiste Rozière, Hugh Leather, Patrick Labatut, François Charton, and Gabriel Synnaeve. 2023 · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023 · 2023
Later among the works it cites.
All languages matter: On the multilingual safety of large language models
Wenxuan Wang, Zhaopeng Tu, Chang Chen, Youliang Yuan, Jen-tse Huang, Wenxiang Jiao, and Michael R Lyu. 2023 · 2023
Later among the works it cites.
The multilingual alignment prism: Aligning global and local preferences to reduce harm
Aakanksha, Arash Ahmadian, Beyza Ermis, Seraphina Goldfarb-Tarrant, Julia Kreutzer, Marzieh Fadaee, and Sara Hooker. 2024 · 2024
Closest in time.
Tower: An open multilingual large language model for translation-related tasks
Duarte M Alves, José Pombal, Nuno M Guerreiro, Pedro H Martins, João Alves, Amin Farajian, Ben Peters, Ricardo Rei, Patrick Fernandes, Sweta Agrawal, et al. 2024 · 2024
Closest in time.
Aya 23: Open weight releases to further multilingual progress
Viraat Aryabumi, John Dang, Dwarak Talupuru, Saurabh Dash, David Cairuz, Hangyu Lin, Bharat Venkitesh, Madeline Smith, Jon Ander Campos, Yi Chern Tan, et al. 2024 · 2024
Closest in time.
Pinzhen Chen, Simon Yu, Zhicheng Guo, and Barry Haddow. 2024 · 2024
Closest in time.
A new massive multilingual dataset for high-performance language technologies
Ona de Gibert, Graeme Nail, Nikolay Arefyev, Marta Bañón, Jelmer van der Linde, Shaoxiong Ji, Jaume Zaragoza-Bernabeu, Mikko Aulamo, Gema Ramírez-Sánchez, Andrey Kutuzov, et al. 2024 · 2024
Closest in time.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. 2024 · 2024
Closest in time.
Investigating the translation capabilities of large language models trained on parallel data only
Javier García Gilabert, Carlos Escolano, Aleix Sant Savall, Francesca De Luca Fornaciari, Audrey Mash, Xixian Liao, and Maite Melero. 2024 · 2024
Closest in time.
Simple and scalable strategies to continually pre-train large language models
Adam Ibrahim, Benjamin Thérien, Kshitij Gupta, Mats L. Richter, Quentin Anthony, Timothée Lesort, Eugene Belilovsky, and Irina Rish. 2024 · 2024
Closest in time.
Glotscript: A resource and tool for low resource writing system identification
Amir Hossein Kargaran, François Yvon, and Hinrich Schütze. 2024 · 2024
Closest in time.
Code pretraining improves entity tracking abilities of language models
Najoung Kim, Sebastian Schuster, and Shubham Toshniwal. 2024 · 2024
Closest in time.
Minato Kondo, Takehito Utsuro, and Masaaki Nagata. 2024 · 2024
Closest in time.
Madlad-400: A multilingual and document-level large audited dataset
Sneha Kudugunta, Isaac Caswell, Biao Zhang, Xavier Garcia, Derrick Xin, Aditya Kusupati, Romi Stella, Ankur Bapna, and Orhan Firat. 2024 · 2024
Closest in time.
LLMs beyond English: Scaling the multilingual capability of LLMs with cross-lingual feedback
Wen Lai, Mohsen Mesgar, and Alexander Fraser. 2024 · 2024
Closest in time.
MaLA-500: Massive language adaptation of large language models
Peiqin Lin, Shaoxiong Ji, Jörg Tiedemann, André FT Martins, and Hinrich Schütze. 2024 · 2024
Closest in time.
Starcoder 2 and the stack v2: The next generation
Anton Lozhkov, Raymond Li, Loubna Ben Allal, Federico Cassano, Joel Lamy-Poirier, Nouamane Tazi, Ao Tang, Dmytro Pykhtar, Jiawei Liu, Yuxiang Wei, et al. 2024 · 2024
Closest in time.
Yinquan Lu, Wenhao Zhu, Lei Li, Yu Qiao, and Fei Yuan. 2024 · 2024
Closest in time.
At which training stage does code data help llms reasoning?
Yingwei Ma, Yue Liu, Yue Yu, Yuanliang Zhang, Yu Jiang, Changjian Wang, and Shanshan Li. 2024 · 2024
Closest in time.
Taishi Nakamura, Mayank Mishra, Simone Tedeschi, Yekun Chai, Jason T Stillerman, Felix Friedrich, Prateek Yadav, Tanmay Laud, Vu Minh Chien, Terry Yue Zhuo, et al. 2024 · 2024
Closest in time.
Ircoder: Intermediate representations make language models robust multilingual code generators
Indraneil Paul, Goran Glavas, and Iryna Gurevych. 2024 · 2024
Closest in time.
Aya dataset: An open-access collection for multilingual instruction tuning
Shivalika Singh, Freddie Vargus, Daniel Dsouza, Börje F Karlsson, Abinaya Mahendiran, Wei-Yin Ko, Herumb Shandilya, Jay Patel, Deividas Mataciunas, Laura OMahony, et al. 2024 · 2024
Closest in time.
Gemma 2: Improving open language models at a practical size
Gemma Team, Morgane Riviere, Shreya Pathak, Pier Giuseppe Sessa, Cassidy Hardin, Surya Bhupatiraju, Léonard Hussenot, Thomas Mesnard, Bobak Shahriari, Alexandre Ramé, et al. 2024 · 2024
Closest in time.
Exploring design choices for building language-specific llms
Atula Tejaswi, Nilesh Gupta, and Eunsol Choi. 2024 · 2024
Closest in time.
Culturay: A large cleaned multilingual dataset of 75 languages
Huu Nguyen Thuat Nguyen and Thien Nguyen. 2024 · 2024
Closest in time.
Aya model: An instruction finetuned open-access multilingual language model
Ahmet Üstün, Viraat Aryabumi, Zheng-Xin Yong, Wei-Yin Ko, Daniel D’souza, Gbemileke Onilude, Neel Bhandari, Shivalika Singh, Hui-Lee Ooi, Amr Kayid, et al. 2024 · 2024
Closest in time.