Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have significantly advanced natural language processing, but their progress has yet to be equal across languages.
Hellaswag: Can a machine really finish your sentence?
Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y. (2019) · 1905
Earlier work this paper cites.
Energy and policy considerations for deep learning in nlp
Strubell, E., Ganesh, A., and McCallum, A. (2019) · 1906
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V. (2019) · 1907
Earlier work this paper cites.
Multifit: Efficient multi-lingual language model fine-tuning
Eisenschlos, J. M., Ruder, S., Czapla, P., Kardas, M., Gugger, S., and Howard, J. (2019) · 1909
Earlier work this paper cites.
Megatron-lm: Training multi-billion parameter language models using model parallelism
Shoeybi, M., Patwary, M., Puri, R., LeGresley, P., Casper, J., and Catanzaro, B. (2019) · 1909
Earlier work this paper cites.
Quantifying the carbon emissions of machine learning
Lacoste, A., Luccioni, A., Schmidt, V., and Dandres, T. (2019) · 1910
Earlier work this paper cites.
Unsupervised cross-lingual representation learning at scale
Conneau, A., Khandelwal, K., Goyal, N., Chaudhary, V., Wenzek, G., Guzmán, F., Grave, E., Ott, M., Zettlemoyer, L., and Stoyanov, V. (2019) · 1911
Earlier work this paper cites.
Energy usage reports: Environmental awareness as part of algorithmic accountability
Lottick, K., Susai, S., Friedler, S. A., and Wilson, J. P. (2019) · 1911
Earlier work this paper cites.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D. (2020) · 2001
Earlier work this paper cites.
Glu variants improve transformer
Shazeer, N. (2020) · 2002
Earlier work this paper cites.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J. (2020) · 2009
Earlier work this paper cites.
mt5: A massively multilingual pre-trained text-to-text transformer
Xue, L., Constant, N., Roberts, A., Kale, M., Al-Rfou, R., Siddhant, A., Barua, A., and Raffel, C. (2020) · 2010
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Maas, A. L., Daly, R. E., Pham, P. T., Huang, D., Ng, A. Y., and Potts, C. (2011) · 2011
Earlier work this paper cites.
Gottbert: a pure german language model
Scheible, R., Thomczyk, F., Tippmann, P., Jaravine, V., and Boeker, M. (2020) · 2012
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J. (2014) · 2014
Earlier work this paper cites.
Character-level convolutional networks for text classification
Zhang, X., Zhao, J. J., and LeCun, Y. (2015) · 2015
Earlier work this paper cites.
Training deep nets with sublinear memory cost
Chen, T., Xu, B., Zhang, C., and Guestrin, C. (2016) · 2016
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F. (2017) · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. (2017) · 2017
Earlier work this paper cites.
Weights & biases: A tool for visualizing and tracking your machine learning experiments
Weights&Biases (2017) · 2017
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge
Clark, P., Cowhey, I., Etzioni, O., Khot, T., Sabharwal, A., Schoenick, C., and Tafjord, O. (2018) · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. (2018) · 2018
Earlier work this paper cites.
Kudo, T. and Richardson, J. (2018) · 2018
Earlier work this paper cites.
Adafactor: Adaptive learning rates with sublinear memory cost
Shazeer, N. and Stern, M. (2018) · 2018
Earlier work this paper cites.
The brwac corpus: A new open resource for brazilian portuguese
Wagner Filho, J. A., Wilkens, R., Idiart, M., and Villavicencio, A. (2018) · 2018
Earlier work this paper cites.
Codecarbon: Track emissions from compute and recommend ways to reduce their impact on the environment
CodeCarbon (2019) · 2019
Earlier work this paper cites.
Estimation of energy consumption in machine learning
García-Martín, E., Rodrigues, C. F., Riley, G., and Grahn, H. (2019) · 2019
Earlier work this paper cites.
Cosmos QA: Machine reading comprehension with contextual commonsense reasoning
Huang, L., Le Bras, R., Bhagavatula, C., and Choi, Y. (2019) · 2019
Earlier work this paper cites.
Tokenizers: Fast state-of-the-art tokenizers optimized for research and production
HuggingFace (2019) · 2019
Earlier work this paper cites.
Asynchronous pipelines for processing huge corpora on medium to low resource infrastructures
Ortiz Su’arez, P. J., Sagot, B., and Romary, L. (2019) · 2019
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., et al. (2019) · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al. (2019) · 2019
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J. (2019) · 2019
Earlier work this paper cites.
Root mean square layer normalization
Zhang, B. and Sennrich, R. (2019) · 2019
Earlier work this paper cites.
Unsupervised cross-lingual representation learning at scale
Conneau, A., Khandelwal, K., Goyal, N., Chaudhary, V., Wenzek, G., Guzmán, F., Grave, E., Ott, M., Zettlemoyer, L., and Stoyanov, V. (2020) · 2020
Earlier work this paper cites.
The pile: An 800gb dataset of diverse text for language modeling
Gao, L., Biderman, S., Black, S., Golding, L., Hoppe, T., Foster, C., Phang, J., He, H., Thite, A., Nabeshima, N., et al. (2020) · 2020
Earlier work this paper cites.
Gportuguese-2 (portuguese gpt-2 small): a language model for portuguese text generation (and more nlp tasks…)
Guillou, P. (2020) · 2020
Earlier work this paper cites.
CamemBERT: a tasty French language model
Martin, L., Muller, B., Ortiz Suárez, P. J., Dupont, Y., Romary, L., de la Clergerie, É., Seddah, D., and Sagot, B. (2020) · 2020
Earlier work this paper cites.
A monolingual approach to contextualized word embeddings for mid-resource languages
Ortiz Su’arez, P. J., Romary, L., and Sagot, B. (2020) · 2020
Earlier work this paper cites.
Zero: Memory optimizations toward training trillion parameter models
Rajbhandari, S., Rasley, J., Ruwase, O., and He, Y. (2020) · 2020
Cited alongside, same era.
The assin 2 shared task: a quick overview
Real, L., Fonseca, E., and Oliveira, H. G. (2020) · 2020
Cited alongside, same era.
Bertimbau: pretrained bert models for brazilian portuguese
Souza, F., Nogueira, R., and Lotufo, R. (2020) · 2020
Cited alongside, same era.
CCNet: Extracting high quality monolingual datasets from web crawl data
Wenzek, G., Lachaux, M.-A., Conneau, A., Chaudhary, V., Guzmán, F., Joulin, A., and Grave, E. (2020) · 2020
Cited alongside, same era.
Transformers: State-of-the-art natural language processing
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., Davison, J., Shleifer, S., von Platen, P., Ma, C., Jernite, Y., Plu, J., Xu, C., Scao, T. L., Gugger, S., Drame, M., Lhoest, Q., and Rush, A. M. (2020) · 2020
Cited alongside, same era.
Redpajama: an open dataset for training large language models
Computer, T. (2023) · 2023
Later among the works it cites.
Free dolly: Introducing the world’s first truly open instruction-tuned llm
Conover, M., Hayes, M., Mathur, A., Xie, J., Wan, J., Shah, S., Ghodsi, A., Wendell, P., Zaharia, M., and Xin, R. (2023) · 2023
Later among the works it cites.
Efficient and effective text encoding for chinese llama and alpaca
Cui, Y., Yang, Z., and Yao, X. (2023) · 2023
Later among the works it cites.
Flashattention-2: Faster attention with better parallelism and work partitioning
Dao, T. (2023) · 2023
Later among the works it cites.
Cerebras-gpt: Open compute-optimal language models trained on the cerebras wafer-scale cluster
Dey, N., Gosal, G., Khachane, H., Marshall, W., Pathria, R., Tom, M., Hestness, J., et al. (2023) · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Antoun, W., Baly, F., and Hajj, H. (2021) · 2021
Cited alongside, same era.
A general language assistant as a laboratory for alignment
Askell, A., Bai, Y., Chen, A., Drain, D., Ganguli, D., Henighan, T., Jones, A., Joseph, N., Mann, B., DasSarma, N., et al. (2021) · 2021
Cited alongside, same era.
Compute and energy consumption trends in deep learning inference
Desislavov, R., Martínez-Plumed, F., and Hernández-Orallo, J. (2021) · 2021
Cited alongside, same era.
A framework for few-shot language model evaluation
Gao, L., Tow, J., Biderman, S., Black, S., DiPofi, A., Foster, C., Golding, L., Hsu, J., McDonell, K., Muennighoff, N., et al. (2021) · 2021
Cited alongside, same era.
Maria: Spanish language models
Gutiérrez-Fandiño, A., Armengol-Estapé, J., Pàmies, M., Llop-Palao, J., Silveira-Ocampo, J., Carrino, C. P., Gonzalez-Agirre, A., Armentano-Oller, C., Rodriguez-Penagos, C., and Villegas, M. (2021) · 2021
Cited alongside, same era.
Deduplicating training data makes language models better
Lee, K., Ippolito, D., Nystrom, A., Zhang, C., Eck, D., Callison-Burch, C., and Carlini, N. (2021) · 2021
Cited alongside, same era.
Datasets: A community library for natural language processing
Lhoest, Q., Villanova del Moral, A., Jernite, Y., Thakur, A., von Platen, P., Patil, S., Chaumond, J., Drame, M., Plu, J., Tunstall, L., Davison, J., Šaško, M., Chhablani, G., Malik, B., Brandeis, S., Le Scao, T., Sanh, V., Xu, C., Patry, N., McMillan-Major, A., Schmid, P., Gugger, S., Delangue, C., Matussière, T., Debut, L., Bekman, S., Cistac, P., Goehringer, T., Mustar, V., Lagunas, F., Rush, A., and Wolf, T. (2021) · 2021
Cited alongside, same era.
Later among the works it cites.
Enhancing chat language models by scaling high-quality instructional conversations
Ding, N., Chen, Y., Xu, B., Qin, Y., Zheng, Z., Hu, S., Liu, Z., Sun, M., and Zhou, B. (2023) · 2023
Later among the works it cites.
Challenging ai for sustainability: what ought it mean?
Falk, S. and van Wynsberghe, A. (2023) · 2023
Later among the works it cites.
Koala: A dialogue model for academic research
Geng, X., Gudibande, A., Liu, H., Wallace, E., Abbeel, P., Levine, S., and Song, D. (2023) · 2023
Later among the works it cites.
Openllama: An open reproduction of llama
Geng, X. and Liu, H. (2023) · 2023
Later among the works it cites.
Mamba: Linear-time sequence modeling with selective state spaces
Gu, A. and Dao, T. (2023) · 2023
Later among the works it cites.
Gunasekar, S., Zhang, Y., Aneja, J., Mendes, C. C. T., Del Giorno, A., Gopi, S., Javaheripi, M., Kauffmann, P., de Rosa, G., Saarikivi, O., et al. (2023) · 2023
Later among the works it cites.
Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., Casas, D. d. l., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., et al. (2023) · 2023
Later among the works it cites.
A technical report for polyglot-ko: Open-source large-scale korean language models
Ko, H., Yang, K., Ryu, M., Choi, T., Yang, S., Park, S., et al. (2023) · 2023
Later among the works it cites.
Openassistant conversations–democratizing large language model alignment
Köpf, A., Kilcher, Y., von Rütte, D., Anagnostidis, S., Tam, Z.-R., Stevens, K., Barhoum, A., Duc, N. M., Stanley, O., Nagyfi, R., et al. (2023) · 2023
Later among the works it cites.
Open multilingual llm evaluation leaderboard
Lai, V., Ngo, N. T., Veyseh, A. P. B., Dernoncourt, F., and Nguyen, T. H. (2023) · 2023
Later among the works it cites.
adaptmllm: Fine-tuning multilingual language models on low-resource languages with integrated llm playgrounds
Lankford, S., Afli, H., and Way, A. (2023) · 2023
Later among the works it cites.
Cabrita: closing the gap for foreign languages
Larcher, C., Piau, M., Finardi, P., Gengo, P., Esposito, P., and Caridá, V. (2023) · 2023
Later among the works it cites.
Bactrian-x : A multilingual replicable instruction-following model with low-rank adaptation
Li, H., Koto, F., Wu, M., Aji, A. F., and Baldwin, T. (2023) · 2023
Later among the works it cites.
Awq: Activation-aware weight quantization for llm compression and acceleration
Lin, J., Tang, J., Tang, H., Yang, S., Dang, X., and Han, S. (2023) · 2023
Later among the works it cites.
Yayi 2: Multilingual open-source large language models
Luo, Y., Kong, Q., Xu, N., Cao, J., Hao, B., Qu, B., Chen, B., Zhu, C., Zhao, C., Zhang, D., et al. (2023) · 2023
Later among the works it cites.
Scaling data-constrained language models
Muennighoff, N., Rush, A. M., Barak, B., Scao, T. L., Piktus, A., Tazi, N., Pyysalo, S., Wolf, T., and Raffel, C. (2023) · 2023
Later among the works it cites.
Rwkv: Reinventing rnns for the transformer era
Peng, B., Alcaide, E., Anthony, Q., Albalak, A., Arcadinho, S., Cao, H., Cheng, X., Chung, M., Grella, M., GV, K. K., et al. (2023) · 2023
Later among the works it cites.
Sabi \ \backslash ’a: Portuguese large language models
Pires, R., Abonizio, H., Rogério, T., and Nogueira, R. (2023) · 2023
Later among the works it cites.
Advancing neural encoding of portuguese with transformer albertina pt
Rodrigues, J., Gomes, L., Silva, J., Branco, A., Santos, R., Cardoso, H. L., and Osório, T. (2023) · 2023
Later among the works it cites.
Faquad-nli: a benchmark for textual entailment
Rodrigues, R. C. (2023) · 2023
Later among the works it cites.
Code llama: Open foundation models for code
Roziere, B., Gehring, J., Gloeckle, F., Sootla, S., Gat, I., Tan, X. E., Adi, Y., Liu, J., Remez, T., Rapin, J., et al. (2023) · 2023
Later among the works it cites.
Mixture-of-experts meets instruction tuning: A winning combination for large language models
Shen, S., Hou, L., Zhou, Y., Du, N., Longpre, S., Wei, J., Chung, H. W., Zoph, B., Fedus, W., Chen, X., et al. (2023) · 2023
Later among the works it cites.
Alpaca: A strong, replicable instruction-following model
Taori, R., Gulrajani, I., Zhang, T., Dubois, Y., Li, X., Guestrin, C., Liang, P., and Hashimoto, T. B. (2023) · 2023
Later among the works it cites.
To repeat or not to repeat: Insights from scaling llm under token-crisis
Xue, F., Fu, Y., Zhou, W., Zheng, Z., and You, Y. (2023) · 2023
Later among the works it cites.
Deepseek llm: Scaling open-source language models with longtermism
Bi, X., Chen, D., Chen, G., Chen, S., Dai, D., Deng, C., Ding, H., Dong, K., Du, Q., Fu, Z., et al. (2024) · 2024
Closest in time.
Instruct-ptbr-enus-11m
Carlo Moro (2024) · 2024
Closest in time.
Introducing bode: A fine-tuned large language model for portuguese prompt-based task
Garcia, G. L., Paiola, P. H., Morelli, L. H., Candido, G., Júnior, A. C., Jodas, D. S., Afonso, L. C. S., Guilherme, I. R., Penteado, B. E., and Papa, J. P. (2024) · 2024
Closest in time.
Jiang, A. Q., Sablayrolles, A., Roux, A., Mensch, A., Savary, B., Bamford, C., Chaplot, D. S., Casas, D. d. l., Hanna, E. B., Bressand, F., et al. (2024) · 2024
Closest in time.
Nous-hermes-2-yi-34b
NousResearch (2024) · 2024
Closest in time.
Gpt4all-j prompt generations pt
Pablo Filetti Moreira (2024) · 2024
Closest in time.
Stable lm 2 1.6b
Team, S. A. L. (2024) · 2024
Closest in time.
Wikimedia Downloads
Wikimedia Foundation (2024) · 2024
Closest in time.
Tinyllama: An open-source small language model
Zhang, P., Zeng, G., Wang, T., and Lu, W. (2024) · 2024
Closest in time.
Llama beyond english: An empirical study on language capability transfer
Zhao, J., Zhang, Z., Zhang, Q., Gui, T., and Huang, X. (2024) · 2024
Closest in time.