Fetching the paper…
Reading the bibliography…
Despite advancements in Natural Language Processing (NLP) and the growing availability of pretrained models, the English language remains the primary focus of model development.
“SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems”, 2020
Alex Wang et al · 1905
Earlier work this paper cites.
“Unsupervised Cross-lingual Representation Learning at Scale”, 2020
Alexis Conneau et al · 1911
Earlier work this paper cites.
“The Parsing System Palavras: Automatic Grammatical Analysis of Portuguese in a Constraint Grammar Famework”
Eckhard Bick · 2000
Earlier work this paper cites.
“Scaling Laws for Neural Language Models”, 2020
Jared Kaplan et al · 2001
Earlier work this paper cites.
Junjie Hu et al · 2003
Earlier work this paper cites.
“Linguateca: um centro de recursos distribuído para o processamento computacional da língua portuguesa”
Diana Santos et al · 2004
Earlier work this paper cites.
“Efficient corpus development for lexicography: Building the New Corpus for Ireland”
Adam Kilgarriff, Michael Rundell and Elaine Dhonnchadha · 2006
Earlier work this paper cites.
“A WaCky Introduction”, 2008, pp. 9–40
Silvia Bernardini, Marco Baroni and Stefan Evert · 2008
Earlier work this paper cites.
“PTT5: Pretraining and validating the T5 model on Brazilian Portuguese data”, 2020
Diedre Carmo et al · 2008
Earlier work this paper cites.
“mT5: A massively multilingual pre-trained text-to-text transformer”, 2021
Linting Xue et al · 2010
Earlier work this paper cites.
“Teaching Machines to Read and Comprehend”, 2015
Karl Hermann et al · 2015
Earlier work this paper cites.
“SQuAD: 100,000+ Questions for Machine Comprehension of Text”, 2016
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev and Percy Liang · 2016
Earlier work this paper cites.
“MS MARCO: A Human Generated MAchine Reading COmprehension Dataset”, 2018
Payal Bajaj et al · 2018
Earlier work this paper cites.
“Building a Sentiment Corpus of Tweets in Brazilian Portuguese”
Henrico Brum and Mariaças Volpe · 2018
Earlier work this paper cites.
Taku Kudo and John Richardson · 2018
Earlier work this paper cites.
“Adafactor: Adaptive Learning Rates with Sublinear Memory Cost”, 2018
Noam Shazeer and Mitchell Stern · 2018
Earlier work this paper cites.
“The brwac corpus: A new open resource for brazilian portuguese”
Jorge Wagner, Rodrigo Wilkens, Marco Idiart and Aline Villavicencio · 2018
Earlier work this paper cites.
“BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding”, 2019
Jacob Devlin, Ming-Wei Chang, Kenton Lee and Kristina Toutanova · 2019
Earlier work this paper cites.
“Exploring the limits of transfer learning with a unified text-to-text transformer”
Colin Raffel et al · 2019
Cited alongside, same era.
“ftfy” Version 5.5, Zenodo, 2019
Robyn Speer · 2019
Cited alongside, same era.
“GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding”, 2019
Alex Wang et al · 2019
Cited alongside, same era.
“spaCy: Industrial-strength Natural Language Processing in Python”, 2020
Matthew Honnibal, Ines Montani, Sofie Van and Adriane Boyd · 2020
Cited alongside, same era.
“Document Ranking with a Pretrained Sequence-to-Sequence Model”
Rodrigo Nogueira, Zhiying Jiang, Ronak Pradeep and Jimmy Lin · 2020
Cited alongside, same era.
“Scaling Language Models: Methods, Analysis & Insights from Training Gopher”, 2022
Jack. Rae et al · 2022
Later among the works it cites.
“Scaling Up Models and Data with t5x
Adam Roberts et al · 2022
Later among the works it cites.
“IT5: Large-scale Text-to-text Pretraining for Italian Language Understanding and Generation”, 2022
Gabriele Sarti and Malvina Nissim · 2022
Later among the works it cites.
Israel Campiotti et al · 2023
Later among the works it cites.
“BERTabaporu: Assessing a Genre-Specific Language Model for Portuguese NLP”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
In Proceedings of the ASSIN 2 Shared Task: Evaluating Semantic Textual Similarity and Textual Entailment in Portuguese , CEUR Workshop Proceedings 2583, 2020
“Proceedings of the ASSIN 2 Shared Task: Evaluating Semantic Textual Similarity and Textual Entailment in Portuguese, Extended Semantic Web Conference” · 2020
Cited alongside, same era.
“BERTimbau: Pretrained BERT Models for Brazilian Portuguese”
Fábio Souza, Rodrigo Nogueira and Roberto Lotufo · 2020
Cited alongside, same era.
“Leveraging emoji to improve sentiment classification of tweets”
Tiago de Barros, Helio Pedrini and Zanoni Dias · 2021
Cited alongside, same era.
“BERTaú: Itaú BERT for digital customer service”, 2021
Paulo Finardi et al · 2021
Cited alongside, same era.
“Pyserini: A Python Toolkit for Reproducible Information Retrieval Research with Sparse and Dense Representations”
Jimmy Lin et al · 2021
Cited alongside, same era.
“A cost-benefit analysis of cross-lingual transfer methods”, 2021
Guilherme Rosa et al · 2021
Cited alongside, same era.
“mMARCO: A Multilingual Version of the MS MARCO Passage Ranking Dataset”, 2022
Luiz Bonifacio et al · 2022
Cited alongside, same era.
Pablo Costa et al · 2023
Later among the works it cites.
“Deep Learning Brasil at ABSAPT 2022: Portuguese Transformer Ensemble Approaches”, 2023
Juliana Gomes et al · 2023
Later among the works it cites.
“Cabrita: closing the gap for foreign languages”, 2023
Celio Larcher et al · 2023
Later among the works it cites.
“Sub-language Sentiment Analysis in WhatsApp Domain with Deep Learning Approaches”
Leonardo de Morais et al · 2023
Later among the works it cites.
“Sabiá: Portuguese Large Language Models”
Ramon Pires, Hugo Abonizio, Thales Almeida and Rodrigo Nogueira · 2023
Later among the works it cites.
“Advancing Neural Encoding of Portuguese with Transformer Albertina PT-*”, 2023
João Rodrigues et al · 2023
Later among the works it cites.
“Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models”, 2023
Aarohi Srivastava et al · 2023
Later among the works it cites.
“Quati: A Brazilian Portuguese Information Retrieval Dataset from Native Speakers”, 2024
Mirelle Bueno et al · 2024
Closest in time.
“Efficient and Effective Text Encoding for Chinese LLaMA and Alpaca”, 2024
Yiming Cui, Ziqing Yang and Xin Yao · 2024
Closest in time.
“Language models scale reliably with over-training and on downstream tasks”, 2024
Samir Gadre et al · 2024
Closest in time.
“Introducing Bode: A Fine-Tuned Large Language Model for Portuguese Prompt-Based Task”, 2024
Gabriel Garcia et al · 2024
Closest in time.
“GlórIA – A Generative and Open Large Language Model for Portuguese”, 2024
Ricardo Lopes, João Magalhães and David Semedo · 2024
Closest in time.
“Advancing Generative AI for Portuguese with Open Decoder Gervásio PT*”, 2024
Rodrigo Santos et al · 2024
Closest in time.
“Multilingual E5 Text Embeddings: A Technical Report”, 2024
Liang Wang et al · 2024
Closest in time.