Fetching the paper…
Reading the bibliography…
This paper provides a detailed discussion of the multilingual tokenizer used for GPT-SW3.
“Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism”, 2020
Mohammad Shoeybi et al · 1909
Earlier work this paper cites.
“SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing”
Taku Kudo and John Richardson · 2012
Earlier work this paper cites.
“Neural Machine Translation of Rare Words with Subword Units”, 2016
Rico Sennrich, Barry Haddow and Alexandra Birch · 2016
Earlier work this paper cites.
Taku Kudo · 2018
Cited alongside, same era.
“Exploring BERT’s Vocabulary”, 2019
Judit Acs · 2019
Cited alongside, same era.
“How Good is Your Tokenizer? On the Monolingual Performance of Multilingual Language Models”
Phillip Rust et al · 2021
Cited alongside, same era.
“GPT-SW3: An Autoregressive Language Model for the Nordic Languages” in preparation
Ariel Ekgren et al
Cited in the paper.
“GPT-NeoX-20B: An Open-Source Autoregressive Language Model”, 2022
Sid Black et al · 2022
Later among the works it cites.
“The Nordic Pile: A 1.2TB Nordic Dataset for Language Modeling”, 2023
Joey Öhman et al · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…