Fetching the paper…
Reading the bibliography…
Standard pretrained language models operate on sequences of subword tokens without direct access to the characters that compose each token's string representation.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Binary codes capable of correcting deletions, insertions, and reversals
Vladimir I Levenshtein et al. 1966 · 1966
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 1966
Earlier work this paper cites.
NLTK: The natural language toolkit
Edward Loper and Steven Bird. 2002 · 2002
Earlier work this paper cites.
GloVe: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher Manning. 2014 · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba. 2015 · 2015
Cited alongside, same era.
chrF: character n-gram F-score for automatic MT evaluation
Maja Popović. 2015 · 2015
Cited alongside, same era.
Farasa: A new fast and accurate Arabic word segmenter
Kareem Darwish and Hamdy Mubarak. 2016 · 2016
Cited alongside, same era.
Deep contextualized word representations
Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018 · 2018
Cited alongside, same era.
fairseq: A fast, extensible toolkit for sequence modeling
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli. 2019 · 2019
Cited alongside, same era.
Efficient estimation of word representations in vector space
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019 · 2019
Later among the works it cites.
AraBERT: Transformer-based model for Arabic language understanding
Wissam Antoun, Fady Baly, and Hazem Hajj. 2020 · 2020
Later among the works it cites.
CharacterBERT: Reconciling ELMo and BERT for word-level open-vocabulary representations from characters
Hicham El Boukkouri, Olivier Ferret, Thomas Lavergne, Hiroshi Noji, Pierre Zweigenbaum, and Jun’ichi Tsujii. 2020 · 2020
Later among the works it cites.
How to train bert with an academic budget
Peter Izsak, Moshe Berchansky, and Omer Levy. 2021 · 2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013a
Cited in the paper.
Distributed representations of words and phrases and their compositionality
Tomas Mikolov, Ilya Sutskever, Kai Chen, Gregory S. Corrado, and Jeffrey Dean. 2013b
Cited in the paper.