Fetching the paper…
Reading the bibliography…
Byte-pair encoding (BPE) is a ubiquitous algorithm in the subword tokenization process of language models as it provides multiple benefits.
Roberta: A robustly optimized BERT pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
CTRL: A conditional transformer language model for controllable generation
Nitish Shirish Keskar, Bryan McCann, Lav R Varshney, Caiming Xiong, and Richard Socher. 2019 · 1909
Earlier work this paper cites.
What to do about bad language on the internet
Jacob Eisenstein. 2013b · 2013
Earlier work this paper cites.
Better word representations with recursive neural networks for morphology
Thang Luong, Richard Socher, and Christopher Manning. 2013 · 2013
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean. 2013 · 2013
Earlier work this paper cites.
Substring-based machine translation
Graham Neubig, Taro Watanabe, Shinsuke Mori, and Tatsuya Kawahara. 2013 · 2013
Earlier work this paper cites.
word2vec explained: deriving Mikolov et al.’s negative-sampling word-embedding method
Yoav Goldberg and Omer Levy. 2014 · 2014
Earlier work this paper cites.
GloVe: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher Manning. 2014 · 2014
Earlier work this paper cites.
Shared tasks of the 2015 workshop on noisy user-generated text: Twitter lexical normalization and named entity recognition
Timothy Baldwin, Marie Catherine de Marneffe, Bo Han, Young-Bum Kim, Alan Ritter, and Wei Xu. 2015 · 2015
Earlier work this paper cites.
Finding function in form: Compositional character models for open vocabulary word representation
Wang Ling, Chris Dyer, Alan W Black, Isabel Trancoso, Ramón Fermandez, Silvio Amir, Luís Marujo, and Tiago Luís. 2015a · 2015
Earlier work this paper cites.
Rupesh Kumar Srivastava, Klaus Greff, and Jürgen Schmidhuber. 2015 · 2015
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton. 2016 · 2016
Earlier work this paper cites.
Morphological priors for probabilistic neural word embeddings
Parminder Bhatia, Robert Guthrie, and Jacob Eisenstein. 2016 · 2016
Earlier work this paper cites.
A character-level decoder without explicit segmentation for neural machine translation
Junyoung Chung, Kyunghyun Cho, and Yoshua Bengio. 2016 · 2016
Earlier work this paper cites.
Character-based neural machine translation
Marta R. Costa-jussà and José A. R. Fonollosa. 2016 · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Earlier work this paper cites.
Bridging nonlinearities and stochastic regularizers with gaussian error linear units
Dan Hendrycks and Kevin Gimpel. 2016 · 2016
Earlier work this paper cites.
Character-aware neural language models
Yoon Kim, Yacine Jernite, D. Sontag, and Alexander M. Rush. 2016 · 2016
Cited alongside, same era.
Achieving open vocabulary neural machine translation with hybrid word-character models
Minh-Thang Luong and Christopher D. Manning. 2016 · 2016
Cited alongside, same era.
Overview for the second shared task on language identification in code-switched data
Giovanni Molina, Fahad AlGhamdi, Mahmoud Ghoneim, Abdelati Hawwari, Nicolas Rey-Villamizar, Mona Diab, and Thamar Solorio. 2016 · 2016
Cited alongside, same era.
Multilingual part-of-speech tagging with bidirectional long short-term memory models and auxiliary loss
Barbara Plank, Anders Søgaard, and Yoav Goldberg. 2016 · 2016
Cited alongside, same era.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Cited alongside, same era.
A Twitter corpus for Hindi-English code mixed POS tagging
Kushagra Singh, Indira Sen, and Ponnurangam Kumaraguru. 2018 · 2018
Later among the works it cites.
Pooled contextualized embeddings for named entity recognition
Alan Akbik, Tanja Bergmann, and Roland Vollgraf. 2019 · 2019
Later among the works it cites.
What does BERT look at? an analysis of BERT’s attention
Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D. Manning. 2019 · 2019
Later among the works it cites.
Tool contest on POS tagging for code-mixed Indian social media (Facebook, Twitter, and Whatsapp) text
Amitava Das. 2016 · 2019
Later among the works it cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V. Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, Jeff Klingner, Apurva Shah, Melvin Johnson, Xiaobing Liu, Lukasz Kaiser, Stephan Gouws, Yoshikiyo Kato, Taku Kudo, Hideto Kazawa, Keith Stevens, George Kurian, Nishant Patil, Wei Wang, Cliff Young, Jason Smith, Jason Riesa, Alex Rudnick, Oriol Vinyals, Greg Corrado, Macduff Hughes, and Jeffrey Dean. 2016 · 2016
Cited alongside, same era.
Very deep convolutional networks for text classification
Alexis Conneau, Holger Schwenk, Loïc Barrault, and Yann Lecun. 2017 · 2017
Cited alongside, same era.
Target-side word segmentation strategies for neural machine translation
Matthias Huck, Simon Riess, and Alexander Fraser. 2017 · 2017
Cited alongside, same era.
Mimicking word embeddings using subword RNNs
Yuval Pinter, Robert Guthrie, and Jacob Eisenstein. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Named entity recognition on code-switched data: Overview of the CALCS 2018 shared task
Gustavo Aguilar, Fahad AlGhamdi, Victor Soto, Mona Diab, Julia Hirschberg, and Thamar Solorio. 2018 · 2018
Cited alongside, same era.
Contextual string embeddings for sequence labeling
Alan Akbik, Duncan Blythe, and Roland Vollgraf. 2018 · 2018
Cited alongside, same era.
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Later among the works it cites.
Attentive mimicking: Better word embeddings by attending to informative contexts
Timo Schick and Hinrich Schütze. 2019 · 2019
Later among the works it cites.
Multilingual neural machine translation with soft decoupled encoding
Xinyi Wang, Hieu Pham, Philip Arthur, and Graham Neubig. 2019 · 2019
Later among the works it cites.
LinCE: A centralized benchmark for linguistic code-switching evaluation
Gustavo Aguilar, Sudipta Kar, and Thamar Solorio. 2020 · 2020
Closest in time.
From English to code-switching: Transfer learning with strong morphological clues
Gustavo Aguilar and Thamar Solorio. 2020 · 2020
Closest in time.
Byte pair encoding is suboptimal for language model pretraining
Kaj Bostrom and Greg Durrett. 2020 · 2020
Closest in time.
Unsupervised cross-lingual representation learning at scale
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020 · 2020
Closest in time.
Don’t stop pretraining: Adapt language models to domains and tasks
Suchin Gururangan, Ana Marasović, Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, and Noah A. Smith. 2020 · 2020
Closest in time.
Robust encodings: A framework for combating adversarial typos
Erik Jones, Robin Jia, Aditi Raghunathan, and Percy Liang. 2020 · 2020
Closest in time.
TextAttack: A framework for adversarial attacks, data augmentation, and adversarial training in NLP
John Morris, Eli Lifland, Jin Yong Yoo, Jake Grigsby, Di Jin, and Yanjun Qi. 2020 · 2020
Closest in time.
A primer in BERTology: What we know about how BERT works
Anna Rogers, Olga Kovaleva, and Anna Rumshisky. 2020 · 2020
Closest in time.