Fetching the paper…
Reading the bibliography…
Pipelined NLP systems have largely been superseded by end-to-end neural modeling, yet nearly all commonly-used models still require an explicit tokenization step.
Bridging the Gap for Tokenizer-Free Language Models
Dokook Choe, Rami Al-Rfou, Mandy Guo, Heeyoung Lee, and Noah Constant. 2019 · 1908
Earlier work this paper cites.
Introduction to the CoNLL-2002 shared task: Language-independent named entity recognition
Erik F. Tjong Kim Sang. 2002 · 2002
Earlier work this paper cites.
Adv-BERT: BERT is not robust on misspellings! Generating nature adversarial samples on BERT
Lichao Sun, Kazuma Hashimoto, Wenpeng Yin, Akari Asai, Jia Li, Philip Yu, and Caiming Xiong. 2020 · 2003
Earlier work this paper cites.
Introduction to the CoNLL-2003 shared task: Language-independent named entity recognition
Erik F. Tjong Kim Sang and Fien De Meulder. 2003 · 2003
Earlier work this paper cites.
Bloom maps
David Talbot and John Talbot. 2008 · 2008
Earlier work this paper cites.
TweetMotif: Exploratory Search and Topic Summarization for Twitter Introduction and Description
Brendan O’Connor, Michel Krieger, and David Ahn. 2010 · 2010
Earlier work this paper cites.
Generating text with recurrent neural networks
Ilya Sutskever, James Martens, and Geoffrey E Hinton. 2011 · 2011
Earlier work this paper cites.
Adaptor Grammars for learning non-concatenative morphology
Jan A. Botha and Phil Blunsom. 2013 · 2013
Earlier work this paper cites.
Learning a part-of-speech tagger from two hours of annotation
Dan Garrette and Jason Baldridge. 2013 · 2013
Earlier work this paper cites.
Generating sequences with recurrent neural networks
Alex Graves. 2013 · 2013
Earlier work this paper cites.
Learning deep structured semantic models for web search using clickthrough data
Po-Sen Huang, Xiaodong He, Jianfeng Gao, Li Deng, Alex Acero, and Larry Heck. 2013 · 2013
Earlier work this paper cites.
Character-level convolutional networks for text classification
Xiang Zhang, Junbo Zhao, and Yann LeCun. 2015 · 2015
Earlier work this paper cites.
Multilingual Language Processing From Bytes
Dan Gillick, Cliff Brunk, Oriol Vinyals, and Amarnag Subramanya. 2016 · 2016
Earlier work this paper cites.
Character-Aware Neural Language Models
Yoon Kim, Yacine Jernite, David Sontag, and Alexander M Rush. 2016 · 2016
Earlier work this paper cites.
Achieving open vocabulary neural machine translation with hybrid word-character models
Minh-Thang Luong and Christopher D. Manning. 2016 · 2016
Earlier work this paper cites.
Neural Machine Translation of Rare Words with Subword Units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Earlier work this paper cites.
Charagram: Embedding words and sentences via character n-grams
John Wieting, Mohit Bansal, Kevin Gimpel, and Karen Livescu. 2016 · 2016
Earlier work this paper cites.
Google’s Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, Jeff Klingner, Apurva Shah, Melvin Johnson, Xiaobing Liu, ukasz Kaiser, Stephan Gouws, Yoshikiyo Kato, Taku Kudo, Hideto Kazawa, Keith Stevens, George Kurian, Nishant Patil, Wei Wang, Cliff Young, Jason Smith, Jason Riesa, Alex Rudnick, Oriol Vinyals, Greg Corrado, Macduff Hughes, and Jeffrey Dean. 2016 · 2016
Earlier work this paper cites.
Enriching word vectors with subword information
Piotr Bojanowski, Edouard Grave, Armand Joulin, and Tomas Mikolov. 2017 · 2017
Earlier work this paper cites.
Hierarchical Multiscale Recurrent Neural Networks
Junyoung Chung, Sungjin Ahn, and Yoshua Bengio. 2017 · 2017
Earlier work this paper cites.
Character-Level Language Modeling with Hierarchical Recurrent Neural Networks
Kyuyeon Hwang and Wonyong Sung. 2017 · 2017
Earlier work this paper cites.
Learning to create and reuse words in open-vocabulary neural language modeling
Kazuya Kawakami, Chris Dyer, and Phil Blunsom. 2017 · 2017
Earlier work this paper cites.
Fully Character-Level Neural Machine Translation without Explicit Segmentation
Jason Lee, Eth Zürich, Kyunghyun Cho, and Thomas Hofmann. 2017 · 2017
Earlier work this paper cites.
Mimicking word embeddings using subword RNNs
Yuval Pinter, Robert Guthrie, and Jacob Eisenstein. 2017 · 2017
Earlier work this paper cites.
Hash Embeddings for Efficient Word Representations
Dan Svenstrup, Jonas Meinertz Hansen, and Ole Winther. 2017 · 2017
Earlier work this paper cites.
Multiscale sequence modeling with a learned dictionary
Bart van Merriënboer, Amartya Sanyal, H. Larochelle, and Yoshua Bengio. 2017 · 2017
Cited alongside, same era.
From Characters to Words to in Between: Do We Capture Morphology?
Clara Vania and Adam Lopez. 2017 · 2017
Cited alongside, same era.
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Which Encoding is the Best for Text Classification in Chinese, English, Japanese and Korean?
Xiang Zhang and Yann LeCun. 2017 · 2017
Cited alongside, same era.
Contextual String Embeddings for Sequence Labeling
Alan Akbik, Duncan Blythe, and Roland Vollgraf. 2018 · 2018
Cited alongside, same era.
How multilingual is Multilingual BERT?
Telmo Pires, Eva Schlinger, and Dan Garrette. 2019 · 2019
Later among the works it cites.
Combating adversarial misspellings with robust word recognition
Danish Pruthi, Bhuwan Dhingra, and Zachary C. Lipton. 2019 · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Later among the works it cites.
BERT rediscovers the classical NLP pipeline
Ian Tenney, Dipanjan Das, and Ellie Pavlick. 2019 · 2019
Later among the works it cites.
ETC: Encoding Long and Structured Inputs in Transformers
Joshua Ainslie, Santiago Ontanon, Chris Alberti, Vaclav Cvicek, Zachary Fisher, Philip Pham, Anirudh Ravula, Sumit Sanghai, Qifan Wang, and Li Yang. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Part-of-Speech Tagging for Code-Switched, Transliterated Texts without Explicit Language Identification
Kelsey Ball and Dan Garrette. 2018 · 2018
Cited alongside, same era.
Neural lattice language models
Jacob Buckman and Graham Neubig. 2018 · 2018
Cited alongside, same era.
Language Modeling for Morphologically Rich Languages: Character-Aware Modeling for Word-Level Prediction
Daniela Gerz, Ivan Vulić, Edoardo Ponti, Jason Naradowsky, Roi Reichart, and Anna Korhonen. 2018 · 2018
Cited alongside, same era.
Byte-Level Machine Reading across Morphologically Varied Languages
Daniel Hewlett, Alexandre Lacoste, Llion Jones, Illia Polosukhin, Andrew Fandrianto, Jay Han, Matthew Kelcey, and David Berthelot. 2018 · 2018
Cited alongside, same era.
Subword Regularization: Improving Neural Network Translation Models with Multiple Subword Candidates
Taku Kudo. 2018 · 2018
Cited alongside, same era.
Taku Kudo and John Richardson. 2018 · 2018
Cited alongside, same era.
Deep contextualized word representations
Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018 · 2018
Cited alongside, same era.
Byte pair encoding is suboptimal for language model pretraining
Kaj Bostrom and Greg Durrett. 2020 · 2020
Later among the works it cites.
Hicham El Boukkouri, Olivier Ferret, Thomas Lavergne, Hiroshi Noji, Pierre Zweigenbaum, and Junichi Tsujii. 2020 · 2020
Later among the works it cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 2020
Later among the works it cites.
Improving Multilingual Models with Language-Clustered Vocabularies
Hyung Won Chung, Dan Garrette, Kiat Chuan Tan, and Jason Riesa. 2020 · 2020
Later among the works it cites.
TyDi QA: A benchmark for Information-Seeking Question Answering in Typologically Diverse Languages
Jonathan H. Clark, Eunsol Choi, Michael Collins, Dan Garrette, Tom Kwiatkowski, Vitaly Nikolaev, and Jennimaria Palomaki. 2020 · 2020
Later among the works it cites.
Unsupervised Cross-lingual Representation Learning at Scale
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020 · 2020
Later among the works it cites.
Funnel-Transformer: Filtering out Sequential Redundancy for Efficient Language Processing
Zihang Dai, Guokun Lai, Yiming Yang, and Quoc V. Le. 2020 · 2020
Later among the works it cites.
Character-level Representations Improve DRS-based Semantic Parsing Even in the Age of BERT
Rik van Noord, Antonio Toral, and Johan Bos. 2020 · 2020
Later among the works it cites.
English Intermediate-Task Training Improves Zero-Shot Cross-Lingual Transfer Too
Jason Phang, Iacer Calixto, Phu Mon Htut, Yada Pruksachatkun, Haokun Liu, Clara Vania, Katharina Kann, and Samuel R Bowman. 2020 · 2020
Later among the works it cites.
BPE-Dropout: Simple and Effective Subword Regularization
Ivan Provilkov, Dmitrii Emelianenko, and Elena Voita. 2020 · 2020
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020 · 2020
Later among the works it cites.
Extending multilingual BERT to low-resource languages
Zihan Wang, Karthikeyan K, Stephen Mayhew, and Dan Roth. 2020 · 2020
Later among the works it cites.
Large Batch Optimization for Deep Learning: Training BERT in 76 minutes
Yang You, Jing Li, Sashank Reddi, Jonathan Hseu, Sanjiv Kumar, Srinadh Bhojanapalli, Xiaodan Song, James Demmel, Kurt Keutzer, and Cho-Jui Hsieh. 2020 · 2020
Later among the works it cites.
Big Bird: Transformers for Longer Sequences
Manzil Zaheer, Guru Guruganesh, Avinava Dubey, Joshua Ainslie, Chris Alberti, Santiago Ontanon, Philip Pham, Anirudh Ravula, Qifan Wang, Li Yang, and Amr Ahmed. 2020 · 2020
Later among the works it cites.
MasakhaNER: Named Entity Recognition for African Languages
David Ifeoluwa Adelani, Jade Abbott, Graham Neubig, Daniel D’souza, Julia Kreutzer, Constantine Lignos, Chester Palen-Michel, Happy Buzaaba, Shruti Rijhwani, Sebastian Ruder, Stephen Mayhew, Israel Abebe Azime, Shamsuddeen H. Muhammad, Chris Chinenye Emezue, Joyce Nakatumba-Nabende, Perez Ogayo, Aremu Anuoluwapo, Catherine Gitau, Derguene Mbaye, Jesujoba Alabi, Seid Muhie Yimam, Tajuddeen Rabiu Gwadabe, Ignatius Ezeani, Rubungo Andre Niyongabo, Jonathan Mukiibi, Verrah Otiende, Iroro Orife, Davis David, Samba Ngom, Tosin Adewumi, Paul Rayson, Mofetoluwa Adeyemi, Gerald Muriuki, Emmanuel Anebi, Chiamaka Chukwuneke, Nkiruka Odu, Eric Peter Wairagala, Samuel Oyerinde, Clemencia Siro, Tobius Saul Bateesa, Temilola Oloyede, Yvonne Wambui, Victor Akinode, Deborah Nabagereka, Maurice Katusiime, Ayodele Awokoya, Mouhamadane MBOUP, Dibora Gebreyohannes, Henok Tilaye, Kelechi Nwaike, Degaga Wolde, Abdoulaye Faye, Blessing Sibanda, Orevaoghene Ahia, Bonaventure F. P. Dossou, Kelechi Ogueji, Thierno Ibrahima DIOP, Abdoulaye Diallo, Adewale Akinfaderin, Tendai Marengereke, and Salomey Osei. 2021 · 2021
Closest in time.
Rethinking Attention with Performers
Krzysztof Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamas Sarlos, Peter Hawkins, Jared Davis, Afroz Mohiuddin, Lukasz Kaiser, David Belanger, Lucy Colwell, and Adrian Weller. 2021 · 2021
Closest in time.
Random feature attention
Hao Peng, Nikolaos Pappas, Dani Yogatama, Roy Schwartz, Noah Smith, and Lingpeng Kong. 2021 · 2021
Closest in time.
Multi-view subword regularization
Xinyi Wang, Sebastian Ruder, and Graham Neubig. 2021 · 2021
Closest in time.
mT5: A massively multilingual pre-trained text-to-text transformer
Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. 2021 · 2021
Closest in time.