Fetching the paper…
Reading the bibliography…
All existing transformer-based approaches to NLP using subword tokenisation algorithms encode whitespace (word boundary information) through the use of special space symbols (such as \#\# or \_) forming part of tokens.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020 · 1901
Earlier work this paper cites.
Transformer-xl: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc V Le, and Ruslan Salakhutdinov. 2019 · 1901
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Limit-bert: Linguistic informed multi-task bert
Junru Zhou, Zhuosheng Zhang, Hai Zhao, and Shuailiang Zhang. 2019 · 1910
Earlier work this paper cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020 · 2001
Earlier work this paper cites.
The Reuters corpus volume 1 -from yesterday’s news to tomorrow’s language resources
Tony Rose, Mark Stevenson, and Miles Whitehead. 2002 · 2002
Earlier work this paper cites.
Introduction to the CoNLL-2003 shared task: Language-independent named entity recognition
Erik F. Tjong Kim Sang and Fien De Meulder. 2003 · 2003
Earlier work this paper cites.
Segatron: Segment-aware transformer for language modeling and understanding
He Bai, Peng Shi, Jimmy Lin, Yuqing Xie, Luchen Tan, Kun Xiong, Wen Gao, and Ming Li. 2020 · 2004
Earlier work this paper cites.
Morpheme segmentation gold standards for finnish and english
Mathias Johan Philip Creutz, Bo Krister Johan Linden, et al. 2004 · 2004
Earlier work this paper cites.
Automatically constructing a corpus of sentential paraphrases
William B Dolan and Chris Brockett. 2005 · 2005
Earlier work this paper cites.
Dagobert: Generating derivational morphology with a pretrained language model
Valentin Hofmann, Janet B Pierrehumbert, and Hinrich Schütze. 2020 · 2005
Earlier work this paper cites.
The fifth PASCAL recognizing textual entailment challenge
Luisa Bentivogli, Ido Dagan, Hoa Trang Dang, Danilo Giampiccolo, and Bernardo Magnini. 2009 · 2009
Earlier work this paper cites.
The winograd schema challenge
Hector Levesque, Ernest Davis, and Leora Morgenstern. 2012 · 2012
Earlier work this paper cites.
Japanese and korean voice search
Mike Schuster and Kaisuke Nakajima. 2012 · 2012
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D Manning, Andrew Ng, and Christopher Potts. 2013 · 2013
Earlier work this paper cites.
Ncbi disease corpus: a resource for disease name recognition and concept normalization
Rezarta Islamaj Doğan, Robert Leaman, and Zhiyong Lu. 2014 · 2014
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2015 · 2015
Earlier work this paper cites.
SQuAD: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Earlier work this paper cites.
SemEval-2017 task 1: Semantic textual similarity multilingual and crosslingual focused evaluation
Daniel Cer, Mona Diab, Eneko Agirre, Iñigo Lopez-Gazpio, and Lucia Specia. 2017 · 2017
Cited alongside, same era.
Results of the wnut2017 shared task on novel and emerging entity recognition
Leon Derczynski, Eric Nichols, Marieke van Erp, and Nut Limsopatham. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Subword regularization: Improving neural network translation models with multiple subword candidates
Taku Kudo. 2018 · 2018
Cited alongside, same era.
Morpholex: A derivational morphological database for 70,000 english words
Claudia H Sánchez-Gutiérrez, Hugo Mailhot, S Hélène Deacon, and Maximiliano A Wilson. 2018 · 2018
Cited alongside, same era.
Learning which features matter: RoBERTa acquires a preference for linguistic generalizations (eventually)
Alex Warstadt, Yian Zhang, Xiaocheng Li, Haokun Liu, and Samuel R. Bowman. 2020 · 2020
Later among the works it cites.
Morphynet: a large multilingual database of derivational and inflectional morphology
Khuyagbaatar Batsuren, Gábor Bella, and Fausto Giunchiglia. 2021 · 2021
Later among the works it cites.
Superbizarre is not superb: Derivational morphology improves BERT’s interpretation of complex words
Valentin Hofmann, Janet Pierrehumbert, and Hinrich Schütze. 2021 · 2021
Later among the works it cites.
Png bert: augmented bert on phonemes and graphemes for neural tts
Ye Jia, Heiga Zen, Jonathan Shen, Yu Zhang, and Yonghui Wu. 2021 · 2021
Later among the works it cites.
Morphology matters: a multilingual language modeling analysis
Hyunji Hayley Park, Katherine J Zhang, Coleman Haley, Kenneth Steimel, Han Liu, and Lane Schwartz. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. 2018 · 2018
Cited alongside, same era.
Neural network acceptability judgments
Alex Warstadt, Amanpreet Singh, and Samuel R. Bowman. 2018 · 2018
Cited alongside, same era.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel R. Bowman. 2018 · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Ladec: the large database of english compounds
Christina L Gagné, Thomas L Spalding, and Daniel Schmidtke. 2019 · 2019
Cited alongside, same era.
Superglue: A stickier benchmark for general-purpose language understanding systems
Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. 2019 · 2019
Cited alongside, same era.
Byte pair encoding is suboptimal for language model pretraining
Kaj Bostrom and Greg Durrett. 2020 · 2020
Cited alongside, same era.
Frustratingly simple pretraining alternatives to masked language modeling
Atsuki Yamaguchi, George Chrysostomou, Katerina Margatina, and Nikolaos Aletras. 2021 · 2021
Later among the works it cites.
Word order does matter and shuffled language models know it
Mostafa Abdou, Vinit Ravishankar, Artur Kulmizev, and Anders Søgaard. 2022 · 2022
Later among the works it cites.
Lert: A linguistically-motivated pre-trained language model
Yiming Cui, Wanxiang Che, Shijin Wang, and Ting Liu. 2022 · 2022
Later among the works it cites.
Improving tokenisation by alternative treatment of spaces
Edward Gow-Smith, Harish Tayyar Madabushi, Carolina Scarton, and Aline Villavicencio. 2022 · 2022
Later among the works it cites.
Training compute-optimal large language models
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al. 2022 · 2022
Later among the works it cites.
An embarrassingly simple method to mitigate undesirable properties of pretrained language model tokenizers
Valentin Hofmann, Hinrich Schuetze, and Janet Pierrehumbert. 2022 · 2022
Later among the works it cites.
Cassandra L Jacobs and Yuval Pinter. 2022 · 2022
Later among the works it cites.
Impact of morphological segmentation on pre-trained language models
Matheus Westhelle, Luciana Bencke, and Viviane P Moreira. 2022 · 2022
Later among the works it cites.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023 · 2023
Later among the works it cites.
McGill BabyLM shared task submission: The effects of data formatting and structural biases
Ziling Cheng, Rahul Aralikatte, Ian Porada, Cesare Spinoso-Di Piano, and Jackie CK Cheung. 2023 · 2023
Later among the works it cites.
Biomedical language models are robust to sub-optimal tokenization
Bernal Jimenez Gutierrez, Huan Sun, and Yu Su. 2023 · 2023
Later among the works it cites.
Finnsentiment: a finnish social media corpus for sentiment polarity annotation
Krister Lindén, Tommi Jauhiainen, and Sam Hardwick. 2023 · 2023
Later among the works it cites.
Findings of the babylm challenge: Sample-efficient pretraining on developmentally plausible corpora
Alex Warstadt, Aaron Mueller, Leshem Choshen, Ethan Wilcox, Chengxu Zhuang, Juan Ciro, Rafael Mosquera, Bhargavi Paranjabe, Adina Williams, Tal Linzen, et al. 2023 · 2023
Later among the works it cites.