Fetching the paper…
Reading the bibliography…
Language tasks involving character-level manipulations (e.g., spelling corrections, arithmetic operations, word games) are challenging for models operating on subword units.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 1901
Earlier work this paper cites.
RoBERTa: A robustly optimized BERT pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Wordnet: A lexical database for English
George A Miller. 1995 · 1995
Earlier work this paper cites.
Japanese and korean voice search
Mike Schuster and Kaisuke Nakajima. 2012 · 2012
Earlier work this paper cites.
Learning character-level representations for part-of-speech tagging
Cicero Dos Santos and Bianca Zadrozny. 2014 · 2014
Earlier work this paper cites.
Achieving open vocabulary neural machine translation with hybrid word-character models
Minh-Thang Luong and Christopher D. Manning. 2016 · 2016
Earlier work this paper cites.
End-to-end sequence labeling via bi-directional LSTM-CNNs-CRF
Xuezhe Ma and Eduard Hovy. 2016 · 2016
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Earlier work this paper cites.
Mimicking word embeddings using subword RNNs
Yuval Pinter, Robert Guthrie, and Jacob Eisenstein. 2017 · 2017
Earlier work this paper cites.
Synthetic and natural noise both break neural machine translation
Yonatan Belinkov and Yonatan Bisk. 2018 · 2018
Earlier work this paper cites.
Subword regularization: Improving neural network translation models with multiple subword candidates
Taku Kudo. 2018 · 2018
Earlier work this paper cites.
Deep contextualized word representations
Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018 · 2018
Earlier work this paper cites.
Abstracting causal models
Sander Beckers and Joseph Y. Halpern. 2019 · 2019
Earlier work this paper cites.
Attentive mimicking: Better word embeddings by attending to informative contexts
Timo Schick and Hinrich Schütze. 2019 · 2019
Earlier work this paper cites.
Approximate causal abstractions
Sander Beckers, Frederick Eberhardt, and Joseph Y. Halpern. 2020 · 2020
Cited alongside, same era.
Byte pair encoding is suboptimal for language model pretraining
Kaj Bostrom and Greg Durrett. 2020 · 2020
Cited alongside, same era.
CharacterBERT: Reconciling ELMo and BERT for word-level open-vocabulary representations from characters
Hicham El Boukkouri, Olivier Ferret, Thomas Lavergne, Hiroshi Noji, Pierre Zweigenbaum, and Jun’ichi Tsujii. 2020 · 2020
Cited alongside, same era.
Injecting numerical reasoning skills into language models
Mor Geva, Ankit Gupta, and Jonathan Berant. 2020 · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020 · 2020
Cited alongside, same era.
Homophonic pun generation with lexically constrained rewriting
Noisy UGC translation at the character level: Revisiting open-vocabulary capabilities and robustness of char-based models
José Carlos Rosales Núñez, Guillaume Wisniewski, and Djamé Seddah. 2021 · 2021
Later among the works it cites.
Decrypting cryptic crosswords: Semantically complex wordplay puzzles as a target for NLP
Joshua Rozner, Christopher Potts, and Kyle Mahowald. 2021 · 2021
Later among the works it cites.
GPT-NeoX-20B: An open-source autoregressive language model
Sid Black, Stella Biderman, Eric Hallahan, Quentin Anthony, Leo Gao, Laurence Golding, Horace He, Connor Leahy, Kyle McDonell, Jason Phang, et al. 2022 · 2022
Closest in time.
Canine: Pre-training an Efficient Tokenization-Free Encoder for Language Representation
Jonathan H. Clark, Dan Garrette, Iulia Turc, and John Wieting. 2022 · 2022
Closest in time.
Models in a spelling bee: Language models implicitly learn the character composition of tokens
Itay Itzhak and Omer Levy. 2022 · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zhiwei Yu, Hongyu Zang, and Xiaojun Wan. 2020 · 2020
Cited alongside, same era.
Char2Subword: Extending the subword embedding space using robust character compositionality
Gustavo Aguilar, Bryan McCann, Tong Niu, Nazneen Rajani, Nitish Shirish Keskar, and Thamar Solorio. 2021 · 2021
Cited alongside, same era.
Cryptonite: A cryptic crossword benchmark for extreme ambiguity in language
Avia Efrat, Uri Shaham, Dan Kilman, and Omer Levy. 2021 · 2021
Cited alongside, same era.
Causal abstractions of neural networks
Atticus Geiger, Hanson Lu, Thomas Icard, and Christopher Potts. 2021 · 2021
Cited alongside, same era.
Deberta: decoding-enhanced bert with disentangled attention
Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen. 2021 · 2021
Cited alongside, same era.
Why don’t people use character-level machine translation?
Jindřich Libovickỳ, Helmut Schmid, and Alexander Fraser. 2021 · 2021
Cited alongside, same era.
Between words and characters: A brief history of open-vocabulary modeling and tokenization in NLP
Sabrina J Mielke, Zaid Alyafeai, Elizabeth Salesky, Colin Raffel, Manan Dey, Matthias Gallé, Arun Raja, Chenglei Si, Wilson Y Lee, Benoît Sagot, et al. 2021 · 2021
Cited alongside, same era.
Closest in time.
What do tokens know about their characters and how do they know it?
Ayush Kaushal and Kyle Mahowald. 2022 · 2022
Closest in time.
Ambipun: Generating humorous puns with ambiguous context
Anirudh Mittal, Yufei Tian, and Nanyun Peng. 2022 · 2022
Closest in time.
BLOOM: A 176b-parameter open-access multilingual language model
Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilić, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, et al. 2022 · 2022
Closest in time.
Charformer: Fast character transformers via gradient-based subword tokenization
Yi Tay, Vinh Q. Tran, Sebastian Ruder, Jai Gupta, Hyung Won Chung, Dara Bahri, Zhen Qin, Simon Baumgartner, Cong Yu, and Donald Metzler. 2022 · 2022
Closest in time.
Automated crossword solving
Eric Wallace, Nicholas Tomlin, Albert Xu, Kevin Yang, Eshaan Pathak, Matthew Ginsberg, and Dan Klein. 2022 · 2022
Closest in time.
Finetuned language models are zero-shot learners
Jason Wei, Maarten Bosma, Vincent Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V Le. 2022 · 2022
Closest in time.
Causal Proxy Models for concept-based model explanations
Zhengxuan Wu, Karel D’Oosterlinck, Atticus Geiger, Amir Zur, and Christopher Potts. 2022 · 2022
Closest in time.
ByT5: Towards a Token-Free Future with Pre-trained Byte-to-Byte Models
Linting Xue, Aditya Barua, Noah Constant, Rami Al-Rfou, Sharan Narang, Mihir Kale, Adam Roberts, and Colin Raffel. 2022 · 2022
Closest in time.
OPT: Open Pre-trained Transformer language models
Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al. 2022 · 2022
Closest in time.