Fetching the paper…
Reading the bibliography…
Byte-Pair Encoding (BPE) is a popular algorithm used for tokenizing data in NLP, despite being devised initially as a compression method.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 1901
Earlier work this paper cites.
Submodular functions and convexity
László Lovász. 1983 · 1983
Earlier work this paper cites.
A technique for high-performance data compression
Terry A. Welch. 1984 · 1984
Earlier work this paper cites.
A new algorithm for data compression
Philip Gage. 1994 · 1994
Earlier work this paper cites.
Elements of Information Theory , 2 edition
Thomas M. Cover and Joy A. Thomas. 2006 · 2006
Earlier work this paper cites.
Combinatorial Problems in Online Advertising
Azarakhsh Malekian. 2009 · 2009
Earlier work this paper cites.
Maximizing sequence-submodular functions and its application to online advertising
Saeed Alaei, Ali Makhdoumi, and Azarakhsh Malekian. 2010 · 2010
Earlier work this paper cites.
Tokenization and the noiseless channel
Vilém Zouhar, Clara Meister, Juan Gastaldi, Li Du, Mrinmaya Sachan, and Ryan Cotterell. 2023 · 2010
Cited alongside, same era.
Submodular function maximization
Andreas Krause and Daniel Golovin. 2014 · 2014
Cited alongside, same era.
String submodular functions with curvature constraints
Zhenliang Zhang, Edwin KP Chong, Ali Pezeshki, and William Moran. 2015 · 2015
Cited alongside, same era.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Cited alongside, same era.
SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing
Taku Kudo and John Richardson. 2018 · 2018
Cited alongside, same era.
A call for prudent choice of subword merge operations in neural machine translation
Shuoyang Ding, Adithya Renduchintala, and Kevin Duh. 2019 · 2019
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Later among the works it cites.
Byte pair encoding is suboptimal for language model pretraining
Kaj Bostrom and Greg Durrett. 2020 · 2020
Later among the works it cites.
CCAligned: A massive collection of cross-lingual web-document pairs
Ahmed El-Kishky, Vishrav Chaudhary, Francisco Guzmán, and Philipp Koehn. 2020 · 2020
Later among the works it cites.
DIALOGPT: Large-scale generative pre-training for conversational response generation
Yizhe Zhang, Siqi Sun, Michel Galley, Yen-Chun Chen, Chris Brockett, Xiang Gao, Jianfeng Gao, Jingjing Liu, and William B Dolan. 2020 · 2020
Later among the works it cites.
Submodularity in machine learning and artificial intelligence
Jeff Bilmes. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilić, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, et al. 2022 · 2022
Later among the works it cites.