Fetching the paper…
Reading the bibliography…
We find that existing language modeling datasets contain many near-duplicate examples and long repetitive substrings.
Transformer-xl: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc V Le, and Ruslan Salakhutdinov. 2019 · 1901
Earlier work this paper cites.
Defending against neural fake news
Rowan Zellers, Ari Holtzman, Hannah Rashkin, Yonatan Bisk, Ali Farhadi, Franziska Roesner, and Yejin Choi. 2019 · 1905
Earlier work this paper cites.
Universal adversarial triggers for attacking and analyzing nlp
Eric Wallace, Shi Feng, Nikhil Kandpal, Matt Gardner, and Sameer Singh. 2019 · 1908
Earlier work this paper cites.
The distribution of the flora in the alpine zone
Paul Jaccard. 1912 · 1912
Earlier work this paper cites.
Space/time trade-offs in hash coding with allowable errors
Burton H Bloom. 1970 · 1970
Earlier work this paper cites.
Linear pattern matching algorithms
Peter Weiner. 1973 · 1973
Earlier work this paper cites.
Suffix arrays: a new method for on-line string searches
Udi Manber and Gene Myers. 1993 · 1993
Earlier work this paper cites.
On the resemblance and containment of documents
Andrei Z Broder. 1997 · 1997
Earlier work this paper cites.
Using suffix arrays to compute term frequency and document frequency for all substrings in a corpus
Mikio Yamamoto and Kenneth W Church. 2001 · 2001
Earlier work this paper cites.
English gigaword
David Graff, Junbo Kong, Ke Chen, and Kazuaki Maeda. 2003 · 2003
Earlier work this paper cites.
Simple linear work suffix array construction
Juha Kärkkäinen and Peter Sanders. 2003 · 2003
Earlier work this paper cites.
Space efficient linear time construction of suffix arrays
Pang Ko and Srinivas Aluru. 2003 · 2003
Earlier work this paper cites.
Towards controllable biases in language generation
Emily Sheng, Kai-Wei Chang, Premkumar Natarajan, and Nanyun Peng. 2020 · 2005
Earlier work this paper cites.
A new suffix tree similarity measure for document clustering
Hung Chim and Xiaotie Deng. 2007 · 2007
Earlier work this paper cites.
Linear suffix array construction by almost pure induced-sorting
Ge Nong, Sen Zhang, and Wai Hong Chan. 2009 · 2009
Earlier work this paper cites.
mt5: A massively multilingual pre-trained text-to-text transformer
Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. 2020 · 2010
Earlier work this paper cites.
Not just bigger: Towards better-quality web corpora
Yannick Versley and Yana Panchenko. 2012 · 2012
Cited alongside, same era.
One billion word benchmark for measuring progress in statistical language modeling
Ciprian Chelba, Tomas Mikolov, Mike Schuster, Qi Ge, Thorsten Brants, Phillipp Koehn, and Tony Robinson. 2013 · 2013
Cited alongside, same era.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Yukun Zhu, Ryan Kiros, Rich Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler. 2015 · 2015
Cited alongside, same era.
Min-hash sketches: A brief survey
Edith Cohen. 2016 · 2016
Cited alongside, same era.
A closer look at memorization in deep networks
Devansh Arpit, Stanisław Jastrzębski, Nicolas Ballas, David Krueger, Emmanuel Bengio, Maxinder S Kanwal, Tegan Maharaj, Asja Fischer, Aaron Courville, Yoshua Bengio, et al. 2017 · 2017
Cited alongside, same era.
Language models are few-shot learners
Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020 · 2020
Later among the works it cites.
Extracting training data from large language models
Nicholas Carlini, Florian Tramer, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, Ulfar Erlingsson, Alina Oprea, and Colin Raffel. 2020 · 2020
Later among the works it cites.
What neural networks memorize and why: Discovering the long tail via influence estimation
Vitaly Feldman and Chiyuan Zhang. 2020 · 2020
Later among the works it cites.
The Pile: An 800gb dataset of diverse text for language modeling
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser, and Connor Leahy. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Membership inference attacks against machine learning models
Reza Shokri, Marco Stronati, Congzheng Song, and Vitaly Shmatikov. 2017 · 2017
Cited alongside, same era.
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Data statements for natural language processing: Toward mitigating system bias and enabling better science
Emily M. Bender and Batya Friedman. 2018 · 2018
Cited alongside, same era.
Identifying and characterizing highly similar notes in big clinical note datasets
Rodney A. Gabriel, Tsung-Ting Kuo, Julian McAuley, and Chun-Nan Hsu. 2018 · 2018
Cited alongside, same era.
Adafactor: Adaptive learning rates with sublinear memory cost
Noam Shazeer and Mitchell Stern. 2018 · 2018
Cited alongside, same era.
A simple method for commonsense reasoning
Trieu H Trinh and Quoc V Le. 2018 · 2018
Cited alongside, same era.
Connected components at scale via local contractions
Jakub Łącki, Vahab Mirrokni, and Michał Włodarczyk. 2018 · 2018
Cited alongside, same era.
Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumé III au2, and Kate Crawford. 2020 · 2020
Later among the works it cites.
Wiki-40b: Multilingual language model dataset
Mandy Guo, Zihang Dai, Denny Vrandecic, and Rami Al-Rfou. 2020 · 2020
Later among the works it cites.
Deduplication of scholarly documents using locality sensitive hashing and word embeddings
Bikash Gyawali, Lucas Anastasiou, and Petr Knoth. 2020 · 2020
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020 · 2020
Later among the works it cites.
Jack Bandy and Nicholas Vincent. 2021 · 2021
Closest in time.
On the dangers of stochastic parrots: Can language models be too big?
Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021 · 2021
Closest in time.
GPT-Neo: Large scale autoregressive language modeling with mesh-tensorflow
Sid Black, Leo Gao, Phil Wang, Connor Leahy, and Stella Biderman. 2021 · 2021
Closest in time.
Carbon emissions and large neural network training
David Patterson, Joseph Gonzalez, Quoc Le, Chen Liang, Lluis-Miquel Munguia, Daniel Rothchild, David So, Maud Texier, and Jeff Dean. 2021 · 2021
Closest in time.
On the geometry of generalization and memorization in deep neural networks
Cory Stephenson, Suchismita Padhy, Abhinav Ganesh, Yue Hui, Hanlin Tang, and SueYeon Chung. 2021 · 2021
Closest in time.
Understanding invariance via feedforward inversion of discriminatively trained classifiers
Piotr Teterwak, Chiyuan Zhang, Dilip Krishnan, and Michael C Mozer. 2021 · 2021
Closest in time.
Wei Zeng, Xiaozhe Ren, Teng Su, Hui Wang, Yi Liao, Zhiwei Wang, Xin Jiang, ZhenZhang Yang, Kaisheng Wang, Xiaoda Zhang, Chen Li, Ziyan Gong, Yifan Yao, Xinjing Huang, Jun Wang, Jianfeng Yu, Qi Guo, Yue Yu, Yan Zhang, Jin Wang, Hengtao Tao, Dasen Yan, Zexuan Yi, Fang Peng, Fangqing Jiang, Han Zhang, Lingfeng Deng, Yehong Zhang, Zhe Lin, Chao Zhang, Shaojie Zhang, Mingyue Guo, Shanzhi Gu, Gaojun Fan, Yaowei Wang, Xuefeng Jin, Qun Liu, and Yonghong Tian. 2021 · 2021
Closest in time.
What does it mean for a language model to preserve privacy?
Hannah Brown, Katherine Lee, Fatemehsadat Mireshghallah, Reza Shokri, and Florian Tramèr. 2022 · 2022
Closest in time.