Fetching the paper…
Reading the bibliography…
Identifying near duplicates within large, noisy text corpora has a myriad of applications that range from de-duplicating training datasets, reducing privacy risk, and evaluating test set leakage, to identifying reproduced news articles and literature within large corpora.
Comparing partitions
Lawrence Hubert and Phipps Arabie · 1985
Earlier work this paper cites.
Dimensionality reduction by learning an invariant mapping
Raia Hadsell, Sumit Chopra, and Yann LeCun · 2006
Earlier work this paper cites.
Chronicling america: Historic american newspapers
Jetta Culpepper · 2007
Earlier work this paper cites.
Reprinting, circulation, and the network author in antebellum newspapers
Ryan Cordell · 2015
Earlier work this paper cites.
Computational methods for uncovering reprinted texts in antebellum newspapers
David A Smith, Ryan Cordell, and Abby Mullen · 2015
Earlier work this paper cites.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Yukun Zhu, Ryan Kiros, Rich Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler · 2015
Earlier work this paper cites.
Newsprint Metropolis
Julia Guarneri · 2017
Earlier work this paper cites.
Mask r-cnn
Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross Girshick · 2017
Earlier work this paper cites.
Efficient natural language response suggestion for smart reply
Matthew Henderson, Rami Al-Rfou, Brian Strope, Yun-Hsuan Sung, László Lukács, Ruiqi Guo, Sanjiv Kumar, Balint Miklos, and Ray Kurzweil · 2017
Earlier work this paper cites.
Quantifying the effects of text duplication on semantic models
Alexandra Schofield, Laure Thompson, and David Mimno · 2017
Earlier work this paper cites.
Deduplication in a massive clinical note dataset
Sanjeev Shenoy, Tsung-Ting Kuo, Rodney Gabriel, Julian McAuley, and Chun-Nan Hsu · 2017
Earlier work this paper cites.
A system for identifying and exploring text repetition in large historical document corpora
Aleksi Vesanto, Filip Ginter, Hannu Salmi, Asko Nivala, and Tapio Salakoski · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Denoising distantly supervised open-domain question answering
Yankai Lin, Haozhe Ji, Zhiyuan Liu, and Maosong Sun · 2018
Earlier work this paper cites.
R 3: Reinforced ranker-reader for open-domain question answering
Shuohang Wang, Mo Yu, Xiaoxiao Guo, Zhiguo Wang, Tim Klinger, Wei Zhang, Shiyu Chang, Gerry Tesauro, Bowen Zhou, and Jing Jiang · 2018
Cited alongside, same era.
The adverse effects of code duplication in machine learning models of code
Miltiadis Allamanis · 2019
Cited alongside, same era.
Aggregating the news: Secondhand Knowledge and the Erosion of Journalistic Authority
Mark Coddington · 2019
Cited alongside, same era.
Billion-scale similarity search with gpus
Jeff Johnson, Matthijs Douze, and Hervé Jégou · 2019
Cited alongside, same era.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Cited alongside, same era.
Mining of massive data sets
Jure Leskovec, Anand Rajaraman, and Jeffrey David Ullman · 2020
Later among the works it cites.
Superglue: Learning feature matching with graph neural networks
Paul-Edouard Sarlin, Daniel DeTone, Tomasz Malisiewicz, and Andrew Rabinovich · 2020
Later among the works it cites.
Mpnet: Masked and permuted pre-training for language understanding
Kaitao Song, Xu Tan, Tao Qin, Jianfeng Lu, and Tie-Yan Liu · 2020
Later among the works it cites.
Jack Bandy and Nicholas Vincent · 2021
Later among the works it cites.
Documenting large webtext corpora: A case study on the colossal clean crawled corpus
Jesse Dodge, Maarten Sap, Ana Marasović, William Agnew, Gabriel Ilharco, Dirk Groeneveld, Margaret Mitchell, and Matt Gardner · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu · 2019
Cited alongside, same era.
Sentence-bert: Sentence embeddings using siamese bert-networks
Nils Reimers and Iryna Gurevych · 2019
Cited alongside, same era.
Scalable zero-shot entity linking with dense entity retrieval
Ledell Wu, Fabio Petroni, Martin Josifoski, Sebastian Riedel, and Luke Zettlemoyer · 2019
Cited alongside, same era.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Cited alongside, same era.
Deduplication of scholarly documents using locality sensitive hashing and word embeddings
Bikash Gyawali, Lucas Anastasiou, and Petr Knoth · 2020
Cited alongside, same era.
Dense passage retrieval for open-domain question answering
Vladimir Karpukhin, Barlas Oğuz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih · 2020
Cited alongside, same era.
Deduplicating training data makes language models better
Katherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang, Douglas Eck, Chris Callison-Burch, and Nicholas Carlini · 2021
Later among the works it cites.
Layoutparser: A unified toolkit for deep learning based document image analysis
Zejiang Shen, Ruochen Zhang, Melissa Dell, Benjamin Charles Germain Lee, Jacob Carlson, and Weining Li · 2021
Later among the works it cites.
Beir: A heterogenous benchmark for zero-shot evaluation of information retrieval models
Nandan Thakur, Nils Reimers, Andreas Rücklé, Abhishek Srivastava, and Iryna Gurevych · 2021
Later among the works it cites.
Improving language models by retrieving from trillions of tokens
Sebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai, Eliza Rutherford, Katie Millican, George Bm Van Den Driessche, Jean-Baptiste Lespiau, Bogdan Damoc, Aidan Clark, et al · 2022
Closest in time.
Quantifying memorization across neural language models
Nicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee, Florian Tramer, and Chiyuan Zhang · 2022
Closest in time.
Deduplicating training data mitigates privacy risks in language models
Nikhil Kandpal, Eric Wallace, and Colin Raffel · 2022
Closest in time.
Do language models plagiarize?
Jooyoung Lee, Thai Le, Jinghui Chen, and Dongwon Lee · 2022
Closest in time.
“note bloat” impacts deep learning-based nlp models for clinical prediction tasks
Jinghui Liu, Daniel Capurro, Anthony Nguyen, and Karin Verspoor · 2022
Closest in time.