Fetching the paper…
Reading the bibliography…
This paper introduces RETSim (Resilient and Efficient Text Similarity), a lightweight, multilingual deep learning model trained to produce robust metric embeddings for near-duplicate text retrieval, clustering, and dataset deduplication tasks.
Multilingual Universal Sentence Encoder for Semantic Retrieval, July 2019
Yinfei Yang, Daniel Cer, Amin Ahmad, Mandy Guo, Jax Law, Noah Constant, Gustavo Hernandez Abrego, Steve Yuan, Chris Tar, Yun-Hsuan Sung, Brian Strope, and Ray Kurzweil · 1907
Earlier work this paper cites.
Min-wise independent permutations
Andrei Z. Broder, Moses Charikar, Alan M. Frieze, and Michael Mitzenmacher · 1998
Earlier work this paper cites.
Similarity estimation techniques from rounding algorithms
Moses S. Charikar · 2002
Earlier work this paper cites.
John X. Morris, Eli Lifland, Jin Yong Yoo, Jake Grigsby, Di Jin, and Yanjun Qi · 2005
Earlier work this paper cites.
Language-agnostic BERT Sentence Embedding, March 2022
Fangxiaoyu Feng, Yinfei Yang, Daniel Cer, Naveen Arivazhagan, and Wei Wang · 2007
Earlier work this paper cites.
Near Duplicate Text Detection Using Frequency-Biased Signatures
Yifang Sun, Jianbin Qin, and Wei Wang · 2013
Earlier work this paper cites.
Analysis of phishing attacks and countermeasures, 2014
B. Issac, R. Chiong, and S. M. Jacob · 2014
Earlier work this paper cites.
Deep Unordered Composition Rivals Syntactic Methods for Text Classification
Mohit Iyyer, Varun Manjunatha, Jordan Boyd-Graber, and Hal Daumé Iii · 2015
Earlier work this paper cites.
FaceNet: A Unified Embedding for Face Recognition and Clustering
Florian Schroff, Dmitry Kalenichenko, and James Philbin · 2015
Earlier work this paper cites.
A large-scale query spelling correction corpus
Matthias Hagen, Martin Potthast, Marcel Gohsen, Anja Rathgeber, and Benno Stein · 2017
Earlier work this paper cites.
Generating Natural Language Adversarial Examples, September 2018
Moustafa Alzantot, Yash Sharma, Ahmed Elgohary, Bo-Jhang Ho, Mani Srivastava, and Kai-Wei Chang · 2018
Earlier work this paper cites.
Daniel Cer, Yinfei Yang, Sheng-yi Kong, Nan Hua, Nicole Limtiaco, Rhomni St John, Noah Constant, Mario Guajardo-Cespedes, Steve Yuan, and Chris Tar · 2018
Cited alongside, same era.
Black-box Generation of Adversarial Text Sequences to Evade Deep Learning Classifiers, May 2018
Ji Gao, Jack Lanchantin, Mary Lou Soffa, and Yanjun Qi · 2018
Cited alongside, same era.
Fine-tuning CNN image retrieval with no human annotation
Filip Radenović, Giorgos Tolias, and Ondřej Chum · 2018
Cited alongside, same era.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding, May 2019
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
Sentence-bert: Sentence embeddings using siamese bert-networks
Transformer quality in linear time
Weizhe Hua, Zihang Dai, Hanxiao Liu, and Quoc Le · 2022
Later among the works it cites.
Deduplicating Training Data Mitigates Privacy Risks in Language Models
Nikhil Kandpal, Eric Wallace, and Colin Raffel · 2022
Later among the works it cites.
The Stack: 3 TB of permissively licensed source code, November 2022
Denis Kocetkov, Raymond Li, Loubna Ben Allal, Jia Li, Chenghao Mou, Carlos Muñoz Ferrandis, Yacine Jernite, Margaret Mitchell, Sean Hughes, Thomas Wolf, Dzmitry Bahdanau, Leandro von Werra, and Harm de Vries · 2022
Later among the works it cites.
Deduplicating Training Data Makes Language Models Better, March 2022
Katherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang, Douglas Eck, Chris Callison-Burch, and Nicholas Carlini · 2022
Later among the works it cites.
A convnet for the 2020s
Zhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer, Trevor Darrell, and Saining Xie · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Nils Reimers and Iryna Gurevych · 2019
Cited alongside, same era.
Multi-similarity loss with general pair weighting for deep metric learning
Xun Wang, Xintong Han, Weilin Huang, Dengke Dong, and Matthew R. Scott · 2019
Cited alongside, same era.
Large batch optimization for deep learning: Training bert in 76 minutes
Yang You, Jing Li, Sashank Reddi, Jonathan Hseu, Sanjiv Kumar, Srinadh Bhojanapalli, Xiaodan Song, James Demmel, Kurt Keutzer, and Cho-Jui Hsieh · 2019
Cited alongside, same era.
Wiki-40b: Multilingual language model dataset
Mandy Guo, Zihang Dai, Denny Vrandečić, and Rami Al-Rfou · 2020
Cited alongside, same era.
Deduplication of Scholarly Documents using Locality Sensitive Hashing and Word Embeddings
Bikash Gyawali, Lucas Anastasiou, and Petr Knoth · 2020
Cited alongside, same era.
mT5: A massively multilingual pre-trained text-to-text transformer
Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel · 2020
Cited alongside, same era.
Machine Learning Techniques for Spam Detection in Email and IoT Platforms: Analysis and Research Challenges
Naeem Ahmed, Rashid Amin, Hamza Aldabbas, Deepika Koundal, Bader Alouffi, and Tariq Shah · 2022
Cited alongside, same era.
MTEB: Massive Text Embedding Benchmark, October 2022
Niklas Muennighoff, Nouamane Tazi, Loïc Magne, and Nils Reimers · 2022
Later among the works it cites.
Noise-Robust De-Duplication at Scale, October 2022
Emily Silcock, Luca D’Amico-Wong, Jinglin Yang, and Melissa Dell · 2022
Later among the works it cites.
Text Embeddings by Weakly-Supervised Contrastive Pre-training, December 2022
Liang Wang, Nan Yang, Xiaolong Huang, Binxing Jiao, Linjun Yang, Daxin Jiang, Rangan Majumder, and Furu Wei · 2022
Later among the works it cites.
PaLM 2 Technical Report, September 2023
Rohan Anil, Andrew M. Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, Eric Chu, Jonathan H. Clark, Laurent El Shafey, Yanping Huang, Kathy Meier-Hellstern, Gaurav Mishra, Erica Moreira, Mark Omernick, Kevin Robinson, Sebastian Ruder, Yi Tay, Kefan Xiao, Yuanzhong Xu, Yujing Zhang, Gustavo Hernandez Abrego, Junwhan Ahn, Jacob Austin, Paul Barham, Jan Botha, James Bradbury, Siddhartha Brahma, Kevin Brooks, Michele Catasta, Yong Cheng, Colin Cherry, Christopher A. Choquette-Choo, Aakanksha Chowdhery, Clément Crepy, Shachi Dave, Mostafa Dehghani, Sunipa Dev, Jacob Devlin, Mark Díaz, Nan Du, Ethan Dyer, Vlad Feinberg, Fangxiaoyu Feng, Vlad Fienber, Markus Freitag, Xavier Garcia, Sebastian Gehrmann, Lucas Gonzalez, Guy Gur-Ari, Steven Hand, Hadi Hashemi, Le Hou, Joshua Howland, Andrea Hu, Jeffrey Hui, Jeremy Hurwitz, Michael Isard, Abe Ittycheriah, Matthew Jagielski, Wenhao Jia, Kathleen Kenealy, Maxim Krikun, Sneha Kudugunta, Chang Lan, Katherine Lee, Benjamin Lee, Eric Li, Music Li, Wei Li, YaGuang Li, Jian Li, Hyeontaek Lim, Hanzhao Lin, Zhongtao Liu, Frederick Liu, Marcello Maggioni, Aroma Mahendru, Joshua Maynez, Vedant Misra, Maysam Moussalem, Zachary Nado, John Nham, Eric Ni, Andrew Nystrom, Alicia Parrish, Marie Pellat, Martin Polacek, Alex Polozov, Reiner Pope, Siyuan Qiao, Emily Reif, Bryan Richter, Parker Riley, Alex Castro Ros, Aurko Roy, Brennan Saeta, Rajkumar Samuel, Renee Shelby, Ambrose Slone, Daniel Smilkov, David R. So, Daniel Sohn, Simon Tokumine, Dasha Valter, Vijay Vasudevan, Kiran Vodrahalli, Xuezhi Wang, Pidong Wang, Zirui Wang, Tao Wang, John Wieting, Yuhuai Wu, Kelvin Xu, Yunhan Xu, Linting Xue, Pengcheng Yin, Jiahui Yu, Qiao Zhang, Steven Zheng, Ce Zheng, Weikang Zhou, Denny Zhou, Slav Petrov, and Yonghui Wu · 2023
Closest in time.
RetVec: Resilient and Efficient Text Vectorizer
Eli Bursztein, Marina Zhang, Owen vallis, Xinyu Jia, and Alexey Kurakin · 2023
Closest in time.
USearch by Unum Cloud, October 2023
Ash Vardanian · 2023
Closest in time.