Fetching the paper…
Reading the bibliography…
Metrics for set similarity are a core aspect of several data mining tasks.
The distribution of the flora in the alpine zone. 1
Paul Jaccard. 1912 · 1912
Earlier work this paper cites.
A statistical interpretation of term specificity and its application in retrieval
Karen Sparck Jones. 1972 · 1972
Earlier work this paper cites.
Sparse distributed memory
Pentti Kanerva. 1988 · 1988
Earlier work this paper cites.
On the resemblance and containment of documents. In Proceedings. Compression and Complexity of SEQUENCES 1997 (Cat. No. 97TB100171) . IEEE, 21–29
Andrei Z Broder. 1997 · 1997
Earlier work this paper cites.
Locality-preserving hashing in multidimensional spaces. In Proceedings of the Twenty-Ninth Annual ACM Symposium on Theory of Computing (STOC) . 618–625
Piotr Indyk, Rajeev Motwani, Prabhakar Raghavan, and Santosh Vempala. 1997 · 1997
Earlier work this paper cites.
Approximate nearest neighbors: towards removing the curse of dimensionality. In Proceedings of the Thirtieth Annual ACM Symposium on Theory of Computing (STOC) . 604–613
Piotr Indyk and Rajeev Motwani. 1998 · 1998
Earlier work this paper cites.
Min-wise independent permutations
Andrei Z Broder, Moses Charikar, Alan M Frieze, and Michael Mitzenmacher. 2000 · 2000
Earlier work this paper cites.
Similarity estimation techniques from rounding algorithms. In Proceedings of the thiry-fourth annual ACM Symposium on Theory of computing (STOC) . 380–388
Moses S Charikar. 2002 · 2002
Earlier work this paper cites.
Jianhan Zhu, Jun Hong, and John G Hughes. 2002 · 2002
Earlier work this paper cites.
Friends and neighbors on the web
Lada A Adamic and Eytan Adar. 2003 · 2003
Earlier work this paper cites.
Vector symbolic architectures answer Jackendoff’s challenges for cognitive neuroscience. In International Conference on Cognitive Science (ICCS) . 133–138
Ross W Gayler. 2003 · 2003
Earlier work this paper cites.
Query-free news search. In Proceedings of the 12th International Conference on World Wide Web (WWW) . 1–10
Monika Henzinger, Bay-Wei Chang, Brian Milch, and Sergey Brin. 2003 · 2003
Earlier work this paper cites.
Trust Management for the Semantic Web. In The Semantic Web - ISWC 2003 , Dieter Fensel, Katia Sycara, and John Mylopoulos (Eds.). Springer Berlin Heidelberg, Berlin, Heidelberg, 351–368
Matthew Richardson, Rakesh Agrawal, and Pedro Domingos. 2003 · 2003
Earlier work this paper cites.
Understanding inverse document frequency: on theoretical arguments for IDF
Stephen Robertson. 2004 · 2004
Earlier work this paper cites.
A network analysis model for disambiguation of names in lists
Bradley Malin, Edoardo Airoldi, and Kathleen M Carley. 2005 · 2005
Earlier work this paper cites.
WikiRelate! Computing semantic relatedness using Wikipedia. In AAAI , Vol. 6. 1419–1424
Michael Strube and Simone Paolo Ponzetto. 2006 · 2006
Earlier work this paper cites.
Google news personalization: scalable online collaborative filtering. In Proceedings of the 16th International Conference on World Wide Web (WWW) . 271–280
Abhinandan S Das, Mayur Datar, Ashutosh Garg, and Shyam Rajaram. 2007 · 2007
Earlier work this paper cites.
The link-prediction problem for social networks
David Liben-Nowell and Jon Kleinberg. 2007 · 2007
Earlier work this paper cites.
Detecting near-duplicates for web crawling. In Proceedings of the 16th International Conference on World Wide Web (WWW) . 141–150
Gurmeet Singh Manku, Arvind Jain, and Anish Das Sarma. 2007 · 2007
Earlier work this paper cites.
Improving similarity measures for short segments of text. In AAAI , Vol. 7. 1489–1494
Wen-tau Yih and Christopher Meek. 2007 · 2007
Earlier work this paper cites.
Near duplicate image detection: Min-hash and TF-IDF weighting.. In The British Machine Vision Conference (BMVC) , Vol. 810. 812–815
Ondrej Chum, James Philbin, Andrew Zisserman, et al · 2008
Earlier work this paper cites.
Jure Leskovec, Kevin J. Lang, Anirban Dasgupta, and Michael W. Mahoney. 2008 · 2008
Earlier work this paper cites.
A parallel algorithm for accurate dot product
Naoya Yamanaka, Takeshi Ogita, Siegfried M Rump, and Shin’ichi Oishi. 2008 · 2008
Earlier work this paper cites.
A graph-based semi-supervised algorithm for protein function prediction from interaction maps. In International Conference on Learning and Intelligent Optimization (LION) . Springer, 249–258
Valerio Freschi. 2009 · 2009
Earlier work this paper cites.
Hyperdimensional computing: An introduction to computing in distributed representation with high-dimensional random vectors
Pentti Kanerva. 2009 · 2009
Earlier work this paper cites.
Recommendation as link prediction: a graph kernel-based machine learning approach. In Proceedings of the 9th ACM/IEEE-CS joint conference on Digital Libraries . 213–216
Xin Li and Hsinchun Chen. 2009 · 2009
Cited alongside, same era.
A survey of binary similarity and distance measures
Seung-Seok Choi, Sung-Hyuk Cha, and Charles C Tappert. 2010 · 2010
Cited alongside, same era.
Improved consistent sampling, weighted minhash and l1 sketching. In 2010 IEEE International Conference on Data Mining (ICDM) . IEEE, 246–255
Sergey Ioffe. 2010 · 2010
Cited alongside, same era.
Synopses for massive data: Samples, histograms, wavelets, sketches
Graham Cormode, Minos Garofalakis, Peter J Haas, Chris Jermaine, et al · 2011
Cited alongside, same era.
Link prediction based on local information. In 2011 International Conference on Advances in Social Networks Analysis and Mining . IEEE, 382–386
Yuxiao Dong, Qing Ke, Bai Wang, and Bin Wu. 2011 · 2011
Cited alongside, same era.
Selectivity estimation for range predicates using lightweight models
Anshuman Dutt, Chi Wang, Azade Nazi, Srikanth Kandula, Vivek Narasayya, and Surajit Chaudhuri. 2019 · 2019
Later among the works it cites.
Uncertainty in big data analytics: survey, opportunities, and challenges
Reihaneh H Hariri, Erik M Fredericks, and Kate M Bowers. 2019 · 2019
Later among the works it cites.
Improving minhash via the containment index with applications to metagenomic analysis
David Koslicki and Hooman Zabeti. 2019 · 2019
Later among the works it cites.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al · 2019
Later among the works it cites.
Multi-scale Attributed Node Embedding
Benedek Rozemberczki, Carl Allen, and Rik Sarkar. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A survey of link prediction in social networks
Mohammad Al Hasan and Mohammed J Zaki. 2011 · 2011
Cited alongside, same era.
Link prediction in complex networks: A survey
Linyuan Lü and Tao Zhou. 2011 · 2011
Cited alongside, same era.
Fast computation of min-hash signatures for image collections. In 2012 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) . IEEE, 3077–3084
Ondřej Chum and Jiří Matas. 2012 · 2012
Cited alongside, same era.
Discrete mathematics and its applications
Kenneth H Rosen. 2012 · 2012
Cited alongside, same era.
Min-hash fingerprints for graph kernels: A trade-off among accuracy, efficiency, and compression
Carlos HC Teixeira, Arlei Silva, and Wagner Meira Jr. 2012 · 2012
Cited alongside, same era.
Network science
Albert-László Barabási. 2013 · 2013
Cited alongside, same era.
From link-prediction in brain connectomes and protein interactomes to the local-community-paradigm in complex networks
Carlo Vittorio Cannistraci, Gregorio Alanis-Lobato, and Timothy Ravasi. 2013 · 2013
Cited alongside, same era.
Cleaning noisy and heterogeneous metadata for record linking across scholarly big datasets. In Proceedings of the AAAI Conference on Artificial Intelligence (AAAI) , Vol. 33. 9601–9606
Athar Sefid, Jian Wu, C Ge Allen, Jing Zhao, Lu Liu, Cornelia Caragea, Prasenjit Mitra, and C Lee Giles. 2019 · 2019
Later among the works it cites.
STRING v11: protein–protein association networks with increased coverage, supporting functional discovery in genome-wide experimental datasets
Damian Szklarczyk, Annika L. Gable, David Lyon, Alexander Junge, Stefan Wyder, Jaime Huerta-Cepas, Milan Simonovic, Nadezhda Tsankova Doncheva, John H. Morris, Peer Bork, Lars Juhl Jensen, and Christian von Mering. 2019 · 2019
Later among the works it cites.
DartMinHash: Fast Sketching for Weighted Sets
Tobias Christiani. 2020 · 2020
Later among the works it cites.
Applications of link prediction in social networks: A review
Nur Nasuha Daud, Siti Hafizah Ab Hamid, Muntadher Saadoon, Firdaus Sahran, and Nor Badrul Anuar. 2020 · 2020
Later among the works it cites.
Deduplication of Scholarly Documents using Locality Sensitive Hashing and Word Embeddings. In Proceedings of 12th Language Resources and Evaluation Conference . France European Language Resources Association, 894–903
Bikash Gyawali, Lucas Anastasiou, and Petr Knoth. 2020 · 2020
Later among the works it cites.
Deep learning models for selectivity estimation of multi-attribute queries. In Proceedings of the 2020 ACM SIGMOD International Conference on Management of Data . 1035–1050
Shohedul Hasan, Saravanan Thirumuruganathan, Jees Augustine, Nick Koudas, and Gautam Das. 2020 · 2020
Later among the works it cites.
Open graph benchmark: Datasets for machine learning on graphs
Weihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong, Hongyu Ren, Bowen Liu, Michele Catasta, and Jure Leskovec. 2020 · 2020
Later among the works it cites.
Mining of massive data sets
Jure Leskovec, Anand Rajaraman, and Jeffrey David Ullman. 2020 · 2020
Later among the works it cites.
Deep learning for generic object detection: A survey
Li Liu, Wanli Ouyang, Xiaogang Wang, Paul Fieguth, Jie Chen, Xinwang Liu, and Matti Pietikäinen. 2020 · 2020
Later among the works it cites.
A survey on performance metrics for object-detection algorithms. In 2020 International Conference on Systems, Signals, and Image Processing (IWSSIP) . IEEE, 237–242
Rafael Padilla, Sergio L Netto, and Eduardo AB Da Silva. 2020 · 2020
Later among the works it cites.
A review for weighted minhash algorithms
Wei Wu, Bin Li, Ling Chen, Junbin Gao, and Chengqi Zhang. 2020 · 2020
Later among the works it cites.
A comparison of 71 binary similarity coefficients: The effect of base rates
Michael Brusco, J Dennis Cradit, and Douglas Steinley. 2021 · 2021
Later among the works it cites.
Deduplicating training data makes language models better
Katherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang, Douglas Eck, Chris Callison-Burch, and Nicholas Carlini. 2021 · 2021
Later among the works it cites.
Rejection sampling for weighted jaccard similarity revisited. In Proceedings of the AAAI Conference on Artificial Intelligence (AAAI) , Vol. 35. 4197–4205
Xiaoyun Li and Ping Li. 2021 · 2021
Later among the works it cites.
Parallel Architectures for Hyperdimensional Computing
Ryan Moughan. 2021 · 2021
Later among the works it cites.
Are we ready for learned cardinality estimation?
Xiaoying Wang, Changbo Qu, Weiyuan Wu, Jiannan Wang, and Qingqing Zhou. 2021 · 2021
Later among the works it cites.
Duplicates Detection
Abdulrazzak Ali. 2022 · 2022
Later among the works it cites.
Torchhd: An Open-Source Python Library to Support Hyperdimensional Computing Research
Mike Heddes, Igor Nunes, Pere Vergés, Dheyay Desai, Tony Givargis, and Alexandru Nicolau. 2022 · 2022
Later among the works it cites.
Learned cardinality estimation: An in-depth study. In Proceedings of the 2022 International Conference on Management of Data . 1214–1227
Kyoungmin Kim, Jisung Jung, In Seo, Wook-Shin Han, Kangwoo Choi, and Jaehyok Chong. 2022 · 2022
Later among the works it cites.