Fetching the paper…
Reading the bibliography…
Finding joinable tables in data lakes is key procedure in many applications such as data integration, data augmentation, data analysis, and data market.
Estimating the recall performance of web search engines
C. S. J. and W. Peter · 1997
Earlier work this paper cites.
Integration of heterogeneous databases without common domains using queries based on textual similarity
W. W. Cohen · 1998
Earlier work this paper cites.
Efficient query evaluation using a two-level retrieval process
A. Z. Broder, D. Carmel, M. Herscovici, A. Soffer, and J. Y. Zien · 2003
Earlier work this paper cites.
Efficient index-based KNN join processing for high-dimensional data
C. Yu, B. Cui, S. Wang, and J. Su · 2007
Earlier work this paper cites.
Product quantization for nearest neighbor search
H. Jégou, M. Douze, and C. Schmid · 2011
Earlier work this paper cites.
Pivot selection: Dimension reduction for distance-based indexing
R. Mao, W. L. Miranker, and D. P. Miranker · 2012
Earlier work this paper cites.
Quicker similarity joins in metric spaces
K. Fredriksson and B. Braithwaite · 2013
Earlier work this paper cites.
Extreme pivots for faster metric indexes
G. Ruiz, F. Santoyo, E. Chávez, K. Figueroa, and E. S. Tellez · 2013
Earlier work this paper cites.
Extending string similarity join to tolerant fuzzy token matching
J. Wang, G. Li, and J. Feng · 2014
Earlier work this paper cites.
SEMA-JOIN: joining semantically-related tables using big table corpora
Y. He, K. Ganjam, and X. Chu · 2015
Earlier work this paper cites.
Faster cover trees
M. Izbicki and C. R. Shelton · 2015
Earlier work this paper cites.
Learning generalized linear models over normalized data
A. Kumar, J. F. Naughton, and J. M. Patel · 2015
Earlier work this paper cites.
WDC web table corpus
D. Ritze, O. Lehmberg, R. Meusel, C. Bizer, and S. Zope · 2015
Earlier work this paper cites.
An empirical evaluation of set similarity join techniques
W. Mann, N. Augsten, and P. Bouros · 2016
Cited alongside, same era.
LSH ensemble: Internet-scale domain search
E. Zhu, F. Nargesian, K. Q. Pu, and R. J. Miller · 2016
Cited alongside, same era.
Efficient metric indexing for similarity search and similarity joins
L. Chen, Y. Gao, X. Li, C. S. Jensen, and G. Chen · 2017
Cited alongside, same era.
Pivot-based metric indexing
L. Chen, Y. Gao, B. Zheng, C. S. Jensen, H. Yang, and K. Yang · 2017
Cited alongside, same era.
The data civilizer system
D. Deng, R. C. Fernandez, Z. Abedjan, S. Wang, M. Stonebraker, A. K. Elmagarmid, I. F. Ilyas, S. Madden, M. Ouzzani, and N. Tang · 2017
Cited alongside, same era.
Silkmoth: An efficient method for finding related sets with maximum matching constraints
Mf-join: Efficient fuzzy string similarity join with multi-level filtering
J. Wang, C. Lin, and C. Zaniolo · 2019
Later among the works it cites.
JOSIE: overlap set similarity search for finding joinable tables in data lakes
E. Zhu, D. Deng, F. Nargesian, and R. J. Miller · 2019
Later among the works it cites.
Dataset discovery in data lakes
A. Bogatu, A. A. A. Fernandes, N. W. Paton, and N. Konstantinou · 2020
Closest in time.
ARDA: automatic relational data augmentation for machine learning
N. Chepurko, R. Marcus, E. Zgraggen, R. C. Fernandez, T. Kraska, and D. Karger · 2020
Closest in time.
fastText: Library for efficient text classification and representation learning
Facebook AI Research Lab · 2020
Closest in time.
Company classification
Kaggle · 2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
D. Deng, A. Kim, S. Madden, and M. Stonebraker · 2017
Cited alongside, same era.
Approximate string joins with abbreviations
W. Tao, D. Deng, and M. Stonebraker · 2017
Cited alongside, same era.
Auto-join: Joining tables by leveraging transformations
E. Zhu, Y. He, and S. Chaudhuri · 2017
Cited alongside, same era.
Seeping semantics: Linking datasets using word embeddings for data discovery
R. C. Fernandez, E. Mansour, A. A. Qahtan, A. K. Elmagarmid, I. F. Ilyas, S. Madden, M. Ouzzani, M. Stonebraker, and N. Tang · 2018
Cited alongside, same era.
Deep learning for entity matching: A design space exploration
S. Mudgal, H. Li, T. Rekatsinas, A. Doan, Y. Park, G. Krishnan, R. Deep, E. Arcaute, and V. Raghavendra · 2018
Cited alongside, same era.
Table union search on open data
F. Nargesian, E. Zhu, K. Q. Pu, and R. J. Miller · 2018
Cited alongside, same era.
Pigeonring: A principle for faster thresholded similarity search
J. Qin and C. Xiao · 2018
Cited alongside, same era.
Toy products on amazon
Kaggle · 2020
Closest in time.
Video game sales
Kaggle · 2020
Closest in time.
Nano product quantization
Y. Matsui · 2020
Closest in time.
Similarity query processing for high-dimensional data
J. Qin, W. Wang, C. Xiao, and Y. Zhang · 2020
Closest in time.
Coveratree
P. Varilly · 2020
Closest in time.
Sato: Contextual semantic type detection in tables
D. Zhang, Y. Suhara, J. Li, M. Hulsebos, Ç. Demiralp, and W. Tan · 2020
Closest in time.
Finding related tables in data lakes for interactive data science
Y. Zhang and Z. G. Ives · 2020
Closest in time.