Fetching the paper…
Reading the bibliography…
Due to the usefulness in data enrichment for data analysis tasks, joinable table discovery has become an important operation in data lake management.
On the resemblance and containment of documents
A. Z. Broder · 1997
Earlier work this paper cites.
Estimating the recall performance of web search engines
C. S. J. and W. Peter · 1997
Earlier work this paper cites.
A primitive operator for similarity joins in data cleaning
S. Chaudhuri, V. Ganti, and R. Kaushik · 2006
Earlier work this paper cites.
Introduction to information retrieval
C. D. Manning, P. Raghavan, and H. Schütze · 2008
Earlier work this paper cites.
Product quantization for nearest neighbor search
H. Jégou, M. Douze, and C. Schmid · 2011
Earlier work this paper cites.
Efficient similarity joins for near-duplicate detection
C. Xiao, W. Wang, X. Lin, J. X. Yu, and G. Wang · 2011
Earlier work this paper cites.
Principles of Data Integration
A. Doan, A. Y. Halevy, and Z. G. Ives · 2012
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
T. Mikolov, I. Sutskever, K. Chen, G. S. Corrado, and J. Dean · 2013
Earlier work this paper cites.
Tabel: Entity linking in web tables
C. S. Bhagavatula, T. Noraset, and D. Downey · 2015
Earlier work this paper cites.
Wikitables
C. S. Bhagavatula, T. Noraset, and D. Downey · 2015
Earlier work this paper cites.
fastText: Library for efficient text classification and representation learning
Facebook AI Research Lab · 2015
Earlier work this paper cites.
SEMA-JOIN: joining semantically-related tables using big table corpora
Y. He, K. Ganjam, and X. Chu · 2015
Earlier work this paper cites.
WDC web table corpus
D. Ritze, O. Lehmberg, R. Meusel, C. Bizer, and S. Zope · 2015
Earlier work this paper cites.
Deep neural networks for youtube recommendations
P. Covington, J. Adams, and E. Sargin · 2016
Earlier work this paper cites.
LSH ensemble: Internet-scale domain search
E. Zhu, F. Nargesian, K. Q. Pu, and R. J. Miller · 2016
Earlier work this paper cites.
Pivot-based metric indexing
L. Chen, Y. Gao, B. Zheng, C. S. Jensen, H. Yang, and K. Yang · 2017
Earlier work this paper cites.
Silkmoth: An efficient method for finding related sets with maximum matching constraints
D. Deng, A. Kim, S. Madden, and M. Stonebraker · 2017
Earlier work this paper cites.
Efficient natural language response suggestion for smart reply
M. L. Henderson, R. Al-Rfou, B. Strope, Y. Sung, L. Lukács, R. Guo, S. Kumar, B. Miklos, and R. Kurzweil · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Auto-join: Joining tables by leveraging transformations
E. Zhu, Y. He, and S. Chaudhuri · 2017
Earlier work this paper cites.
Table union search on open data
F. Nargesian, E. Zhu, K. Q. Pu, and R. J. Miller · 2018
Earlier work this paper cites.
BERT: pre-training of deep bidirectional transformers for language understanding
J. Devlin, M. Chang, K. Lee, and K. Toutanova · 2019
Earlier work this paper cites.
Tabular cell classification using pre-trained cell embeddings
M. Ghasemi-Gol, J. Pujara, and P. A. Szekely · 2019
Cited alongside, same era.
Sherlock: A deep learning approach to semantic data type detection
M. Hulsebos, K. Z. Hu, M. A. Bakker, E. Zgraggen, A. Satyanarayan, T. Kraska, Ç. Demiralp, and C. A. Hidalgo · 2019
Cited alongside, same era.
Text generation from knowledge graphs with graph transformers
R. Koncel-Kedziorski, D. Bekal, Y. Luan, M. Lapata, and H. Hajishirzi · 2019
Cited alongside, same era.
Sentencetransformers.losses
Nils Reimers · 2019
Cited alongside, same era.
Sentence-bert: Sentence embeddings using siamese bert-networks
N. Reimers and I. Gurevych · 2019
Cited alongside, same era.
Distilbert, a distilled version of BERT: smaller, faster, cheaper and lighter
Sato: Contextual semantic type detection in tables
D. Zhang, Y. Suhara, J. Li, M. Hulsebos, Ç. Demiralp, and W. Tan · 2020
Later among the works it cites.
Finding related tables in data lakes for interactive data science
Y. Zhang and Z. G. Ives · 2020
Later among the works it cites.
Discovering related data at scale
S. Bharadwaj, P. Gupta, R. Bhagwan, and S. Guha · 2021
Later among the works it cites.
Efficient joinable table discovery in data lakes: A high-dimensional similarity-based approach
Y. Dong, K. Takeoka, C. Xiao, and M. Oyamada · 2021
Later among the works it cites.
Effective and scalable data discovery with nextiajd
J. Flores, S. Nadal, and O. Romero · 2021
Later among the works it cites.
TABBIE: pretrained representations of tabular data
H. Iida, D. Thai, V. Manjunatha, and M. Iyyer · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
V. Sanh, L. Debut, J. Chaumond, and T. Wolf · 2019
Cited alongside, same era.
Table2vec: Neural word and entity embeddings for table population and retrieval
L. Zhang, S. Zhang, and K. Balog · 2019
Cited alongside, same era.
JOSIE: overlap set similarity search for finding joinable tables in data lakes
E. Zhu, D. Deng, F. Nargesian, and R. J. Miller · 2019
Cited alongside, same era.
Dataset discovery in data lakes
A. Bogatu, A. A. A. Fernandes, N. W. Paton, and N. Konstantinou · 2020
Cited alongside, same era.
Table search using a deep contextualized language model
Z. Chen, M. Trabelsi, J. Heflin, Y. Xu, and B. D. Davison · 2020
Cited alongside, same era.
ARDA: automatic relational data augmentation for machine learning
N. Chepurko, R. Marcus, E. Zgraggen, R. C. Fernandez, T. Kraska, and D. Karger · 2020
Cited alongside, same era.
TURL: table understanding through representation learning
X. Deng, H. Sun, A. Lees, Y. Wu, and C. Yu · 2020
Cited alongside, same era.
Valentine: Evaluating matching techniques for dataset discovery
C. Koutras, G. Siachamis, A. Ionescu, K. Psarakis, J. Brons, M. Fragkoulis, C. Lofi, A. Bonifati, and A. Katsifodimos · 2021
Later among the works it cites.
High-dimensional similarity query processing for data science
J. Qin, W. Wang, C. Xiao, Y. Zhang, and Y. Wang · 2021
Later among the works it cites.
Auto-validate: Unsupervised data validation using data-domain patterns inferred from data lakes
J. Song and Y. He · 2021
Later among the works it cites.
Ember: No-code context enrichment via similarity-based keyless joins
S. Suri, I. F. Ilyas, C. Ré, and T. Rekatsinas · 2021
Later among the works it cites.
RPT: relational pre-trained transformer is almost all you need towards democratizing data preparation
N. Tang, J. Fan, F. Li, J. Tu, X. Du, G. Li, S. Madden, and M. Ouzzani · 2021
Later among the works it cites.
Retrieving complex tables with multi-granular graph representation learning
F. Wang, K. Sun, M. Chen, J. Pujara, and P. A. Szekely · 2021
Later among the works it cites.
TUTA: tree-based transformers for generally structured table pre-training
Z. Wang, H. Dong, R. Jia, J. Li, Z. Fu, S. Han, and D. Zhang · 2021
Later among the works it cites.
https://huggingface.co/docs/transformers/index , 2022
Hugging face transformers · 2022
Closest in time.
Table enrichment system for machine learning
Y. Dong and M. Oyamada · 2022
Closest in time.
Faiss: Facebook ai similarity search
facebook AI research · 2022
Closest in time.
Integrating data lake tables
A. Khatiwada, R. Shraga, W. Gatterbauer, and R. J. Miller · 2022
Closest in time.
Can foundation models wrangle your data?
A. Narayan, I. Chami, L. J. Orr, and C. Ré · 2022
Closest in time.
Sentence bert
N. Reimers · 2022
Closest in time.
Annotating columns with pre-trained language models
Y. Suhara, J. Li, Y. Li, D. Zhang, Ç. Demiralp, C. Chen, and W. Tan · 2022
Closest in time.
Chain of thought prompting elicits reasoning in large language models
J. Wei, X. Wang, D. Schuurmans, M. Bosma, E. H. Chi, Q. Le, and D. Zhou · 2022
Closest in time.