Fetching the paper…
Reading the bibliography…
Dataset discovery from data lakes is essential in many real application scenarios.
Finding Related Tables in Data Lakes for Interactive Data Science. In SIGMOD . 1951–1966
Yi Zhang and Zachary G. Ives. 2020 · 1966
Earlier work this paper cites.
Data Lake Management: Challenges and Opportunities
Fatemeh Nargesian, Erkang Zhu, Renée J. Miller, Ken Q. Pu, and Patricia C. Arocena. 2019 · 1989
Earlier work this paper cites.
Similarity Search in High Dimensions via Hashing. In VLDB . Morgan Kaufmann, 518–529
Aristides Gionis, Piotr Indyk, and Rajeev Motwani. 1999 · 1999
Earlier work this paper cites.
Similarity estimation techniques from rounding algorithms. In STOC . 380–388
Moses Charikar. 2002 · 2002
Earlier work this paper cites.
WebTables: exploring the power of tables on the web
Michael J. Cafarella, Alon Y. Halevy, Daisy Zhe Wang, Eugene Wu, and Yang Zhang. 2008 · 2008
Earlier work this paper cites.
Efficient Merging and Filtering Algorithms for Approximate String Searches. In ICDE . 257–266
Chen Li, Jiaheng Lu, and Yiming Lu. 2008 · 2008
Earlier work this paper cites.
Introduction to information retrieval
Christopher D. Manning, Prabhakar Raghavan, and Hinrich Schütze. 2008 · 2008
Earlier work this paper cites.
Data Integration for the Relational Web
Michael J. Cafarella, Alon Y. Halevy, and Nodira Khoussainova. 2009 · 2009
Earlier work this paper cites.
Annotating and Searching Web Tables Using Entities, Types and Relationships
Girija Limaye, Sunita Sarawagi, and Soumen Chakrabarti. 2010 · 2010
Earlier work this paper cites.
Recovering Semantics of Tables on the Web
Petros Venetis, Alon Y. Halevy, Jayant Madhavan, Marius Pasca, Warren Shen, Fei Wu, Gengxin Miao, and Chung Wu. 2011 · 2011
Earlier work this paper cites.
Finding related tables. In SIGMOD . 817–828
Anish Das Sarma, Lujun Fang, Nitin Gupta, Alon Y. Halevy, Hongrae Lee, Fei Wu, Reynold Xin, and Cong Yu. 2012 · 2012
Earlier work this paper cites.
InfoGather: entity augmentation and attribute discovery by holistic matching with web tables. In SIGMOD . ACM, 97–108
Mohamed Yakout, Kris Ganjam, Kaushik Chakrabarti, and Surajit Chaudhuri. 2012 · 2012
Earlier work this paper cites.
Schema Extraction for Tabular Data on the Web
Marco D. Adelfio and Hanan Samet. 2013 · 2013
Earlier work this paper cites.
Synthesizing Union Tables from the Web. In IJCAI . 2677–2683
Xiao Ling, Alon Y. Halevy, Fei Wu, and Cong Yu. 2013 · 2013
Earlier work this paper cites.
XGBoost: A Scalable Tree Boosting System. In KDD . ACM, 785–794
Tianqi Chen and Carlos Guestrin. 2016 · 2016
Earlier work this paper cites.
A Large Public Corpus of Web Tables containing Time and Context Metadata. In WWW (Companion Volume) . ACM, 75–76
Oliver Lehmberg, Dominique Ritze, Robert Meusel, and Christian Bizer. 2016 · 2016
Earlier work this paper cites.
Visualizing Semantic Table Annotations with TableMiner+. In ISWC , Vol. 1690
Suvodeep Mazumdar and Ziqi Zhang. 2016 · 2016
Earlier work this paper cites.
LSH Ensemble: Internet-Scale Domain Search
Erkang Zhu, Fatemeh Nargesian, Ken Q. Pu, and Renée J. Miller. 2016 · 2016
Earlier work this paper cites.
Stitching Web Tables for Improving Matching Quality
Oliver Lehmberg and Christian Bizer. 2017 · 2017
Earlier work this paper cites.
Attention is All you Need. In NeurIPS . 5998–6008
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Effective and efficient Semantic Table Interpretation using TableMiner + {}^{\mbox{+}}
Ziqi Zhang. 2017 · 2017
Cited alongside, same era.
Open Data Integration
Renée J. Miller. 2018 · 2018
Cited alongside, same era.
Making Open Data Transparent: Data Discovery on Open Data
Renée J. Miller, Fatemeh Nargesian, Erkang Zhu, Christina Christodoulakis, Ken Q. Pu, and Periklis Andritsos. 2018 · 2018
Cited alongside, same era.
Table Union Search on Open Data
Fatemeh Nargesian, Erkang Zhu, Ken Q. Pu, and Renée J. Miller. 2018 · 2018
Cited alongside, same era.
Google Dataset Search: Building a search engine for datasets in an open Web ecosystem. In WWW . 1365–1375
Dan Brickley, Matthew Burgess, and Natasha F. Noy. 2019 · 2019
Efficient and Robust Approximate Nearest Neighbor Search Using Hierarchical Navigable Small World Graphs
Yury A. Malkov and Dmitry A. Yashunin. 2020 · 2020
Later among the works it cites.
Data-Driven Domain Discovery for Structured Datasets
Masayo Ota, Heiko Mueller, Juliana Freire, and Divesh Srivastava. 2020 · 2020
Later among the works it cites.
Transformers: State-of-the-Art Natural Language Processing. In EMNLP . 38–45
Thomas Wolf, Lysandre Debut, Victor Sanh, and et al. 2020 · 2020
Later among the works it cites.
TaBERT: Pretraining for Joint Understanding of Textual and Tabular Data. In ACL . 8413–8426
Pengcheng Yin, Graham Neubig, Wen-tau Yih, and Sebastian Riedel. 2020 · 2020
Later among the works it cites.
Sato: Contextual Semantic Type Detection in Tables
Dan Zhang, Yoshihiko Suhara, Jinfeng Li, Madelon Hulsebos, Çagatay Demiralp, and Wang-Chiew Tan. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In NAACL-HLT . 4171–4186
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Sherlock: A Deep Learning Approach to Semantic Data Type Detection. In KDD . 1500–1508
Madelon Hulsebos, Kevin Zeng Hu, Michiel A. Bakker, Emanuel Zgraggen, Arvind Satyanarayan, Tim Kraska, Çagatay Demiralp, and César A. Hidalgo. 2019 · 2019
Cited alongside, same era.
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 2019
Cited alongside, same era.
Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. In EMNLP . Association for Computational Linguistics, 3980–3990
Nils Reimers and Iryna Gurevych. 2019 · 2019
Cited alongside, same era.
MF-Join: Efficient Fuzzy String Similarity Join with Multi-level Filtering. In ICDE . 386–397
Jin Wang, Chunbin Lin, and Carlo Zaniolo. 2019 · 2019
Cited alongside, same era.
Scalable Metric Similarity Join Using MapReduce. In ICDE . 1662–1665
Jiacheng Wu, Yong Zhang, Jin Wang, Chunbin Lin, Yingjia Fu, and Chunxiao Xing. 2019 · 2019
Cited alongside, same era.
Sonia Castelo, Rémi Rampin, Aécio S. R. Santos, Aline Bessa, Fernando Chirigati, and Juliana Freire. 2021 · 2021
Later among the works it cites.
Efficient Joinable Table Discovery in Data Lakes: A High-Dimensional Similarity-Based Approach. In ICDE . 456–467
Yuyang Dong, Kunihiro Takeoka, Chuan Xiao, and Masafumi Oyamada. 2021 · 2021
Later among the works it cites.
Relational Header Discovery using Similarity Search in a Table Corpus. In ICDE . 444–455
Hazar Harmouch, Thorsten Papenbrock, and Felix Naumann. 2021 · 2021
Later among the works it cites.
TABBIE: Pretrained Representations of Tabular Data. In NAACL-HLT . 3446–3456
Hiroshi Iida, Dung Thai, Varun Manjunatha, and Mohit Iyyer. 2021 · 2021
Later among the works it cites.
Valentine: Evaluating Matching Techniques for Dataset Discovery. In ICDE . 468–479
Christos Koutras, George Siachamis, Andra Ionescu, Kyriakos Psarakis, Jerry Brons, Marios Fragkoulis, Christoph Lofi, Angela Bonifati, and Asterios Katsifodimos. 2021 · 2021
Later among the works it cites.
DomainNet: Homograph Detection for Data Lake Disambiguation. In EDBT . 13–24
Aristotelis Leventidis, Laura Di Rocco, Wolfgang Gatterbauer, Renée J. Miller, and Mirek Riedewald. 2021 · 2021
Later among the works it cites.
Deep Entity Matching: Challenges and Opportunities
Yuliang Li, Jinfeng Li, Yoshihiko Suhara, Jin Wang, Wataru Hirota, and Wang-Chiew Tan. 2021 · 2021
Later among the works it cites.
TCN: Table Convolutional Network for Web Table Interpretation. In WWW . 4020–4032
Daheng Wang, Prashant Shiralkar, Colin Lockard, Binxuan Huang, Xin Luna Dong, and Meng Jiang. 2021 · 2021
Later among the works it cites.
Grace Fan, Jin Wang, Yuliang Li, Dan Zhang, and Renée J. Miller. 2022 · 2022
Closest in time.
A Sketch-based Index for Correlated Dataset Search. In ICDE . 2928–2941
Aécio S. R. Santos, Aline Bessa, Christopher Musco, and Juliana Freire. 2022 · 2022
Closest in time.
Annotating Columns with Pre-trained Language Models. In SIGMOD . 1493–1503
Yoshihiko Suhara, Jinfeng Li, Yuliang Li, Dan Zhang, Çagatay Demiralp, Chen Chen, and Wang-Chiew Tan. 2022 · 2022
Closest in time.
Leva: Boosting Machine Learning Performance with Relational Embedding Data Augmentation. In SIGMOD . 1504–1517
Zixuan Zhao and Raul Castro Fernandez. 2022 · 2022
Closest in time.
SANTOS: Relationship-based Semantic Table Union Search. In SIGMOD
Aamod Khatiwada, Grace Fan, Roee Shraga, Zixuan Chen, Wolfgang Gatterbauer, Renée J. Miller, and Mirek Riedewald. 2023 · 2023
Closest in time.
CLAMS: Bringing Quality to Data Lakes. In SIGMOD . 2089–2092
Mina H. Farid, Alexandra Roatis, Ihab F. Ilyas, Hella-Franziska Hoffmann, and Xu Chu. 2016 · 2092
Closest in time.