Fetching the paper…
Reading the bibliography…
Structured data, or data that adheres to a pre-defined schema, can suffer from fragmented context: information describing a single entity can be scattered across multiple datasets or tables tailored for specific business needs, with no explicit linking keys (e.g., primary key-foreign key relationships or heuristic functions).
A survey of information retrieval and filtering methods
Christos Faloutsos and Douglas W Oard. 1998 · 1998
Earlier work this paper cites.
A taxonomy of web search. In ACM Sigir forum , Vol. 36. ACM New York, NY, USA, 3–10
Andrei Broder. 2002 · 2002
Earlier work this paper cites.
Crossing the Structure Chasm
Alon Y Halevy, Oren Etzioni, AnHai Doan, Zachary G Ives, Jayant Madhavan, Luke K McDowell, and Igor Tatarinov. 2003 · 2003
Earlier work this paper cites.
Challenges in Enterprise Search.. In ADC , Vol. 4. Citeseer, 15–24
David Hawking. 2004 · 2004
Earlier work this paper cites.
Enterprise Search: Tough Stuff: Why is it that searching an intranet is so much harder than searching the Web?
Rajat Mukherjee and Jianchang Mao. 2004 · 2004
Earlier work this paper cites.
Data integration for the relational web
Michael J Cafarella, Alon Halevy, and Nodira Khoussainova. 2009 · 2009
Earlier work this paper cites.
The probabilistic relevance framework: BM25 and beyond
Stephen Robertson and Hugo Zaragoza. 2009 · 2009
Earlier work this paper cites.
Distance metric learning for large margin nearest neighbor classification
Kilian Q Weinberger and Lawrence K Saul. 2009 · 2009
Earlier work this paper cites.
Google fusion tables: data management, integration and collaboration in the cloud. In Proceedings of the 1st ACM symposium on Cloud computing . 175–180
Hector Gonzalez, Alon Halevy, Christian S Jensen, Anno Langen, Jayant Madhavan, Rebecca Shapley, and Warren Shen. 2010 · 2010
Earlier work this paper cites.
Frameworks for entity matching: A comparison
Hanna Köpcke and Erhard Rahm. 2010 · 2010
Earlier work this paper cites.
Fast-join: An efficient method for fuzzy token matching based string similarity join. In 2011 IEEE 27th International Conference on Data Engineering . IEEE, 458–469
Jiannan Wang, Guoliang Li, and Jianhua Fe. 2011 · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. 2012 · 2012
Earlier work this paper cites.
Mctest: A challenge dataset for the open-domain machine comprehension of text. In Proceedings of the 2013 conference on empirical methods in natural language processing . 193–203
Matthew Richardson, Christopher JC Burges, and Erin Renshaw. 2013 · 2013
Earlier work this paper cites.
Datahub: Collaborative data science & dataset version management at scale
Anant Bhardwaj, Souvik Bhattacherjee, Amit Chavan, Amol Deshpande, Aaron J Elmore, Samuel Madden, and Aditya G Parameswaran. 2014 · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba. 2014 · 2014
Earlier work this paper cites.
Constructing an interactive natural language interface for relational databases
Fei Li and HV Jagadish. 2014 · 2014
Earlier work this paper cites.
Entity linking with a knowledge base: Issues, techniques, and solutions
Wei Shen, Jianyong Wang, and Jiawei Han. 2014 · 2014
Earlier work this paper cites.
Recommender system application developments: a survey
Jie Lu, Dianshuang Wu, Mingsong Mao, Wei Wang, and Guangquan Zhang. 2015 · 2015
Earlier work this paper cites.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books. In Proceedings of the IEEE international conference on computer vision . 19–27
Yukun Zhu, Ryan Kiros, Rich Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler. 2015 · 2015
Earlier work this paper cites.
MS MARCO: A human generated machine reading comprehension dataset
Payal Bajaj, Daniel Campos, Nick Craswell, Li Deng, Jianfeng Gao, Xiaodong Liu, Rangan Majumder, Andrew McNamara, Bhaskar Mitra, Tri Nguyen, et al · 2016
Earlier work this paper cites.
Goods: Organizing google’s datasets. In Proceedings of the 2016 International Conference on Management of Data . 795–806
Alon Halevy, Flip Korn, Natalya F Noy, Christopher Olston, Neoklis Polyzotis, Sudip Roy, and Steven Euijong Whang. 2016 · 2016
Earlier work this paper cites.
Magellan: Toward building entity matching management systems
Pradap Konda, Sanjib Das, Paul Suganthan GC, AnHai Doan, Adel Ardalan, Jeffrey R Ballard, Han Li, Fatemah Panahi, Haojun Zhang, Jeff Naughton, et al · 2016
Earlier work this paper cites.
To join or not to join? Thinking twice about joins before feature selection. In Proceedings of the 2016 International Conference on Management of Data . 19–34
Arun Kumar, Jeffrey Naughton, Jignesh M Patel, and Xiaojin Zhu. 2016 · 2016
Earlier work this paper cites.
Squad: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Cited alongside, same era.
String similarity search and join: a survey
Minghe Yu, Guoliang Li, Dong Deng, and Jianhua Feng. 2016 · 2016
Cited alongside, same era.
A demo of the data civilizer system. In Proceedings of the 2017 ACM International Conference on Management of Data . 1639–1642
Raul Castro Fernandez, Dong Deng, Essam Mansour, Abdulhakim A Qahtan, Wenbo Tao, Ziawasch Abedjan, Ahmed Elmagarmid, Ihab F Ilyas, Samuel Madden, Mourad Ouzzani, et al · 2017
Cited alongside, same era.
Billion-scale similarity search with GPUs
Jeff Johnson, Matthijs Douze, and Hervé Jégou. 2017 · 2017
Cited alongside, same era.
Are Key-Foreign Key Joins Safe to Avoid when Learning High-Capacity Classifiers?
Vraj Shah, Arun Kumar, and Xiaojin Zhu. 2017 · 2017
ARDA: Automatic Relational Data Augmentation for Machine Learning
Nadiia Chepurko, Ryan Marcus, Emanuel Zgraggen, Raul Castro Fernandez, Tim Kraska, and David Karger. 2020 · 2020
Later among the works it cites.
Yingqi Qu Yuchen Ding, Jing Liu, Kai Liu, Ruiyang Ren, Xin Zhao, Daxiang Dong, Hua Wu, and Haifeng Wang. 2020 · 2020
Later among the works it cites.
Nearest neighbor machine translation
Urvashi Khandelwal, Angela Fan, Dan Jurafsky, Luke Zettlemoyer, and Mike Lewis. 2020 · 2020
Later among the works it cites.
Colbert: Efficient and effective passage search via contextualized late interaction over bert. In Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval . 39–48
Omar Khattab and Matei Zaharia. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Attention is All you Need. In NIPS
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
End-to-end neural ad-hoc ranking with kernel pooling. In Proceedings of the 40th International ACM SIGIR conference on research and development in information retrieval . 55–64
Chenyan Xiong, Zhuyun Dai, Jamie Callan, Zhiyuan Liu, and Russell Power. 2017 · 2017
Cited alongside, same era.
Datasets for DeepMatcher paper
2018 · 2018
Cited alongside, same era.
Distributed Representations of Tuples for Entity Resolution
Muhammad Ebraheem, Saravanan Thirumuruganathan, Shafiq Joty, Mourad Ouzzani, and Nan Tang. 2018 · 2018
Cited alongside, same era.
Aurum: A data discovery system. In 2018 IEEE 34th International Conference on Data Engineering (ICDE) . IEEE, 1001–1012
Raul Castro Fernandez, Ziawasch Abedjan, Famien Koko, Gina Yuan, Samuel Madden, and Michael Stonebraker. 2018 · 2018
Cited alongside, same era.
Deep learning for entity matching: A design space exploration. In Proceedings of the 2018 International Conference on Management of Data . 19–34
Sidharth Mudgal, Han Li, Theodoros Rekatsinas, AnHai Doan, Youngchoon Park, Ganesh Krishnan, Rohit Deep, Esteban Arcaute, and Vijay Raghavendra. 2018 · 2018
Cited alongside, same era.
Hotpotqa: A dataset for diverse, explainable multi-hop question answering
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W Cohen, Ruslan Salakhutdinov, and Christopher D Manning. 2018 · 2018
Cited alongside, same era.
Yuliang Li, Jinfeng Li, Yoshihiko Suhara, AnHai Doan, and Wang-Chiew Tan. 2020 · 2020
Later among the works it cites.
Bridging Textual and Tabular Data for Cross-Domain Text-to-SQL Semantic Parsing
Xi Victoria Lin, Richard Socher, and Caiming Xiong. 2020 · 2020
Later among the works it cites.
Picket: Self-supervised Data Diagnostics for ML Pipelines
Zifan Liu, Zhechun Zhou, and Theodoros Rekatsinas. 2020 · 2020
Later among the works it cites.
Bootleg: Chasing the Tail with Self-Supervised Named Entity Disambiguation
Laurel Orr, Megan Leszczynski, Simran Arora, Sen Wu, Neel Guha, Xiao Ling, and Christopher Re. 2020 · 2020
Later among the works it cites.
Blocking and filtering techniques for entity resolution: A survey
George Papadakis, Dimitrios Skoutas, Emmanouil Thanos, and Themis Palpanas. 2020 · 2020
Later among the works it cites.
Relational Pretrained Transformers towards Democratizing Data Preparation [Vision]
Nan Tang, Ju Fan, Fangyi Li, Jianhong Tu, Xiaoyong Du, Guoliang Li, Sam Madden, and Mourad Ouzzani. 2020 · 2020
Later among the works it cites.
James Thorne, Majid Yazdani, Marzieh Saeidi, Fabrizio Silvestri, Sebastian Riedel, and Alon Halevy. 2020 · 2020
Later among the works it cites.
TaBERT: Pretraining for Joint Understanding of Textual and Tabular Data. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics . 8413–8426
Pengcheng Yin, Graham Neubig, Wen-tau Yih, and Sebastian Riedel. 2020 · 2020
Later among the works it cites.
Photon: A Robust Cross-Domain Text-to-SQL System
Jichuan Zeng, Xi Victoria Lin, Caiming Xiong, Richard Socher, Michael R Lyu, Irwin King, and Steven CH Hoi. 2020 · 2020
Later among the works it cites.
Cloud AutoML
2021 · 2021
Closest in time.
Data Robot
2021 · 2021
Closest in time.
IMDb Datasets
2021 · 2021
Closest in time.
MS MARCO
2021 · 2021
Closest in time.
rank-bm25
2021 · 2021
Closest in time.
Wikimedia Downloads
2021 · 2021
Closest in time.
Auctus: A Dataset Search Engine for Data Augmentation
Fernando Chirigati, Rémi Rampin, Aécio Santos, Aline Bessa, and Juliana Freire. 2021 · 2021
Closest in time.
Auto-FuzzyJoin: Auto-Program Fuzzy Similarity Joins Without Labeled Examples
Peng Li, Xiang Cheng, Xu Chu, Yeye He, and Surajit Chaudhuri. 2021 · 2021
Closest in time.
Retrieving and Reading: A Comprehensive Survey on Open-domain Question Answering
Fengbin Zhu, Wenqiang Lei, Chao Wang, Jianming Zheng, Soujanya Poria, and Tat-Seng Chua. 2021 · 2021
Closest in time.