Fetching the paper…
Reading the bibliography…
Correctly detecting the semantic type of data columns is crucial for data science tasks such as automated data cleaning, schema matching, and data discovery.
A Survey of Approaches to Automatic Schema Matching
Erhard Rahm and Philip A. Bernstein. 2001 · 2001
Earlier work this paper cites.
Potter’s Wheel: An Interactive Data Cleaning System. In Proceedings of the 27th International Conference on Very Large Data Bases (VLDB ’01) . Morgan Kaufmann Publishers Inc., San Francisco, CA, USA, 381–390
Vijayshankar Raman and Joseph M. Hellerstein. 2001 · 2001
Earlier work this paper cites.
DBpedia: A nucleus for a web of open data
Sören Auer, Christian Bizer, Georgi Kobilarov, Jens Lehmann, Richard Cyganiak, and Zachary Ives. 2007 · 2007
Earlier work this paper cites.
Freebase: A Collaboratively Created Graph Database for Structuring Human Knowledge. In Proceedings of the 2008 ACM SIGMOD International Conference on Management of Data (SIGMOD ’08) . ACM, New York, NY, USA, 1247–1250
Kurt Bollacker, Colin Evans, Praveen Paritosh, Tim Sturge, and Jamie Taylor. 2008 · 2008
Earlier work this paper cites.
WebTables: Exploring the Power of Tables on the Web
Michael J. Cafarella, Alon Halevy, Daisy Zhe Wang, Eugene Wu, and Yang Zhang. 2008 · 2008
Earlier work this paper cites.
Crowdsourcing User Studies with Mechanical Turk. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems (CHI ’08) . ACM, New York, NY, USA, 453–456
Aniket Kittur, Ed H. Chi, and Bongwon Suh. 2008 · 2008
Earlier work this paper cites.
Annotating and searching web tables using entities, types and relationships
Girija Limaye, Sunita Sarawagi, and Soumen Chakrabarti. 2010 · 2010
Earlier work this paper cites.
Software Framework for Topic Modelling with Large Corpora. In Proceedings of the LREC 2010 Workshop on New Challenges for NLP Frameworks . ELRA, 45–50
Radim Řehůřek and Petr Sojka. 2010 · 2010
Earlier work this paper cites.
Exploiting a web of semantic data for interpreting tables. In Proceedings of the Second Web Science Conference
Zareen Syed, Tim Finin, Varish Mulwad, Anupam Joshi, et al · 2010
Earlier work this paper cites.
About WordNet
Princeton University. 2010 · 2010
Earlier work this paper cites.
Wrangler: Interactive Visual Specification of Data Transformation Scripts. In ACM Human Factors in Computing Systems (CHI)
Sean Kandel, Andreas Paepcke, Joseph Hellerstein, and Jeffrey Heer. 2011 · 2011
Earlier work this paper cites.
Scikit-learn: Machine Learning in Python
Fabian Pedregosa et al · 2011
Cited alongside, same era.
Recovering semantics of tables on the web
Petros Venetis, Alon Halevy, Jayant Madhavan, Marius Paşca, Warren Shen, Fei Wu, Gengxin Miao, and Chung Wu. 2011 · 2011
Cited alongside, same era.
Exploiting structure within data for accurate labeling using conditional random fields. In Proceedings on the International Conference on Artificial Intelligence (ICAI)
Aman Goel, Craig A Knoblock, and Kristina Lerman. 2012 · 2012
Cited alongside, same era.
A Specialist Approach for Classification of Column Data
Nikhil Waman Puranik. 2012 · 2012
Cited alongside, same era.
Utilizing Regular Expressions for Instance-Based Schema Matching
Benjamin Zapilko, Matthäus Zloch, and Johann Schaible. 2012 · 2012
Cited alongside, same era.
Distributed representations of sentences and documents. In International Conference on Machine Learning . 1188–1196
Inference of regular expressions for text extraction from examples
Alberto Bartoli, Andrea De Lorenzo, Eric Medvet, and Fabiano Tarlao. 2016 · 2016
Later among the works it cites.
Semantic labeling: a domain-independent approach. In International Semantic Web Conference . Springer, 446–462
Minh Pham, Suresh Alse, Craig A Knoblock, and Pedro Szekely. 2016 · 2016
Later among the works it cites.
Matching web tables to DBpedia – a feature utility study
Dominique Ritze and Christian Bizer. 2017 · 2017
Later among the works it cites.
Seeping Semantics: Linking Datasets Using Word Embeddings for Data Discovery
Raul Castro Fernandez, Essam Mansour, Abdulhakim Qahtan, Ahmed Elmagarmid, Ihab Ilyas, Samuel Madden, Mourad Ouzzani, Michael Stonebraker, and Nan Tang. 2018b · 2018
Later among the works it cites.
Synthesizing type-detection logic for rich semantic data types using open-source code. In Proceedings of the 2018 International Conference on Management of Data . ACM, 35–50
Cong Yan and Yeye He. 2018 · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Quoc Le and Tomas Mikolov. 2014 · 2014
Cited alongside, same era.
Glove: Global vectors for word representation. In Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP) . 1532–1543
Jeffrey Pennington, Richard Socher, and Christopher Manning. 2014 · 2014
Cited alongside, same era.
Document embedding with paragraph vectors
Andrew M Dai, Christopher Olah, and Quoc V Le. 2015 · 2015
Cited alongside, same era.
Short text similarity with word embeddings. In Proceedings of the 24th ACM international on conference on information and knowledge management . ACM, 1411–1420
Tom Kenter and Maarten De Rijke. 2015 · 2015
Cited alongside, same era.
Assigning semantic labels to data sources. In European Semantic Web Conference . Springer, 403–417
S Krishnamurthy Ramnandan, Amol Mittal, Craig A Knoblock, and Pedro Szekely. 2015 · 2015
Cited alongside, same era.
TensorFlow: A system for large-scale machine learning. In 12th USENIX Symposium on Operating Systems Design and Implementation (OSDI 16) . 265–283
Martín Abadi et al · 2016
Cited alongside, same era.
Aurum: A Data Discovery System. 1001–1012
Raul Castro Fernandez, Ziawasch Abedjan, Famien Koko, Gina Yuan, Samuel Madden, and Michael Stonebraker. 2018a
Cited in the paper.
Datalib: JavaScript Data Utilities
Interactive Data Lab. 2019 · 2019
Closest in time.
messytables ⋅ \cdot PyPi
Open Knowledge Foundation. 2019 · 2019
Closest in time.
Google Data Studio
Google. 2019 · 2019
Closest in time.
VizNet: Towards a large-scale visualization learning and benchmarking repository. In Proceedings of the 2019 Conference on Human Factors in Computing Systems (CHI) . ACM
Kevin Hu, Neil Gaikwad, Michiel Bakker, Madelon Hulsebos, Emanuel Zgraggen, César Hidalgo, Tim Kraska, Guoliang Li, Arvind Satyanarayan, and Çağatay Demiralp. 2019 · 2019
Closest in time.
Power BI | Interactive Data Visualization BI
Microsoft. 2019 · 2019
Closest in time.
Data Wrangling Tools & Software
Trifacta. 2019 · 2019
Closest in time.