Fetching the paper…
Reading the bibliography…
Can AI help automate human-easy but computer-hard data preparation tasks that burden data scientists, practitioners, and crowd workers? We answer this question by presenting RPT, a denoising auto-encoder for tuple-to-X models (X could be tuple, token, label, JSON, and so on).
Baran: Effective Error Correction via a Unified Context Representation and Transfer Learning
Mohammad Mahdavi and Ziawasch Abedjan. 2020 · 1961
Earlier work this paper cites.
Learning to Learn: Introduction and Overview
Sebastian Thrun and Lorien Y. Pratt. 1998 · 1998
Earlier work this paper cites.
Retrospective Reader for Machine Reading Comprehension
Zhuosheng Zhang, Junjie Yang, and Hai Zhao. 2020 · 2001
Earlier work this paper cites.
A Cost-Based Model and Effective Heuristic for Repairing Constraints by Value Modification. In SIGMOD , Fatma Özcan (Ed.). 143–154
Philip Bohannon, Michael Flaster, Wenfei Fan, and Rajeev Rastogi. 2005 · 2005
Earlier work this paper cites.
Duplicate Record Detection: A Survey
Ahmed K. Elmagarmid, Panagiotis G. Ipeirotis, and Vassilios S. Verykios. 2007 · 2007
Earlier work this paper cites.
Web-scale extraction of structured data
Michael J. Cafarella, Jayant Madhavan, and Alon Y. Halevy. 2008 · 2008
Earlier work this paper cites.
Conditional functional dependencies for capturing data inconsistencies
Wenfei Fan, Floris Geerts, Xibei Jia, and Anastasios Kementsietsidis. 2008 · 2008
Earlier work this paper cites.
Towards Certain Fixes with Editing Rules and Master Data
Wenfei Fan, Jianzhong Li, Shuai Ma, Nan Tang, and Wenyuan Yu. 2010 · 2010
Earlier work this paper cites.
Automatic Rule Refinement for Information Extraction
Bin Liu, Laura Chiticariu, Vivian Chu, H. V. Jagadish, and Frederick Reiss. 2010 · 2010
Earlier work this paper cites.
Guided data repair
Mohamed Yakout, Ahmed K. Elmagarmid, Jennifer Neville, Mourad Ouzzani, and Ihab F. Ilyas. 2011 · 2011
Earlier work this paper cites.
Principles of Data Integration
AnHai Doan, Alon Y. Halevy, and Zachary G. Ives. 2012 · 2012
Earlier work this paper cites.
Foundations of Data Quality Management
Wenfei Fan and Floris Geerts. 2012 · 2012
Earlier work this paper cites.
Active Sampling for Entity Matching with Guarantees
Kedar Bellare, Suresh Iyengar, Aditya G. Parameswaran, and Vibhor Rastogi. 2013 · 2013
Earlier work this paper cites.
Discovering Denial Constraints
Xu Chu, Ihab F. Ilyas, and Paolo Papotti. 2013 · 2013
Earlier work this paper cites.
Inferring data currency and consistency for conflict resolution. In ICDE , Christian S. Jensen, Christopher M. Jermaine, and Xiaofang Zhou (Eds.). 470–481
Wenfei Fan, Floris Geerts, Nan Tang, and Wenyuan Yu. 2013 · 2013
Earlier work this paper cites.
Don’t be SCAREd: use SCalable Automatic REpairing with maximal likelihood and bounded changes. In SIGMOD , Kenneth A. Ross, Divesh Srivastava, and Dimitris Papadias (Eds.). ACM, 553–564
Mohamed Yakout, Laure Berti-Équille, and Ahmed K. Elmagarmid. 2013 · 2013
Earlier work this paper cites.
Towards dependable data repairing with fixing rules. In SIGMOD . ACM, 457–468
Jiannan Wang and Nan Tang. 2014 · 2014
Earlier work this paper cites.
KATARA: A Data Cleaning System Powered by Knowledge Bases and Crowdsourcing. In SIGMOD . 1247–1261
Xu Chu, John Morcos, Ihab F. Ilyas, Mourad Ouzzani, Paolo Papotti, Nan Tang, and Yin Ye. 2015 · 2015
Earlier work this paper cites.
Proof positive and negative in data cleaning. In ICDE . 18–29
Matteo Interlandi and Nan Tang. 2015 · 2015
Earlier work this paper cites.
Combining Quantitative and Logical Data Cleaning
Nataliya Prokoshyna, Jaroslaw Szlichta, Fei Chiang, Renée J. Miller, and Divesh Srivastava. 2015 · 2015
Earlier work this paper cites.
Detecting Data Errors: Where are we and what needs to be done?
Ziawasch Abedjan, Xu Chu, Dong Deng, Raul Castro Fernandez, Ihab F. Ilyas, Mourad Ouzzani, Paolo Papotti, Michael Stonebraker, and Nan Tang. 2016 · 2016
Earlier work this paper cites.
Interactive and Deterministic Data Cleaning. In SIGMOD . 893–907
Jian He, Enzo Veltri, Donatello Santoro, Guoliang Li, Giansalvatore Mecca, Paolo Papotti, and Nan Tang. 2016 · 2016
Earlier work this paper cites.
Magellan: Toward Building Entity Matching Management Systems
Pradap Konda, Sanjib Das, Paul Suganthan G. C., AnHai Doan, Adel Ardalan, Jeffrey R. Ballard, Han Li, Fatemah Panahi, Haojun Zhang, Jeffrey F. Naughton, Shishir Prasad, Ganesh Krishnan, Rohit Deep, and Vijay Raghavendra. 2016 · 2016
Cited alongside, same era.
SQuAD: 100, 000+ Questions for Machine Comprehension of Text. In EMNLP , Jian Su, Xavier Carreras, and Kevin Duh (Eds.). 2383–2392
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Cited alongside, same era.
"Why Should I Trust You?": Explaining the Predictions of Any Classifier. In SIGKDD . 1135–1144
Marco Túlio Ribeiro, Sameer Singh, and Carlos Guestrin. 2016 · 2016
Cited alongside, same era.
The Data Civilizer System. In CIDR
Dong Deng, Raul Castro Fernandez, Ziawasch Abedjan, Sibo Wang, Michael Stonebraker, Ahmed K. Elmagarmid, Ihab F. Ilyas, Samuel Madden, Mourad Ouzzani, and Nan Tang. 2017 · 2017
Cited alongside, same era.
A Roadmap for a Rigorous Science of Interpretability
Finale Doshi-Velez and Been Kim. 2017 · 2017
Raha: A Configuration-Free Error Detection System. In SIGMOD . 865–882
Mohammad Mahdavi, Ziawasch Abedjan, Raul Castro Fernandez, Samuel Madden, Mourad Ouzzani, Michael Stonebraker, and Nan Tang. 2019 · 2019
Later among the works it cites.
Federated Machine Learning: Concept and Applications
Qiang Yang, Yang Liu, Tianjian Chen, and Yongxin Tong. 2019 · 2019
Later among the works it cites.
Language Models are Few-Shot Learners. In NeurIPS
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 2020
Closest in time.
Entity Matching with Transformer Architectures - A Step Forward in Data Integration. In EDBT . 463–473
Ursin Brunner and Kurt Stockinger. 2020 · 2020
Closest in time.
Creating Embeddings of Heterogeneous Relational Datasets for Data Integration Tasks. In SIGMOD . 1335–1349
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Data Integration: After the Teenage Years. In PODS , Emanuel Sallinger, Jan Van den Bussche, and Floris Geerts (Eds.). ACM, 101–106
Behzad Golshan, Alon Y. Halevy, George A. Mihaila, and Wang-Chiew Tan. 2017 · 2017
Cited alongside, same era.
Cleaning Relations Using Knowledge Bases. In ICDE . 933–944
Shuang Hao, Nan Tang, Guoliang Li, and Jian Li. 2017 · 2017
Cited alongside, same era.
Understanding Workers, Developing Effective Tasks, and Enhancing Marketplace Dynamics: A Study of a Large Crowdsourcing Marketplace
Ayush Jain, Akash Das Sarma, Aditya G. Parameswaran, and Jennifer Widom. 2017 · 2017
Cited alongside, same era.
Interpretable Machine Learning: The fuss, the concrete and the questions. In ICML Tutorial
Been Kim and Finale Doshi-Velez. 2017 · 2017
Cited alongside, same era.
HoloClean: Holistic Data Repairs with Probabilistic Inference
Theodoros Rekatsinas, Xu Chu, Ihab F. Ilyas, and Christopher Ré. 2017 · 2017
Cited alongside, same era.
Attention is All you Need. In NIPS . 5998–6008
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Distributed Representations of Tuples for Entity Resolution
Muhammad Ebraheem, Saravanan Thirumuruganathan, Shafiq R. Joty, Mourad Ouzzani, and Nan Tang. 2018 · 2018
Cited alongside, same era.
Riccardo Cappuzzo, Paolo Papotti, and Saravanan Thirumuruganathan. 2020 · 2020
Closest in time.
A Simple Framework for Contrastive Learning of Visual Representations. In ICML
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey E. Hinton. 2020 · 2020
Closest in time.
TURL: Table Understanding through Representation Learning
Xiang Deng, Huan Sun, Alyssa Lees, You Wu, and Cong Yu. 2020 · 2020
Closest in time.
Data Preparation: A Survey of Commercial Tools
Mazhar Hameed and Felix Naumann. 2020 · 2020
Closest in time.
Record fusion: A learning approach
Alireza Heidari, George Michalopoulos, Shrinu Kushagra, Ihab F. Ilyas, and Theodoros Rekatsinas. 2020 · 2020
Closest in time.
TaPas: Weakly Supervised Table Parsing via Pre-training. In ACL , Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel R. Tetreault (Eds.). 4320–4333
Jonathan Herzig, Pawel Krzysztof Nowak, Thomas Müller, Francesco Piccinno, and Julian Martin Eisenschlos. 2020 · 2020
Closest in time.
SpanBERT: Improving Pre-training by Representing and Predicting Spans
Mandar Joshi, Danqi Chen, Yinhan Liu, Daniel S. Weld, Luke Zettlemoyer, and Omer Levy. 2020 · 2020
Closest in time.
BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension. In ACL . 7871–7880
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020 · 2020
Closest in time.
A Secure Federated Transfer Learning Framework
Yang Liu, Yan Kang, Chaoping Xing, Tianjian Chen, and Qiang Yang. 2020 · 2020
Closest in time.
Blocking and Filtering Techniques for Entity Resolution: A Survey
George Papadakis, Dimitrios Skoutas, Emmanouil Thanos, and Themis Palpanas. 2020 · 2020
Closest in time.
Pattern Functional Dependencies for Data Cleaning
Abdulhakim Ali Qahtan, Nan Tang, Mourad Ouzzani, Yang Cao, and Michael Stonebraker. 2020 · 2020
Closest in time.
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020 · 2020
Closest in time.
Exploiting Cloze Questions for Few-Shot Text Classification and Natural Language Inference
Timo Schick and Hinrich Schütze. 2020 · 2020
Closest in time.
sigmod-2020-contest. [n.d.]
2020
Closest in time.
CorDEL: A Contrastive Deep Learning Approach for Entity Linkage
Zhengyang Wang, Bunyamin Sisman, Hao Wei, Xin Luna Dong, and Shuiwang Ji. 2020 · 2020
Closest in time.
ZeroER: Entity Resolution using Zero Labeled Examples. In SIGMOD . 1149–1164
Renzhi Wu, Sanya Chaba, Saurabh Sawlani, Xu Chu, and Saravanan Thirumuruganathan. 2020 · 2020
Closest in time.
TaBERT: Pretraining for Joint Understanding of Textual and Tabular Data. In ACL . 8413–8426
Pengcheng Yin, Graham Neubig, Wen-tau Yih, and Sebastian Riedel. 2020 · 2020
Closest in time.
Siamese Neural Networks: An Overview
Davide Chicco. 2021 · 2021
Closest in time.