Fetching the paper…
Reading the bibliography…
The analyst effort in data cleaning is gradually shifting away from the design of hand-written scripts to building and tuning complex pipelines of automated data cleaning libraries.
A relational model of data for large shared data banks
E. F. Codd · 1970
Earlier work this paper cites.
The theory of joins in relational databases
A. V. Aho, C. Beeri, and J. D. Ullman · 1979
Earlier work this paper cites.
Data integration using self-maintainable views
A. Gupta, H. V. Jagadish, and I. S. Mumick · 1996
Earlier work this paper cites.
Min-wise independent permutations
A. Z. Broder, M. Charikar, A. M. Frieze, and M. Mitzenmacher · 2000
Earlier work this paper cites.
Data cleaning: Problems and current approaches
E. Rahm and H. H. Do · 2000
Earlier work this paper cites.
Declarative data cleaning: Language, model, and algorithms
H. Galhardas, D. Florescu, D. Shasha, E. Simon, and C. Saita · 2001
Earlier work this paper cites.
Potter’s wheel: An interactive data cleaning system
V. Raman and J. M. Hellerstein · 2001
Earlier work this paper cites.
The chase revisited
A. Deutsch, A. Nash, and J. B. Remmel · 2008
Earlier work this paper cites.
Database Repairing and Consistent Query Answering
L. E. Bertossi · 2011
Earlier work this paper cites.
Wrangler: interactive visual specification of data transformation scripts
S. Kandel, A. Paepcke, J. Hellerstein, and J. Heer · 2011
Earlier work this paper cites.
Guided data repair
M. Yakout, A. K. Elmagarmid, J. Neville, M. Ouzzani, and I. F. Ilyas · 2011
Earlier work this paper cites.
Hyperopt: A python library for optimizing the hyperparameters of machine learning algorithms
J. Bergstra, D. Yamins, and D. D. Cox · 2013
Earlier work this paper cites.
Scorpion: Explaining away outliers in aggregate queries
E. Wu and S. Madden · 2013
Earlier work this paper cites.
Don’t be scared: use scalable automatic repairing with maximal likelihood and bounded changes
M. Yakout, L. Berti-Equille, and A. K. Elmagarmid · 2013
Earlier work this paper cites.
http://www.nytimes.com/2014/08/18/technology/for-big-data-scientists-hurdle-to-insights-is-janitor-work.html
For big-data scientists, ’janitor work’ is key hurdle to insights · 2014
Earlier work this paper cites.
Progressive approach to relational entity resolution
Y. Altowim, D. V. Kalashnikov, and S. Mehrotra · 2014
Earlier work this paper cites.
A data quality metric (dqm): How to estimate the number of undetected errors in data sets
Y. Chung, S. Krishnan, and T. Kraska · 2014
Cited alongside, same era.
Incremental detection of inconsistencies in distributed data
W. Fan, J. Li, N. Tang, et al · 2014
Cited alongside, same era.
Corleone: Hands-off crowdsourcing for entity matching
C. Gokhale, S. Das, A. Doan, J. F. Naughton, N. Rampalli, J. Shavlik, and X. Zhu · 2014
Cited alongside, same era.
Scaling up crowd-sourcing to very large datasets: A case for active learning
B. Mozafari, P. Sarkar, M. J. Franklin, M. I. Jordan, and S. Madden · 2014
Cited alongside, same era.
A sample-and-clean framework for fast and accurate query processing on dirty data
J. Wang, S. Krishnan, M. J. Franklin, K. Goldberg, T. Kraska, and T. Milo · 2014
Cited alongside, same era.
Query: a framework for integrating entity resolution with query processing
Artificial intelligence: a modern approach
S. J. Russell and P. Norvig · 2016
Later among the works it cites.
Taking the human out of the loop: A review of bayesian optimization
B. Shahriari, K. Swersky, Z. Wang, R. P. Adams, and N. De Freitas · 2016
Later among the works it cites.
Macrobase: Prioritizing attention in fast data
P. Bailis, E. Gan, S. Madden, D. Narayanan, K. Rong, and S. Suri · 2017
Later among the works it cites.
Tfx: A tensorflow-based production-scale machine learning platform
D. Baylor, E. Breck, H.-T. Cheng, N. Fiedel, C. Y. Foo, Z. Haque, S. Haykal, M. Ispir, V. Jain, L. Koc, et al · 2017
Later among the works it cites.
Google vizier: A service for black-box optimization
D. Golovin, B. Solnik, S. Moitra, G. Kochanski, J. Karro, and D. Sculley · 2017
Later among the works it cites.
Foofah: Transforming data by example
Z. Jin, M. R. Anderson, M. Cafarella, and H. Jagadish · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
H. Altwaijry, S. Mehrotra, and D. V. Kalashnikov · 2015
Cited alongside, same era.
Query-oriented data cleaning with oracles
M. Bergman, T. Milo, S. Novgorodov, and W. C. Tan · 2015
Cited alongside, same era.
Wisteria: Nurturing scalable data cleaning infrastructure
D. Haas, S. Krishnan, J. Wang, M. J. Franklin, and E. Wu · 2015
Cited alongside, same era.
Trends in cleaning relational data: Consistency and deduplication
I. F. Ilyas, X. Chu, et al · 2015
Cited alongside, same era.
Bigdansing: A system for big data cleansing
Z. Khayyat, I. F. Ilyas, A. Jindal, S. Madden, M. Ouzzani, P. Papotti, J.-A. Quiané-Ruiz, N. Tang, and S. Yin · 2015
Cited alongside, same era.
Sampleclean: Fast and reliable analytics on dirty data
S. Krishnan, J. Wang, M. J. Franklin, K. Goldberg, T. Kraska, T. Milo, and E. Wu · 2015
Cited alongside, same era.
Data cleaning: Overview and emerging challenges
X. Chu, I. F. Ilyas, S. Krishnan, and J. Wang · 2016
Cited alongside, same era.
Later among the works it cites.
Hyperband: A novel bandit-based approach to hyperparameter optimization
L. Li, K. Jamieson, G. DeSalvo, A. Rostamizadeh, and A. Talwalkar · 2017
Later among the works it cites.
Holoclean: Holistic data repairs with probabilistic inference
T. Rekatsinas, X. Chu, I. F. Ilyas, and C. Ré · 2017
Later among the works it cites.
Keystoneml: Optimizing pipelines for large-scale advanced analytics
E. R. Sparks, S. Venkataraman, T. Kaftan, M. J. Franklin, and B. Recht · 2017
Later among the works it cites.
Combining design and performance in a data visualization management system
E. Wu, F. Psallidas, Z. Miao, H. Zhang, and L. Rettig · 2017
Later among the works it cites.
Toward a system building agenda for data integration (and data science)
A. Doan, P. Konda, A. Ardalan, J. R. Ballard, S. Das, Y. Govind, H. Li, P. Martinkus, S. Mudgal, E. Paulson, et al · 2018
Later among the works it cites.
Data cleaning is a machine learning problem
I. Ilyas · 2018
Later among the works it cites.
Tune: A research platform for distributed model selection and training
R. Liaw, E. Liang, R. Nishihara, P. Moritz, J. E. Gonzalez, and I. Stoica · 2018
Later among the works it cites.
Deep learning for entity matching: A design space exploration
S. Mudgal, H. Li, T. Rekatsinas, A. Doan, Y. Park, G. Krishnan, R. Deep, E. Arcaute, and V. Raghavendra · 2018
Later among the works it cites.
Pyod: A python toolbox for scalable outlier detection
Y. Zhao, Z. Nasrullah, and Z. Li · 2019
Closest in time.