Fetching the paper…
Reading the bibliography…
Data cleaning is often an important step to ensure that predictive models, such as regression and classification, are not affected by systematic errors such as inconsistent, out-of-date, or outlier data.
The interpretation of interaction in contingency tables
E. H. Simpson · 1951
Earlier work this paper cites.
The elements of statistical learning
J. Friedman, T. Hastie, and R. Tibshirani · 2001
Earlier work this paper cites.
Model-driven data acquisition in sensor networks
A. Deshpande, C. Guestrin, S. Madden, J. M. Hellerstein, and W. Hong · 2004
Earlier work this paper cites.
Declarative support for sensor data cleaning
S. R. Jeffery, G. Alonso, M. J. Franklin, W. Hong, and J. Widom · 2006
Earlier work this paper cites.
Active learning as non-convex optimization
A. Guillory, E. Chastain, and J. Bilmes · 2009
Earlier work this paper cites.
ERACER: a database approach for statistical inference and data cleaning
C. Mayfield, J. Neville, and S. Prabhakar · 2010
Earlier work this paper cites.
A survey on transfer learning
S. J. Pan and Q. Yang · 2010
Earlier work this paper cites.
Active learning literature survey
B. Settles · 2010
Earlier work this paper cites.
Guided data repair
M. Yakout, A. K. Elmagarmid, J. Neville, M. Ouzzani, and I. F. Ilyas · 2011
Earlier work this paper cites.
Stochastic gradient descent tricks
L. Bottou · 2012
Earlier work this paper cites.
Optimal distributed online prediction using mini-batches
O. Dekel, R. Gilad-Bachrach, O. Shamir, and L. Xiao · 2012
Earlier work this paper cites.
Fast approximation of matrix coherence and statistical leverage
P. Drineas, M. Magdon-Ismail, M. W. Mahoney, and D. P. Woodruff · 2012
Earlier work this paper cites.
Enterprise data analysis and visualization: An interview study
S. Kandel, A. Paepcke, J. M. Hellerstein, and J. Heer · 2012
Cited alongside, same era.
Query strategies for evading convex-inducing classifiers
B. Nelson, B. I. P. Rubinstein, L. Huang, A. D. Joseph, S. J. Lee, S. Rao, and J. D. Tygar · 2012
Cited alongside, same era.
Don’t be scared: use scalable automatic repairing with maximal likelihood and bounded changes
M. Yakout, L. Berti-Equille, and A. K. Elmagarmid · 2013
Cited alongside, same era.
http://www.nytimes.com/2014/08/18/technology/for-big-data-scientists-hurdle-to-insights-is-janitor-work.html
For big-data scientists, ’janitor work’ is key hurdle to insights · 2014
Cited alongside, same era.
The stratosphere platform for big data analytics
A. Alexandrov, R. Bergmann, S. Ewen, J. Freytag, F. Hueske, A. Heise, O. Kao, M. Leich, U. Leser, V. Markl, F. Naumann, M. Peters, A. Rheinländer, M. J. Sax, S. Schelter, M. Höger, K. Tzoumas, and D. Warneke · 2014
Cited alongside, same era.
A methodology for learning, analyzing, and mitigating social influence bias in recommender systems
S. Krishnan, J. Patel, M. J. Franklin, and K. Goldberg · 2014
Later among the works it cites.
Learning accurate kinematic control of cable-driven surgical robots using data cleaning and gaussian process regression
J. Mahler, S. Krishnan, M. Laskey, S. Sen, A. Murali, B. Kehoe, S. Patil, J. Wang, M. Franklin, P. Abbeel, and K. Y. Goldberg · 2014
Later among the works it cites.
Scaling up crowd-sourcing to very large datasets: A case for active learning
B. Mozafari, P. Sarkar, M. J. Franklin, M. I. Jordan, and S. Madden · 2014
Later among the works it cites.
A sample-and-clean framework for fast and accurate query processing on dirty data
J. Wang, S. Krishnan, M. J. Franklin, K. Goldberg, T. Kraska, and T. Milo · 2014
Later among the works it cites.
Incremental entity resolution on rules and data
S. E. Whang and H. Garcia-Molina · 2014
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Progressive approach to relational entity resolution
Y. Altowim, D. V. Kalashnikov, and S. Mehrotra · 2014
Cited alongside, same era.
Tupleware: Distributed machine learning on small clusters
A. Crotty, A. Galakatos, and T. Kraska · 2014
Cited alongside, same era.
Robust logistic regression and classification
J. Feng, H. Xu, S. Mannor, and S. Yan · 2014
Cited alongside, same era.
Corleone: Hands-off crowdsourcing for entity matching
C. Gokhale, S. Das, A. Doan, J. F. Naughton, N. Rampalli, J. Shavlik, and X. Zhu · 2014
Cited alongside, same era.
Incremental record linkage
A. Gruenheid, X. L. Dong, and D. Srivastava · 2014
Cited alongside, same era.
Estimating the number and sizes of fuzzy-duplicate clusters
A. Heise, G. Kasneci, and F. Naumann · 2014
Cited alongside, same era.
https://amplab.cs.berkeley.edu/software/
Berkeley data analytics stack
Cited in the paper.
Query-oriented data cleaning with oracles
M. Bergman, T. Milo, S. Novgorodov, and W. C. Tan · 2015
Later among the works it cites.
Incremental gradient, subgradient, and proximal methods for convex optimization: A survey
D. P. Bertsekas · 2015
Later among the works it cites.
Stale view cleaning: Getting fresh answers from stale materialized views
S. Krishnan, J. Wang, M. J. Franklin, K. Goldberg, and T. Kraska · 2015
Later among the works it cites.
Progressive duplicate detection
T. Papenbrock, A. Heise, and F. Naumann · 2015
Later among the works it cites.
Is feature selection secure against training data poisoning?
H. Xiao, B. Biggio, G. Brown, G. Fumera, C. Eckert, and F. Roli · 2015
Later among the works it cites.
Stochastic optimization with importance sampling for regularized loss minimization
P. Zhao and T. Zhang · 2015
Later among the works it cites.