Fetching the paper…
Reading the bibliography…
Software organizations are increasingly incorporating machine learning (ML) into their product offerings, driving a need for new data management tools.
F. J. Massey Jr, “The kolmogorov-smirnov test for goodness of fit,” Journal of the American statistical Association , vol. 46, no. 253, pp. 68–78, 1951
1951
Earlier work this paper cites.
S. Acharya, P. B. Gibbons, V. Poosala, and S. Ramaswamy, “The aqua approximate query answering system,” in Proceedings of the 1999 ACM SIGMOD international conference on Management of data , 1999, pp. 574–576
1999
Earlier work this paper cites.
S. Chaudhuri, R. Motwani, and V. Narasayya, “On random sampling over joins,” SIGMOD Rec. , vol. 28, no. 2, p. 263–274, jun 1999. [Online]. Available: https://doi.org/10.1145/304181.304206
1999
Earlier work this paper cites.
B. Babcock, M. Datar, and R. Motwani, “Sampling from a moving window over streaming data,” in 2002 Annual ACM-SIAM Symposium on Discrete Algorithms (SODA 2002) . Stanford InfoLab, 2001
2001
Earlier work this paper cites.
M. Datar, A. Gionis, P. Indyk, and R. Motwani, “Maintaining stream statistics over sliding windows,” SIAM journal on computing , vol. 31, no. 6, pp. 1794–1813, 2002
2002
Earlier work this paper cites.
P. B. Gibbons, Y. Matias, and V. Poosala, “Fast incremental maintenance of approximate histograms,” ACM Trans. Database Syst. , vol. 27, no. 3, p. 261–298, sep 2002. [Online]. Available: https://doi.org/10.1145/581751.581753
2002
Earlier work this paper cites.
C. X. Ling, J. Huang, H. Zhang et al. , “Auc: a statistically consistent and more discriminating measure than accuracy,” in Ijcai , vol. 3, 2003, pp. 519–524
2003
Earlier work this paper cites.
J. H. Chang and W. S. Lee, “Finding recent frequent itemsets adaptively over online data streams,” in Proceedings of the ninth ACM SIGKDD international conference on Knowledge discovery and data mining , 2003, pp. 487–492
2003
Earlier work this paper cites.
T. Oinn, M. Addis, J. Ferris, D. Marvin, M. Senger, M. Greenwood, T. Carver, K. Glover, M. R. Pocock, A. Wipat et al. , “Taverna: a tool for the composition and enactment of bioinformatics workflows,” Bioinformatics , vol. 20, no. 17, pp. 3045–3054, 2004
2004
Earlier work this paper cites.
R. Jin and G. Agrawal, “An algorithm for in-core frequent itemset mining on streaming data,” in Fifth IEEE International Conference on Data Mining (ICDM’05) , 2005, pp. 8 pp.–
2005
Earlier work this paper cites.
C. C. Aggarwal, “On biased reservoir sampling in the presence of stream evolution,” in Proceedings of the 32nd international conference on Very large data bases . Citeseer, 2006, pp. 607–618
2006
Earlier work this paper cites.
D. François, V. Wertz, and M. Verleysen, “The permutation test for feature selection by mutual information,” in ESANN , 2006
2006
Earlier work this paper cites.
M. Sugiyama et al. , “Covariate shift adaptation by importance weighted cross validation,” in JMLR , 2007
2007
Earlier work this paper cites.
R. Gemulla and W. Lehner, “Sampling time-based sliding windows in bounded space,” in Proceedings of the 2008 ACM SIGMOD international conference on Management of data , 2008, pp. 379–392
2008
Earlier work this paper cites.
A. Thusoo et al. , “Hive: A warehousing solution over a map-reduce framework,” in VLDB , 2009
2009
Earlier work this paper cites.
J. Cheney, L. Chiticariu, W.-C. Tan et al. , “Provenance in databases: Why, how, and where,” Foundations and Trends® in Databases , vol. 1, no. 4, pp. 379–474, 2009
2009
Earlier work this paper cites.
M. A. B. Tobji, B. B. Yaghlane, and K. Mellouli, “Incremental maintenance of frequent itemsets in evidential databases,” in ECSQARU , 2009
2009
Earlier work this paper cites.
V. Chandola, A. Banerjee, and V. Kumar, “Anomaly detection: A survey,” ACM Comput. Surv. , vol. 41, no. 3, jul 2009. [Online]. Available: https://doi.org/10.1145/1541880.1541882
2009
Earlier work this paper cites.
S. Kolahi and L. V. S. Lakshmanan, “On approximating optimum repairs for functional dependency violations,” in Proceedings of the 12th International Conference on Database Theory , ser. ICDT ’09. New York, NY, USA: Association for Computing Machinery, 2009, p. 53–62. [Online]. Available: https://doi.org/10.1145/1514894.1514901
2009
Earlier work this paper cites.
L. Bertossi, S. Kolahi, and L. V. S. Lakshmanan, “Data cleaning and query answering with matching dependencies and matching functions,” in Proceedings of the 14th International Conference on Database Theory , ser. ICDT ’11. New York, NY, USA: Association for Computing Machinery, 2011, p. 268–279. [Online]. Available: https://doi.org/10.1145/1938551.1938585
2011
Earlier work this paper cites.
J. G. Moreno-Torres, T. Raeder, R. Alaiz-Rodríguez, N. V. Chawla, and F. Herrera, “A unifying view on dataset shift in classification,” Pattern Recognition , vol. 45, no. 1, pp. 521–530, 2012. [Online]. Available: https://www.sciencedirect.com/science/article/pii/S0031320311002901
2012
Earlier work this paper cites.
J. Freire and C. T. Silva, “Making computations and publications reproducible with vistrails,” Computing in Science & Engineering , vol. 14, no. 4, pp. 18–25, 2012
2012
Earlier work this paper cites.
P. J. Guo and M. I. Seltzer, “Burrito: Wrapping your lab notebook in computational infrastructure,” 2012
2012
Earlier work this paper cites.
M. Lin, H. Lucas, and G. Shmueli, “Too big to fail: Large samples and the p-value problem,” Information Systems Research , vol. 24, pp. 906–917, 12 2013
2013
Earlier work this paper cites.
S. Agarwal, B. Mozafari, A. Panda, H. Milner, S. Madden, and I. Stoica, “Blinkdb: queries with bounded errors and bounded response times on very large data,” in Proceedings of the 8th ACM European Conference on Computer Systems , 2013, pp. 29–42
2013
Earlier work this paper cites.
J. a. Gama, I. Žliobaitundefined, A. Bifet, M. Pechenizkiy, and A. Bouchachia, “A survey on concept drift adaptation,” ACM Comput. Surv. , vol. 46, no. 4, mar 2014. [Online]. Available: https://doi.org/10.1145/2523813
2014
Earlier work this paper cites.
V. L. Parsons, “Stratified sampling,” Wiley StatsRef: Statistics Reference Online , pp. 1–11, 2014
2014
Earlier work this paper cites.
D. Sculley et al. , “Hidden technical debt in ml systems,” in NIPS , 2015
2015
Earlier work this paper cites.
M. Mousavi, A. A. Bakar, and M. Vakilian, “Data stream clustering algorithms: A review,” in SOCO 2015 , 2015
2015
Earlier work this paper cites.
M. Germain, K. Gregor, I. Murray, and H. Larochelle, “Made: Masked autoencoder for distribution estimation,” in Proceedings of the 32nd International Conference on Machine Learning , ser. Proceedings of Machine Learning Research, F. Bach and D. Blei, Eds., vol. 37. Lille, France: PMLR, 07–09 Jul 2015, pp. 881–889. [Online]. Available: https://proceedings.mlr.press/v37/germain15.html
2015
Earlier work this paper cites.
N. Prokoshyna, J. Szlichta, F. Chiang, R. J. Miller, and D. Srivastava, “Combining quantitative and logical data cleaning,” Proc. VLDB Endow. , vol. 9, no. 4, p. 300–311, dec 2015. [Online]. Available: https://doi.org/10.14778/2856318.2856325
2015
Earlier work this paper cites.
K. Wongsuphasawat, D. Moritz, A. Anand, J. Mackinlay, B. Howe, and J. Heer, “Voyager: Exploratory analysis via faceted browsing of visualization recommendations,” IEEE transactions on visualization and computer graphics , vol. 22, no. 1, pp. 649–658, 2015
2015
Cited alongside, same era.
M. Vartak, “Modeldb: a system for machine learning model management,” in HILDA ’16 , 2016
2016
Cited alongside, same era.
S. Kandula, A. Shanbhag, A. Vitorovic, M. Olma, R. Grandl, S. Chaudhuri, and B. Ding, “Quickr: Lazily approximating complex adhoc queries in bigdata clusters,” in Proceedings of the 2016 International Conference on Management of Data , ser. SIGMOD ’16. New York, NY, USA: Association for Computing Machinery, 2016, p. 631–646. [Online]. Available: https://doi.org/10.1145/2882903.2882940
2016
Cited alongside, same era.
Z. Abedjan et al. , “Detecting data errors: Where are we and what needs to be done?” Proc. VLDB Endow. , vol. 9, no. 12, p. 993–1004, Aug. 2016. [Online]. Available: https://doi.org/10.14778/2994509.2994518
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers) . Minneapolis, Minnesota: Association for Computational Linguistics, Jun. 2019, pp. 4171–4186. [Online]. Available: https://aclanthology.org/N19-1423
2019
Later among the works it cites.
E. Wallace, Y. Wang, S. Li, S. Singh, and M. Gardner, “Do nlp models know numbers? probing numeracy in embeddings,” in EMNLP , 2019
2019
Later among the works it cites.
E. Rezig et al. , “Dagger: A data (not code) debugger,” in CIDR , 2020
2020
Later among the works it cites.
L. Biewald, “Tracking with weights and biases www.wandb.com/
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2016
Cited alongside, same era.
C. Pit-Claudel, Z. E. Mariet, R. Harding, and S. Madden, “Outlier detection in heterogeneous datasets using automatic tuple expansion,” 2016
2016
Cited alongside, same era.
S. Krishnan, J. Wang, E. Wu, M. J. Franklin, and K. Goldberg, “Activeclean: Interactive data cleaning for statistical modeling,” Proc. VLDB Endow. , vol. 9, no. 12, p. 948–959, aug 2016. [Online]. Available: https://doi.org/10.14778/2994509.2994514
2016
Cited alongside, same era.
——, “Towards a general-purpose query language for visualization recommendation,” in Proceedings of the Workshop on Human-In-the-Loop Data Analytics , 2016, pp. 1–6
2016
Cited alongside, same era.
H. Miao, A. Li, L. Davis, and A. Deshpande, “Towards unified data and lifecycle management for deep learning,” in ICDE’17 , 2017
2017
Cited alongside, same era.
E. Breck et al. , “The ml test score: A rubric for ml production readiness and technical debt reduction,” in Big Data’17 , 2017
2017
Cited alongside, same era.
N. Polyzotis, S. Roy, S. E. Whang, and M. Zinkevich, “Data management challenges in production machine learning,” in SIGMOD ’17 , 2017
2017
Cited alongside, same era.
A. N. Modi et al. , “Tfx: A tensorflow-based production-scale machine learning platform,” in KDD 2017 , 2017
2017
Cited alongside, same era.
P. Bailis, E. Gan, S. Madden, D. Narayanan, K. Rong, and S. Suri, “Macrobase: Prioritizing attention in fast data,” in Proceedings of the 2017 ACM International Conference on Management of Data , 2017, pp. 541–556
2017
Cited alongside, same era.
2020
Later among the works it cites.
2020
Later among the works it cites.
C. Ré, F. Niu, P. Gudipati, and C. Srisuwananukorn, “Overton: A data system for monitoring and improving machine-learned products,” in CIDR , 2020
2020
Later among the works it cites.
“Tlc trip record data,” 2020. [Online]. Available: https://www1.nyc.gov/site/tlc/about/tlc-trip-record-data.page
2020
Later among the works it cites.
A. Chapman, P. Missier, G. Simonelli, and R. Torlone, “Capturing and querying fine-grained provenance of preprocessing pipelines in data science,” Proceedings of the VLDB Endowment , vol. 14, no. 4, pp. 507–520, 2020
2020
Later among the works it cites.
M. H. Namaki, A. Floratou, F. Psallidas, S. Krishnan, A. Agrawal, Y. Wu, Y. Zhu, and M. Weimer, “Vamsa: Automated provenance tracking in data science scripts,” in Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining , 2020, pp. 1542–1551
2020
Later among the works it cites.
2020
Later among the works it cites.
Y. Luo, C. Chai, X. Qin, N. Tang, and G. Li, “Visclean: Interactive cleaning for progressive visualization,” Proc. VLDB Endow. , vol. 13, no. 12, p. 2821–2824, aug 2020. [Online]. Available: https://doi.org/10.14778/3415478.3415484
2020
Later among the works it cites.
S. Saisubramanian, S. Galhotra, and S. Zilberstein, Balancing the Tradeoff Between Clustering Value and Interpretability . New York, NY, USA: Association for Computing Machinery, 2020, p. 351–357. [Online]. Available: https://doi.org/10.1145/3375627.3375843
2020
Later among the works it cites.
J. Leskovec, A. Rajaraman, and J. D. Ullman, Mining of massive data sets . Cambridge university press, 2020
2020
Later among the works it cites.
S. Grafberger, S. Guha, J. Stoyanovich, and S. Schelter, “Mlinspect: A data distribution debugger for machine learning pipelines,” in SIGMOD’21 , 2021
2021
Closest in time.
R. Garcia et al. , “Hindsight logging for model training,” in VLDB’21 , 2021
2021
Closest in time.
S. Karumuri, F. Solleza, S. Zdonik, and N. Tatbul, “Towards observability data management at scale,” ACM SIGMOD Record , vol. 49, no. 4, pp. 18–23, 2021
2021
Closest in time.
C. Majors, “Observability: A manifesto,” Jul 2021. [Online]. Available: https://www.honeycomb.io/blog/observability-a-manifesto/
2021
Closest in time.
S. Sagadeeva and M. Boehm, SliceLine: Fast, Linear-Algebra-Based Slice Finding for ML Model Debugging . New York, NY, USA: Association for Computing Machinery, 2021, p. 2290–2299. [Online]. Available: https://doi.org/10.1145/3448016.3457323
2021
Closest in time.
S. Santurkar, D. Tsipras, and A. Madry, “Breeds: Benchmarks for subpopulation shift,” arXiv: Computer Vision and Pattern Recognition , 2021
2021
Closest in time.
2021
Closest in time.
X. Liang, S. Sintos, Z. Shang, and S. Krishnan, Combining Aggregation and Sampling (Nearly) Optimally for Approximate Query Processing . New York, NY, USA: Association for Computing Machinery, 2021, p. 1129–1141. [Online]. Available: https://doi.org/10.1145/3448016.3457277
2021
Closest in time.
C. M. Ellis, “What is adversarial validation?” Jul 2021. [Online]. Available: https://www.kaggle.com/carlmcbrideellis/what-is-adversarial-validation
2021
Closest in time.
D. J.-L. Lee, V. Setlur, M. Tory, K. G. Karahalios, and A. Parameswaran, “Deconstructing categorization in visualization recommendation: A taxonomy and comparative study,” IEEE Transactions on Visualization and Computer Graphics , 2021
2021
Closest in time.
D. J.-L. Lee, D. Tang, K. Agarwal, T. Boonmark, C. Chen, J. Kang, U. Mukhopadhyay, J. Song, M. Yong, M. A. Hearst et al. , “Lux: always-on visualization recommendations for exploratory dataframe workflows,” Proceedings of the VLDB Endowment , vol. 15, no. 3, pp. 727–738, 2021
2021
Closest in time.
“Overview of kubeflow pipelines,” Apr 2021. [Online]. Available: https://www.kubeflow.org/docs/components/pipelines/overview/pipelines-overview/
2021
Closest in time.
S. P. Kanuparthy, P. Kanuparthy, Meta, S. A. Dalakoti, A. Dalakoti, S. K. Bhalla, and K. Bhalla, “Ml monitoring & observability @meta scale,” May 2022. [Online]. Available: https://atscaleconference.com/videos/ml-monitoring-observability-meta-scale/
2022
Closest in time.
M. Mathur, “Full-spectrum ml model monitoring at lyft,” Jun 2022. [Online]. Available: https://eng.lyft.com/full-spectrum-ml-model-monitoring-at-lyft-a4cdaf828e8f
2022
Closest in time.
S. Oppold and M. Herschel, “Provenance-based explanations: are they useful?” in Proceedings of the 14th International Workshop on the Theory and Practice of Provenance , 2022, pp. 1–4
2022
Closest in time.
F. Chirigati, R. Rampin, D. Shasha, and J. Freire, “Reprozip: Computational reproducibility with ease,” in Proceedings of the 2016 international conference on management of data , 2016, pp. 2085–2088
2088
Closest in time.