Fetching the paper…
Reading the bibliography…
Machine learning (ML) applications become increasingly common in many domains.
Potter’s Wheel: An Interactive Data Cleaning System
V. Raman and J. M. Hellerstein · 2001
Earlier work this paper cites.
Modern applied statistics with S, 4th Ed
W. N. Venables and B. D. Ripley · 2002
Earlier work this paper cites.
A Survey of Outlier Detection Methodologies
V. J. Hodge and J. Austin · 2004
Earlier work this paper cites.
Model Management 2.0: Manipulating Richer Mappings
P. A. Bernstein and S. Melnik · 2007
Earlier work this paper cites.
Adaptive Query Processing
A. Deshpande et al · 2007
Earlier work this paper cites.
Anomaly Detection: A Survey
V. Chandola et al · 2009
Earlier work this paper cites.
MAD Skills: New Analysis Practices for Big Data
J. Cohen et al · 2009
Earlier work this paper cites.
Fully homomorphic encryption using ideal lattices
C. Gentry · 2009
Earlier work this paper cites.
An Architecture for Recycling Intermediates in a Column-store
M. Ivanova et al · 2009
Earlier work this paper cites.
SystemML: Declarative Machine Learning on MapReduce
A. Ghoting et al · 2011
Earlier work this paper cites.
Wrangler: Interactive Visual Specification of Data Transformation Scripts
S. Kandel et al · 2011
Earlier work this paper cites.
Scikit-learn: Machine Learning in Python
F. Pedregosa et al · 2011
Earlier work this paper cites.
The Architecture of SciDB
M. Stonebraker et al · 2011
Earlier work this paper cites.
mice: Multivariate Imputation by Chained Equations in R
S. van Buuren and K. Groothuis-Oudshoorn · 2011
Earlier work this paper cites.
The NumPy Array: A Structure for Efficient Numerical Computation
S. van der Walt et al · 2011
Earlier work this paper cites.
NoDB: Efficient Query Execution on Raw Data Files
I. Alagiannis et al · 2012
Earlier work this paper cites.
GLADE: Big Data Analytics Made Easy
Y. Cheng et al · 2012
Earlier work this paper cites.
Towards a Unified Architecture for in-RDBMS Analytics
X. Feng et al · 2012
Earlier work this paper cites.
The MADlib Analytics Library or MAD Skills, the SQL
J. M. Hellerstein et al · 2012
Earlier work this paper cites.
Resilient Distributed Datasets: A Fault-Tolerant Abstraction for In-Memory Cluster Computing
M. Zaharia et al · 2012
Earlier work this paper cites.
Hybrid Parallelization Strategies for Large- Scale Machine Learning in SystemML
M. Boehm et al · 2014
Earlier work this paper cites.
SystemML’s Optimizer: Plan Generation for Large-Scale Machine Learning Programs
M. Boehm et al · 2014
Earlier work this paper cites.
Interpretable and Informative Explanations of Outcomes
K. E. Gebaly et al · 2014
Earlier work this paper cites.
Materialization Optimizations for Feature Selection Workloads
C. Zhang et al · 2014
Earlier work this paper cites.
MXNet: A Flexible and Efficient Machine Learning Library for Heterogeneous Distributed Systems
T. Chen et al · 2015
Earlier work this paper cites.
Predictive Interaction for Data Transformation
J. Heer et al · 2015
Earlier work this paper cites.
Resource Elasticity for Large-Scale Machine Learning
B. Huang et al · 2015
Earlier work this paper cites.
Learning Generalized Linear Models Over Normalized Data
A. Kumar et al · 2015
Cited alongside, same era.
Hidden Technical Debt in Machine Learning Systems
D. Sculley et al · 2015
Cited alongside, same era.
WANalytics: Analytics for a Geo-Distributed Data-Intensive World
A. Vulimiri et al · 2015
Cited alongside, same era.
TensorFlow: A System for Large-Scale Machine Learning
M. Abadi et al · 2016
Cited alongside, same era.
SystemML: Declarative Machine Learning on Spark
M. Boehm et al · 2016
Cited alongside, same era.
Dask: Library for dynamic task scheduling
Dask · 2016
Cited alongside, same era.
Compressed Linear Algebra for Large-Scale Machine Learning
Northstar: An Interactive Data Science System
T. Kraska · 2018
Later among the works it cites.
PRETZEL: Opening the Black Box of Machine Learning Prediction Serving Systems
Y. Lee et al · 2018
Later among the works it cites.
Incremental View Maintenance with Triple Lock Factorization Benefits
M. Nikolic and D. Olteanu · 2018
Later among the works it cites.
Filter Before You Parse: Faster Analytics on Raw Data with Sparser
S. Palkar et al · 2018
Later among the works it cites.
Data Lifecycle Challenges in Production Machine Learning: A Survey
N. Polyzotis et al · 2018
Later among the works it cites.
Automating Large-Scale Data Quality Verification
S. Schelter et al · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Elgohary et al · 2016
Cited alongside, same era.
Fast Queries Over Heterogeneous Data Through Engine Customization
M. Karpathiotakis et al · 2016
Cited alongside, same era.
MLlib: Machine Learning in Apache Spark
X. Meng et al · 2016
Cited alongside, same era.
Samsara: Declarative Machine Learning on Distributed Dataflow Systems
S. Schelter et al · 2016
Cited alongside, same era.
Learning Linear Regression Models over Factorized Joins
M. Schleich et al · 2016
Cited alongside, same era.
TFX: A TensorFlow-Based Production- Scale Machine Learning Platform
D. Baylor et al · 2017
Cited alongside, same era.
AIDA - Abstraction for Advanced In-Database Analytics
J. V. D. silva et al · 2018
Later among the works it cites.
Data Integration: The Current Status and the Way Forward
M. Stonebraker and I. F. Ilyas · 2018
Later among the works it cites.
MISTIQUE: A System to Store and Query Model Intermediates for Model Diagnosis
M. Vartak et al · 2018
Later among the works it cites.
Helix: Holistic Optimization for Accelerating Iterative Machine Learning
D. Xin et al · 2018
Later among the works it cites.
Accelerating the Machine Learning Lifecycle with MLflow
M. Zaharia et al · 2018
Later among the works it cites.
Towards Federated Learning at Scale: System Design
K. Bonawitz et al · 2019
Closest in time.
Slice Finder: Automated Data Slicing for Model Validation
Y. Chung et al · 2019
Closest in time.
AutoAugment: Learning Augmentation Policies from Data
E. D. Cubuk et al · 2019
Closest in time.
Auto-sklearn: Efficient and Robust Automated Machine Learning
M. Feurer et al · 2019
Closest in time.
Accelerating the Unacceleratable: Hybrid CPU/GPU Algorithms for Memory-Bound Database Primitives
M. Gowanlock et al · 2019
Closest in time.
HoloDetect: Few-Shot Learning for Error Detection
A. Heidari et al · 2019
Closest in time.
Sherlock: A Deep Learning Approach to Semantic Data Type Detection
M. Hulsebos et al · 2019
Closest in time.
An Intermediate Representation for Opti- mizing Machine Learning Pipelines
A. Kunft et al · 2019
Closest in time.
Tuple-oriented Compression for Large-scale Mini-batch SGD
F. Li et al · 2019
Closest in time.
AutoGraph: Imperative-style Coding with Graph-based Performance
D. Moldovan et al · 2019
Closest in time.
TPOT: A Tree-Based Pipeline Optimization Tool for Automating Machine Learning
R. S. Olson and J. H. Moore · 2019
Closest in time.
PyTorch: An Imperative Style, High- Performance Deep Learning Library
A. Paszke et al · 2019
Closest in time.
Democratizing Data Science through Interactive Curation of ML Pipelines
Z. Shang et al · 2019
Closest in time.
MNC: Structure-Exploiting Sparsity Estimation for Matrix Expressions
J. Sommer et al · 2019
Closest in time.
Federated AI for the Enterprise: A Web Services Based Implementation
D. C. Verma et al · 2019
Closest in time.
Sato: Contextual Semantic Type Detection in Tables
D. Zhang et al · 2019
Closest in time.