Fetching the paper…
Reading the bibliography…
Currently, a variety of pipeline tools are available for use in data engineering.
“Exploratory data mining and data cleaning”
Tamraparni Dasu and Theodore Johnson · 2003
Earlier work this paper cites.
“Data integration and machine learning: A natural synergy”
Xin Dong and Theodoros Rekatsinas · 2018
Earlier work this paper cites.
“Tensorflow data validation: Data analysis and validation in continuous ml pipelines”
Emily Caveness et al · 2020
Earlier work this paper cites.
“Lineage checkpoint approach for long-lineage problem in apache spark”
Minhyeok Kweun et al · 2020
Earlier work this paper cites.
“Data engineering for data analytics: A classification of the issues, and case studies”
Alfredo Nazabal et al · 2020
Earlier work this paper cites.
“A Machine learning pipeline generation approach for data analysis”
Zhao Ru-tao et al · 2020
Earlier work this paper cites.
“Evaluation of stream processing frameworks”
Giselle Van and Dirk Van · 2020
Earlier work this paper cites.
“High performance data engineering everywhere”
Chathura Widanage et al · 2020
Cited alongside, same era.
“An overview of current trends in data ingestion and integration”
Tomislav Hlupić and Josip Puniš · 2021
Cited alongside, same era.
“The IDEAL household energy dataset, electricity, gas, contextual sensor data and survey data for 255 UK homes”
Martin Pullinger et al · 2021
Cited alongside, same era.
“Cost-Aware Resource Recommendation for DAG-Based Big Data Workflows: An Apache Spark Case Study”
Mohammad-Mohsen Aseman-Manzar et al · 2022
Cited alongside, same era.
“Automating data science”
Tijl De et al · 2022
Cited alongside, same era.
“Operationalizing and automating data governance”
Sergi Nadal, Petar Jovanovic, Besim Bilalli and Oscar Romero · 2022
Cited alongside, same era.
“Real-Time Blood Pressure Prediction Using Apache Spark and Kafka Machine Learning”
Ali Farki and Elham Noughabi · 2023
Later among the works it cites.
“Autokeras: An automl library for deep learning”
Haifeng Jin, François Chollet, Qingquan Song and Xia Hu · 2023
Later among the works it cites.
“Streamlining Enterprise Data Processing, Reporting and Realtime Alerting using Apache Kafka”
Kiran Peddireddy · 2023
Later among the works it cites.
“The pipeline for the continuous development of artificial intelligence models—Current state of research and practice”
Monika Steidl, Michael Felderer and Rudolf Ramler · 2023
Later among the works it cites.
“A data integration tool for the integrated modeling and analysis for east”
Liu Xiaojuan and Zhi Yu · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Performance evaluation of apache kafka–a modern platform for real time data streaming”
Shubham Vyas, Rajesh Tyagi, Charu Jain and Shashank Sahu · 2022
Cited alongside, same era.
Harald Foidl, Valentina Golendukhina, Rudolf Ramler and Michael Felderer · 2023
Later among the works it cites.
“Data pipeline quality: Influencing factors, root causes of data-related issues, and processing problem areas for developers”
Harald Foidl, Valentina Golendukhina, Rudolf Ramler and Michael Felderer · 2024
Closest in time.