Understand
Orchestrating a high-quality data preparation program is essential for successful machine learning (ML), but it is known to be time and effort consuming.
- Despite the impressive capabilities of large language models like ChatGPT in generating programs by interacting with users through natural language prompts, there are still limitations.
- Specifically, a user must provide specific prompts to iteratively guide ChatGPT in improving data preparation programs, which requires a certain level of expertise in programming, the dataset used and the ML task.
- Moreover, once a program has been generated, it is non-trivial to revisit a previous version or make changes to the program without starting the process over again.
Built on
Optimizing machine learning workloads in collaborative environments. In SIGMOD . 1701–1716
Behrouz Derakhshan, Alireza Rezaei Mahdiraji, Ziawasch Abedjan, Tilmann Rabl, and Volker Markl. 2020 · 2020
Earlier work this paper cites.
Similar
Unixcoder: Unified cross-modal pre-training for code representation
Daya Guo, Shuai Lu, Nan Duan, Yanlin Wang, Ming Zhou, and Jian Yin. 2022 · 2022
Cited alongside, same era.
https://www.kaggle.com/datasets/akshaydattatraykhare/diabetes-dataset
Kaggle Diabetes Dataset. [n.d.]
Cited in the paper.
Then
HybridPipe: Combining Human-generated and Machine-generated Pipelines for Data Preparation. In SIGMOD
Sibei Chen, Nan Tang, Ju Fan, Xuemi Yan, Chengliang Chai, Guoliang Li, and Xiaoyong Du. 2023 (to appear) · 2023
Closest in time.
Beyond the bibliography
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…