Fetching the paper…
Reading the bibliography…
Data cleaning is a crucial yet challenging task in data analysis, often requiring significant manual effort.
Data cleaning: Problems and current approaches
Erhard Rahm, Hong Hai Do, et al · 2000
Earlier work this paper cites.
Data quality and the bottom line
Wayne W Eckerson · 2002
Earlier work this paper cites.
Sampling the repairs of functional dependency violations under hard constraints
George Beskales, Ihab F. Ilyas, and Lukasz Golab · 2010
Earlier work this paper cites.
Eracer: a database approach for statistical inference and data cleaning
Chris Mayfield, Jennifer Neville, and Sunil Prabhakar · 2010
Earlier work this paper cites.
Profiler: Integrated statistical analysis and visualization for data quality assessment
Sean Kandel, Ravi Parikh, Andreas Paepcke, Joseph M Hellerstein, and Jeffrey Heer · 2012
Earlier work this paper cites.
Holistic data cleaning: Putting violations into context
Xu Chu, Ihab F Ilyas, and Paolo Papotti · 2013
Earlier work this paper cites.
Data cleaning: Overview and emerging challenges
Xu Chu, Ihab F Ilyas, Sanjay Krishnan, and Jiannan Wang · 2016
Earlier work this paper cites.
Rayyan—a web and mobile app for systematic reviews
Mourad Ouzzani, Hossam Hammady, Zbys Fedorowicz, and Ahmed Elmagarmid · 2016
Earlier work this paper cites.
Holoclean: Holistic data repairs with probabilistic inference
Theodoros Rekatsinas, Xu Chu, Ihab F Ilyas, and Christopher Ré · 2017
Cited alongside, same era.
Fahes: A robust disguised missing values detector
Abdulhakim A Qahtan, Ahmed Elmagarmid, Raul Castro Fernandez, Mourad Ouzzani, and Nan Tang · 2018
Cited alongside, same era.
Raha: A configuration-free error detection system
Mohammad Mahdavi, Ziawasch Abedjan, Raul Castro Fernandez, Samuel Madden, Mourad Ouzzani, Michael Stonebraker, and Nan Tang · 2019
Cited alongside, same era.
Baran: Effective error correction via a unified context representation and transfer learning
Mohammad Mahdavi and Ziawasch Abedjan · 2020
Cited alongside, same era.
“everyone wants to do the model work, not the data work”: Data cascades in high-stakes ai
Nithya Sambasivan, Shivani Kapania, Hannah Highfill, Diana Akrong, Praveen Paritosh, and Lora M Aroyo · 2021
Cited alongside, same era.
Sparks of artificial general intelligence: Early experiments with gpt-4
Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, et al · 2023
Later among the works it cites.
Victor Dibia · 2023
Later among the works it cites.
Dead or alive: Continuous data profiling for interactive data science
Will Epperson, Vaishnavi Gorantla, Dominik Moritz, and Adam Perer · 2023
Later among the works it cites.
Data ambiguity strikes back: How documentation improves gpt’s text-to-sql
Zezhou Huang, Pavan Kalyan Damalapati, and Eugene Wu · 2023
Later among the works it cites.
Cocoon: Semantic table profiling using large language models
Zezhou Huang and Eugene Wu · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tushar Khot, Harsh Trivedi, Matthew Finlayson, Yao Fu, Kyle Richardson, Peter Clark, and Ashish Sabharwal · 2022
Cited alongside, same era.
Retclean: Retrieval-based data cleaning using foundation models and data lakes, 2023
Mohammad Shahmeer Ahmad, Zan Ahmad Naeem, Mohamed Eltabakh, Mourad Ouzzani, and Nan Tang · 2023
Cited alongside, same era.
The magellan data repository
Sanjib Das, AnHai Doan, Paul Suganthan G. C., Chaitanya Gokhale, Pradap Konda, Yash Govind, and Derek Paulsen
Cited in the paper.
Transform table to database using large language models
Zezhou Huang, Jia Guo, and Eugene Wu
Cited in the paper.
Relationalizing tables with large language models: The promise and challenges
Zezhou Huang and Eugene Wu · 2024
Closest in time.
Cleanagent: Automating data standardization with llm-based agents, 2024
Danrui Qi and Jiannan Wang · 2024
Closest in time.