Fetching the paper…
Reading the bibliography…
Practitioners are increasingly turning to Extract-Load-Transform (ELT) pipelines with the widespread adoption of cloud data warehouses.
Data manipulation in heterogeneous databases
Abhirup Chatterjee and Arie Segev. 1991 · 1991
Earlier work this paper cites.
Research problems in data warehousing. In Proceedings of the Fourth International Conference on Information and Knowledge Management (Baltimore, Maryland, USA) (CIKM ’95) . Association for Computing Machinery, New York, NY, USA, 25–30
Jennifer Widom. 1995 · 1995
Earlier work this paper cites.
Real-world Data is Dirty: Data Cleansing and The Merge/Purge Problem
Mauricio A. Hernández and Salvatore J. Stolfo. 1998 · 1998
Earlier work this paper cites.
Conceptual modeling for ETL processes. In Proceedings of the 5th ACM International Workshop on Data Warehousing and OLAP (McLean, Virginia, USA) (DOLAP ’02) . Association for Computing Machinery, New York, NY, USA, 14–21
Panos Vassiliadis, Alkis Simitsis, and Spiros Skiadopoulos. 2002 · 2002
Earlier work this paper cites.
A UML Based Approach for Modeling ETL Processes in Data Warehouses, Vol. 2813. 307–320
Juan Trujillo and Sergio Luján-Mora. 2003 · 2003
Earlier work this paper cites.
The Data Warehouse ETL Toolkit: Practical Techniques for Extracting, Cleaning, Conforming and Delivering Data
Ralph Kimball and Joe Caserta. 2004 · 2004
Earlier work this paper cites.
Data Mapping Diagrams for Data Warehouse Design with UML, Vol. 3288. 191–204
Sergio Luján-Mora, Panos Vassiliadis, and Juan Trujillo. 2004 · 2004
Earlier work this paper cites.
Designing ETL processes using semantic web technologies. In Proceedings of the 9th ACM International Workshop on Data Warehousing and OLAP (Arlington, Virginia, USA) (DOLAP ’06) . Association for Computing Machinery, New York, NY, USA, 67–74
Dimitrios Skoutas and Alkis Simitsis. 2006 · 2006
Earlier work this paper cites.
Automating the loading of business process data warehouses. In Proceedings of the 12th International Conference on Extending Database Technology: Advances in Database Technology (Saint Petersburg, Russia) (EDBT ’09) . Association for Computing Machinery, New York, NY, USA, 612–623
Malu Castellanos, Alkis Simitsis, Kevin Wilkinson, and Umeshwar Dayal. 2009 · 2009
Earlier work this paper cites.
On-Demand ELT Architecture for Right-Time BI: Extending the Vision
Florian Waas, Robert Wrembel, Tobias Freudenreich, Maik Thiele, Christian Koncilia, and Pedro Furtado. 2013 · 2013
Earlier work this paper cites.
TPC-DI: the first industry benchmark for data integration
Meikel Poess, Tilmann Rabl, Hans-Arno Jacobsen, and Brian Caufield. 2014 · 2014
Earlier work this paper cites.
Effort estimation of ETL projects using Forward Stepwise Regression. In 2015 International Conference on Emerging Technologies (ICET) . 1–6
Raza Rasool and Ali Afzal Malik. 2015 · 2015
Earlier work this paper cites.
The Snowflake Elastic Data Warehouse. In Proceedings of the 2016 International Conference on Management of Data (San Francisco, California, USA) (SIGMOD ’16) . Association for Computing Machinery, New York, NY, USA, 215–226
Benoit Dageville, Thierry Cruanes, Marcin Zukowski, Vadim Antonov, Artin Avanes, Jon Bock, Jonathan Claybaugh, Daniel Engovatov, Martin Hentschel, Jiansheng Huang, Allison W. Lee, Ashish Motivala, Abdul Q. Munir, Steven Pelley, Peter Povinec, Greg Rahn, Spyridon Triantafyllis, and Philipp Unterbrunner. 2016 · 2016
Earlier work this paper cites.
Learning a Neural Semantic Parser from User Feedback
Srinivasan Iyer, Ioannis Konstas, Alvin Cheung, Jayant Krishnamurthy, and Luke Zettlemoyer. 2017 · 2017
Earlier work this paper cites.
SQLNet: Generating Structured Queries From Natural Language Without Reinforcement Learning
Xiaojun Xu, Chang Liu, and Dawn Song. 2017 · 2017
Earlier work this paper cites.
SQLizer: query synthesis from natural language
Navid Yaghmazadeh, Yuepeng Wang, Isil Dillig, and Thomas Dillig. 2017 · 2017
Earlier work this paper cites.
Seq2SQL: Generating Structured Queries from Natural Language using Reinforcement Learning
Victor Zhong, Caiming Xiong, and Richard Socher. 2017 · 2017
Earlier work this paper cites.
Improving Text-to-SQL Evaluation Methodology. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . Association for Computational Linguistics, 351–360
Catherine Finegan-Dollak, Jonathan K. Kummerfeld, Li Zhang, Karthik Ramanathan, Sesh Sadasivam, Rui Zhang, and Dragomir Radev. 2018 · 2018
Earlier work this paper cites.
Tao Yu, Rui Zhang, Kai Yang, Michihiro Yasunaga, Dongxu Wang, Zifan Li, James Ma, Irene Li, Qingning Yao, Shanelle Roman, Zilin Zhang, and Dragomir Radev. 2019 · 2019
Earlier work this paper cites.
Role of Machine Learning in ETL Automation. In Proceedings of the 21st International Conference on Distributed Computing and Networking (Kolkata, India) (ICDCN ’20) . Association for Computing Machinery, New York, NY, USA, Article 57, 6 pages
Kartick Chandra Mondal, Neepa Biswas, and Swati Saha. 2020 · 2020
Earlier work this paper cites.
LGESQL: Line Graph Enhanced Text-to-SQL Model with Mixed Local and Non-Local Relations
Ruisheng Cao, Lu Chen, Zhi Chen, Yanbin Zhao, Su Zhu, and Kai Yu. 2021 · 2021
Earlier work this paper cites.
DS-1000: A Natural and Reliable Benchmark for Data Science Code Generation
Yuhang Lai, Chengxi Li, Yiming Wang, Tianyi Zhang, Ruiqi Zhong, Luke Zettlemoyer, Scott Wen tau Yih, Daniel Fried, Sida Wang, and Tao Yu. 2022 · 2022
Cited alongside, same era.
ETL vs ELT: Choosing the right approach for your data warehouse
Dhamotharan Seenivasan. 2022 · 2022
Cited alongside, same era.
ETL, ELT and reverse ETL: a business case Study. In 2022 Second International Conference on Advanced Technologies in Intelligent Control, Environment, Computing & Communication Engineering (ICATIECE) . IEEE, 1–4
Bharat Singhal and Alok Aggarwal. 2022 · 2022
Cited alongside, same era.
C3: Zero-shot Text-to-SQL with ChatGPT
Xuemei Dong, Chao Zhang, Yuhang Ge, Yuren Mao, Yunjun Gao, lu Chen, Jinshu Lin, and Dongfang Lou. 2023 · 2023
Cited alongside, same era.
Data Interpreter: An LLM Agent For Data Science
Sirui Hong, Yizhang Lin, Bang Liu, Bangbang Liu, Binhao Wu, Ceyao Zhang, Chenxing Wei, Danyang Li, Jiaqi Chen, Jiayi Zhang, Jinlin Wang, Li Zhang, Lingyao Zhang, Min Yang, Mingchen Zhuge, Taicheng Guo, Tuo Zhou, Wei Tao, Xiangru Tang, Xiangtao Lu, Xiawu Zheng, Xinbing Liang, Yaying Fei, Yuheng Cheng, Zhibin Gou, Zongze Xu, and Chenglin Wu. 2024 · 2024
Later among the works it cites.
InfiAgent-DABench: Evaluating Agents on Data Analysis Tasks. In Proceedings of the 41st International Conference on Machine Learning (Proceedings of Machine Learning Research) , Ruslan Salakhutdinov, Zico Kolter, Katherine Heller, Adrian Weller, Nuria Oliver, Jonathan Scarlett, and Felix Berkenkamp (Eds.), Vol. 235. PMLR, 19544–19572
Xueyu Hu, Ziyu Zhao, Shuang Wei, Ziwei Chai, Qianli Ma, Guoyin Wang, Xuwu Wang, Jing Su, Jingjing Xu, Ming Zhu, Yao Cheng, Jianbo Yuan, Jiwei Li, Kun Kuang, Yang Yang, Hongxia Yang, and Fei Wu. 2024 · 2024
Later among the works it cites.
DA-Code: Agent Data Science Code Generation Benchmark for Large Language Models. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (Eds.). Association for Computational Linguistics, Miami, Florida, USA, 13487–13521
Yiming Huang, Jianwen Luo, Yan Yu, Yitong Zhang, Fangyu Lei, Yifan Wei, Shizhu He, Lifu Huang, Xiao Liu, Jun Zhao, and Kang Liu. 2024a · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Harald Foidl, Valentina Golendukhina, Rudolf Ramler, and Michael Felderer. 2024 · 2023
Cited alongside, same era.
Large Language Models are Zero-Shot Reasoners
Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2023 · 2023
Cited alongside, same era.
ETL vs ELT: What’s the difference?
Daniel Poppy. 2023 · 2023
Cited alongside, same era.
DIN-SQL: Decomposed In-Context Learning of Text-to-SQL with Self-Correction
Mohammadreza Pourreza and Davood Rafiei. 2023 · 2023
Cited alongside, same era.
ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs
Yujia Qin, Shihao Liang, Yining Ye, Kunlun Zhu, Lan Yan, Yaxi Lu, Yankai Lin, Xin Cong, Xiangru Tang, Bill Qian, Sihan Zhao, Lauren Hong, Runchu Tian, Ruobing Xie, Jie Zhou, Mark Gerstein, Dahai Li, Zhiyuan Liu, and Maosong Sun. 2023 · 2023
Cited alongside, same era.
Toolformer: Language Models Can Teach Themselves to Use Tools
Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom. 2023 · 2023
Cited alongside, same era.
The History, Present, and Future of ETL Technology (invited). In Proceedings of the 25th International Workshop on Design, Optimization, Languages and Analytical Processing of Big Data (DOLAP) co-located with the 26th International Conference on Extending Database Technology and the 26th International Conference on Database Theory (EDBT/ICDT 2023), Ioannina, Greece, March 28, 2023 (CEUR Workshop Proceedings) , Enrico Gallinucci and Lukasz Golab (Eds.), Vol. 3369. CEUR-WS.org, 3–12
Alkis Simitsis, Spiros Skiadopoulos, and Panos Vassiliadis. 2023 · 2023
Cited alongside, same era.
Cloud Analytics Benchmark
Alexander van Renen and Viktor Leis. 2023 · 2023
Cited alongside, same era.
Later among the works it cites.
SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan. 2024 · 2024
Later among the works it cites.
AutoWebGLM: A Large Language Model-based Web Navigating Agent
Hanyu Lai, Xiao Liu, Iat Long Iong, Shuntian Yao, Yuxuan Chen, Pengbo Shen, Hao Yu, Hanchen Zhang, Xiaohan Zhang, Yuxiao Dong, and Jie Tang. 2024 · 2024
Later among the works it cites.
Spider 2.0: Evaluating Language Models on Real-World Enterprise Text-to-SQL Workflows
Fangyu Lei, Jixuan Chen, Yuxiao Ye, Ruisheng Cao, Dongchan Shin, Hongjin Su, Zhaoqing Suo, Hongcheng Gao, Wenjing Hu, Pengcheng Yin, Victor Zhong, Caiming Xiong, Ruoxi Sun, Qian Liu, Sida Wang, and Tao Yu. 2024 · 2024
Later among the works it cites.
A Survey of Pipeline Tools for Data Engineering
Anthony Mbata, Yaji Sripada, and Mingjun Zhong. 2024 · 2024
Later among the works it cites.
OpenAI. 2024 · 2024
Later among the works it cites.
Autonomous Evaluation and Refinement of Digital Agents
Jiayi Pan, Yichi Zhang, Nicholas Tomlin, Yifei Zhou, Sergey Levine, and Alane Suhr. 2024 · 2024
Later among the works it cites.
Executable Code Actions Elicit Better LLM Agents
Xingyao Wang, Yangyi Chen, Lifan Yuan, Yizhe Zhang, Yunzhu Li, Hao Peng, and Heng Ji. 2024 · 2024
Later among the works it cites.
Understanding the Effects of Noise in Text-to-SQL: An Examination of the BIRD-Bench Benchmark
Niklas Wretblad, Fredrik Gordh Riseby, Rahul Biswas, Amin Ahmadi, and Oskar Holmström. 2024 · 2024
Later among the works it cites.
SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering
John Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret, Shunyu Yao, Karthik Narasimhan, and Ofir Press. 2024 · 2024
Later among the works it cites.
τ \tau -bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
Shunyu Yao, Noah Shinn, Pedram Razavi, and Karthik Narasimhan. 2024 · 2024
Later among the works it cites.
AutoCodeRover: Autonomous Program Improvement
Yuntong Zhang, Haifeng Ruan, Zhiyu Fan, and Abhik Roychoudhury. 2024 · 2024
Later among the works it cites.
WebArena: A Realistic Web Environment for Building Autonomous Agents
Shuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig. 2024 · 2024
Later among the works it cites.
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
DeepSeek-AI. 2025 · 2025
Closest in time.
A Preview of XiYan-SQL: A Multi-Generator Ensemble Framework for Text-to-SQL
Yingqi Gao, Yifu Liu, Xiaoxia Li, Xiaorong Shi, Yin Zhu, Yiming Wang, Shiqi Li, Wei Li, Yuntao Hong, Zhiling Luo, Jinyang Gao, Liyu Mou, and Yu Li. 2025 · 2025
Closest in time.
LocalStack: A Fully Functional Local AWS Cloud Stack
LocalStack. 2025 · 2025
Closest in time.
Qwen. 2025 · 2025
Closest in time.
airbyte Provider
Terraform. 2025 · 2025
Closest in time.