Fetching the paper…
Reading the bibliography…
In recent years, data science agents powered by Large Language Models (LLMs), known as "data agents," have shown significant potential to transform the traditional data analysis paradigm.
SPSS Statistics
IBM (1968) · 1968
Earlier work this paper cites.
SAS Software
Inc., S. I. (1976) · 1976
Earlier work this paper cites.
Microsoft Excel
Microsoft (1985) · 1985
Earlier work this paper cites.
Python Programming Language
Foundation, P. S. (1991) · 1991
Earlier work this paper cites.
R: A Language and Environment for Statistical Computing
for Statistical Computing, R. F. (1995) · 1995
Earlier work this paper cites.
Big data: The next frontier for innovation, competition, and productivity
Institute, M. G. (2011) · 2011
Earlier work this paper cites.
Business intelligence and analytics: From big data to big impact
Chen, H., Chiang, R. H., and Storey, V. C. (2012) · 2012
Earlier work this paper cites.
Power BI
Microsoft (2013) · 2013
Earlier work this paper cites.
Data science and its relationship to big data and data-driven decision making
Provost, F. and Fawcett, T. (2013) · 2013
Earlier work this paper cites.
The data revolution: Big data, open data, data infrastructures and their consequences
Kitchin, R. (2014) · 2014
Earlier work this paper cites.
Machine learning: Trends, perspectives, and prospects
Jordan, M. I. and Mitchell, T. M. (2015) · 2015
Earlier work this paper cites.
Psamm: A portable system for the analysis of metabolic models
Steffensen, J. L., Dufault-Thompson, K., and Zhang, Y. (2016) · 2016
Earlier work this paper cites.
Data science, predictive analytics, and big data: A revolution that will transform supply chain design and management
Waller, M. A. and Fawcett, S. E. (2016) · 2016
Earlier work this paper cites.
Data Mining: Practical machine learning tools and techniques
Witten, I. H., Frank, E., and Hall, M. A. (2016) · 2016
Earlier work this paper cites.
Data science: Challenges and directions
Cao, L. (2017) · 2017
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W.-t., Rocktäschel, T., et al. (2020) · 2020
Earlier work this paper cites.
Training and evaluating a jupyter notebook data science assistant
Chandel, S., Clement, C. B., Serrato, G., and Sundaresan, N. (2022) · 2022
Earlier work this paper cites.
A review of some techniques for inclusion of domain-knowledge into deep neural networks
Dash, T., Chitlangia, S., Ahuja, A., and Srinivasan, A. (2022) · 2022
Earlier work this paper cites.
Ds-1000: A natural and reliable benchmark for data science code generation
Lai, Y., Li, C., Wang, Y., Zhang, T., Zhong, R., Zettlemoyer, L., tau Yih, S. W., Fried, D., Wang, S., and Yu, T. (2022) · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al. (2022) · 2022
Earlier work this paper cites.
React: Synergizing reasoning and acting in language models
Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K., and Cao, Y. (2022) · 2022
Earlier work this paper cites.
Chapyter
chapyter (2023) · 2023
Cited alongside, same era.
Seed: Domain-specific data curation with large language models. arxiv 2023
Chen, Z., Cao, L., Madden, S., Kraska, T., Shang, Z., Fan, J., Tang, N., Gu, Z., Liu, C., and Cafarella, M. (2024) · 2023
Cited alongside, same era.
Cheng, L., Li, X., and Bing, L. (2023) · 2023
Cited alongside, same era.
Mathematical capabilities of chatgpt
Frieder, S., Pinchetti, L., Chevalier, A., Griffiths, R.-R., Salvatori, T., Lukasiewicz, T., Petersen, P. C., and Berner, J. (2023) · 2023
Cited alongside, same era.
Chatgpt as your personal data scientist
Hassan, M. M., Knipper, A., and Santu, S. K. K. (2023) · 2023
Cited alongside, same era.
Chatglm: A family of large language models from glm-130b to glm-4 all tools
GLM, T. (2024) · 2024
Closest in time.
Large language models orchestrating structured reasoning achieve kaggle grandmaster level
Grosnit, A., Maraval, A., Doran, J., Paolo, G., Thomas, A., Beevi, R. S. H. N., Gonzalez, J., Khandelwal, K., Iacobacci, I., Benechehab, A., et al. (2024) · 2024
Closest in time.
Ds-agent: Automated data science by empowering large language models with case-based reasoning
Guo, S., Deng, C., Wen, Y., Chen, H., Chang, Y., and Wang, J. (2024) · 2024
Closest in time.
Data interpreter: An llm agent for data science
Hong, S., Lin, Y., Liu, B., Wu, B., Li, D., Chen, J., Zhang, J., Wang, J., Zhang, L., Zhuge, M., et al. (2024) · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jupyter-ai
jupyterlab (2023) · 2023
Cited alongside, same era.
Jarvix: A llm no code platform for tabular data analysis and optimization
Liu, S.-C., Wang, S., Chang, T., Lin, W., Hsiung, C.-W., Hsieh, Y.-C., Cheng, Y.-P., Luo, S.-H., and Zhang, J. (2023) · 2023
Cited alongside, same era.
Insightpilot: An llm-empowered automated data exploration system
Ma, P., Ding, R., Wang, S., Han, S., and Zhang, D. (2023) · 2023
Cited alongside, same era.
Llms for science: Usage for code generation and data analysis
Nejjar, M., Zacharias, L., Stiehle, F., and Weber, I. (2023) · 2023
Cited alongside, same era.
OpenAI (2023) · 2023
Cited alongside, same era.
Taskweaver: A code-first agent framework
Qiao, B., Li, L., Zhang, X., He, S., Kang, Y., Zhang, C., Yang, F., Dong, H., Zhang, J., Wang, L., et al. (2023) · 2023
Cited alongside, same era.
What should data science education do with large language models?
Tu, X., Zou, J., Su, W. J., and Zhang, L. (2023) · 2023
Cited alongside, same era.
Hu, X., Zhao, Z., Wei, S., Chai, Z., Ma, Q., Wang, G., Wang, X., Su, J., Xu, J., Zhu, M., et al. (2024) · 2024
Closest in time.
Data analysis in the era of generative ai
Inala, J. P., Wang, C., Drucker, S., Ramos, G., Dibia, V., Riche, N., Brown, D., Marshall, D., and Gao, J. (2024) · 2024
Closest in time.
AIDE: the Machine Learning CodeGen Agent
Jiang, Z. et al. (2024) · 2024
Closest in time.
Autokaggle: A multi-agent framework for autonomous data science competitions
Li, Z., Zang, Q., Ma, D., Guo, J., Zheng, T., Niu, X., Yue, X., Wang, Y., Yang, J., Liu, J., et al. (2024) · 2024
Closest in time.
Cmat: A multi-agent collaboration tuning framework for enhancing small language models
Liang, X., He, Y., Tao, M., Xia, Y., Wang, J., Shi, T., Wang, J., and Yang, J. (2024) · 2024
Closest in time.
Autom3l: An automated multimodal machine learning framework with large language models
Luo, D., Feng, C., Nong, Y., and Shen, Y. (2024) · 2024
Closest in time.
Gpt-4o system card
OpenAI (2024) · 2024
Closest in time.
Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face
Shen, Y., Song, K., Tan, X., Li, D., Lu, W., and Zhuang, Y. (2024) · 2024
Closest in time.
Lambda: A large model based data agent
Sun, M., Han, R., Jiang, B., Qi, H., Sun, D., Yuan, Y., and Huang, J. (2024) · 2024
Closest in time.
Automl-agent: A multi-agent llm framework for full-pipeline automl
Trirat, P., Jeong, W., and Hwang, S. J. (2024) · 2024
Closest in time.
Xie, L., Zheng, C., Xia, H., Qu, H., and Zhu-Tian, C. (2024) · 2024
Closest in time.
Matplotagent: Method and evaluation for llm-based agentic scientific data visualization
Yang, Z., Zhou, Z., Wang, S., Cong, X., Han, X., Yan, Y., Liu, Z., Tan, Z., Liu, P., Yu, D., et al. (2024) · 2024
Closest in time.
Tree of thoughts: Deliberate problem solving with large language models
Yao, S., Yu, D., Zhao, J., Shafran, I., Griffiths, T., Cao, Y., and Narasimhan, K. (2024) · 2024
Closest in time.
Ufo: A ui-focused agent for windows os interaction
Zhang, C., Li, L., He, S., Zhang, X., Qiao, B., Qin, S., Ma, M., Kang, Y., Lin, Q., Rajmohan, S., et al. (2024) · 2024
Closest in time.
Data science agent in colab with gemini
Google (2025) · 2025
Closest in time.
Julius ai
Julius (2025) · 2025
Closest in time.
Score: Story coherence and retrieval enhancement for ai narratives
Yi, Q., He, Y., Wang, J., Song, X., Qian, S., Yuan, X., Sun, L., Xin, Y., Tang, J., Li, K., et al. (2025) · 2025
Closest in time.