Fetching the paper…
Reading the bibliography…
The adoption of Large Language Models (LLMs) for code generation in data science offers substantial potential for enhancing tasks such as data manipulation, statistical analysis, and visualization.
C. Wohlin, P. Runeson, M. Höst, M. C. Ohlsson, B. Regnell, A. Wesslén et al. , Experimentation in software engineering . Springer, 2012, vol. 236
2012
Earlier work this paper cites.
2021
Earlier work this paper cites.
N. Nguyen and S. Nadi, “An empirical evaluation of github copilot’s code suggestions,” in Proceedings of the 19th International Conference on Mining Software Repositories , 2022, pp. 1–5
2022
Earlier work this paper cites.
A. Halevy, Y. Choi, A. Floratou, M. J. Franklin, N. Noy, and H. Wang, “Will llms reshape, supercharge, or kill data science?(vldb 2023 panel),” Proceedings of the VLDB Endowment , vol. 16, no. 12, pp. 4114–4115, 2023
2023
Earlier work this paper cites.
N. Nascimento, C. Tavares, P. Alencar, and D. Cowan, “Gpt in data science: A practical exploration of model selection,” in 2023 IEEE International Conference on Big Data (BigData) . IEEE, 2023, pp. 4325–4334
2023
Earlier work this paper cites.
Y. Lai, C. Li, Y. Wang, T. Zhang, R. Zhong, L. Zettlemoyer, W.-t. Yih, D. Fried, S. Wang, and T. Yu, “Ds-1000: A natural and reliable benchmark for data science code generation,” in International Conference on Machine Learning . PMLR, 2023, pp. 18 319–18 345
2023
Earlier work this paper cites.
N. Nathalia, A. Paulo, and C. Donald, “Artificial intelligence vs. software engineers: An empirical study on performance and efficiency using chatgpt,” in Proceedings of the 33rd Annual International Conference on Computer Science and Software Engineering , 2023, pp. 24–33
2023
Earlier work this paper cites.
C. Troy, S. Sturley, J. M. Alcaraz-Calero, and Q. Wang, “Enabling generative ai to produce sql statements: A framework for the auto-generation of knowledge based on ebnf context-free grammars,” IEEE Access , vol. 11, pp. 123 543–123 564, 2023
2023
Cited alongside, same era.
J. White, S. Hays, Q. Fu, J. Spencer-Smith, and D. C. Schmidt, “Chatgpt prompt patterns for improving code quality, refactoring, requirements elicitation, and software design,” 2023
2023
Cited alongside, same era.
N. Nascimento, P. Alencar, and D. Cowan, “Gpt-in-the-loop: Supporting adaptation in multiagent systems,” in 2023 IEEE International Conference on Big Data (BigData) . IEEE, 2023, pp. 4674–4683
2023
Cited alongside, same era.
Ritchie Vink, “Polars: Blazingly fast dataframes in rust, python, node.js, r, and sql,” 2023. [Online]. Available: https://github.com/pola-rs/polars
2023
Cited alongside, same era.
M. A. Kuhail, S. S. Mathew, A. Khalil, J. Berengueres, and S. J. H. Shah, ““will i be replaced?” assessing chatgpt’s effect on software development and programmer perceptions of ai tools,” Science of Computer Programming , vol. 235, p. 103111, 2024
2024
Closest in time.
T. Coignion, C. Quinton, and R. Rouvoy, “A performance study of llm-generated code on leetcode,” in Proceedings of the 28th International Conference on Evaluation and Assessment in Software Engineering , 2024, pp. 79–89
2024
Closest in time.
B. Grewal, W. Lu, S. Nadi, and C.-P. Bezemer, “Analyzing developer use of chatgpt generated code in open source github projects,” in 2024 IEEE/ACM 21st International Conference on Mining Software Repositories (MSR) . IEEE, 2024, pp. 157–161
2024
Closest in time.
X. Gu, M. Chen, Y. Lin, Y. Hu, H. Zhang, C. Wan, Z. Wei, Y. Xu, and J. Wang, “On the effectiveness of large language models in domain-specific code generation,” ACM Transactions on Software Engineering and Methodology , 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
G. V. Aher, R. I. Arriaga, and A. T. Kalai, “Using large language models to simulate multiple humans and replicate human subject studies,” in International Conference on Machine Learning . PMLR, 2023, pp. 337–371
2023
Cited alongside, same era.
J. Li, B. Hui, G. Qu, J. Yang, B. Li, B. Li, B. Wang, B. Qin, R. Geng, N. Huo et al. , “Can llm already serve as a database interface? a big bench for large-scale database grounded text-to-sqls,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Cited alongside, same era.
M. Kazemitabaar, J. Williams, I. Drosos, T. Grossman, A. Z. Henley, C. Negreanu, and A. Sarkar, “Improving steering and verification in ai-assisted data analysis with interactive task decomposition,” in Proceedings of the 37th Annual ACM Symposium on User Interface Software and Technology , 2024, pp. 1–19
2024
Cited alongside, same era.
2024
Closest in time.
StrataScratch, “Master coding for data science,” https://www.stratascratch.com/, n.d., accessed: 2024-11-01
2024
Closest in time.
M. Malekpour, N. Shaheen, F. Khomh, and A. Mhedhbi, “Towards optimizing sql generation via llm routing,” in NeurIPS 2024 Third Table Representation Learning Workshop
2024
Closest in time.
S. A. Boominathan, S. S. Chintakunta, N. Nascimento, and E. Guimaraes, “LLM4DS-Benchmark: A Dataset for Assessing LLM Performance in Data Science Coding Tasks,” Nov. 2024. [Online]. Available: https://doi.org/10.5281/zenodo.14064111
2024
Closest in time.