Fetching the paper…
Reading the bibliography…
Agents based on large language models (LLMs) for machine learning engineering (MLE) can automatically implement ML models via code generation.
An introduction to case-based reasoning
J. L. Kolodner · 1992
Earlier work this paper cites.
Case-based reasoning: A review
I. Watson and F. Marir · 1994
Earlier work this paper cites.
Generalized and heuristic-free feature construction for improved accuracy
W. Fan, E. Zhong, J. Peng, O. Verscheure, K. Zhang, J. Ren, R. Yan, and Q. Yang · 2010
Earlier work this paper cites.
Scikit-learn: Machine learning in python
F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, et al · 2011
Earlier work this paper cites.
Deep feature synthesis: Towards automating data science endeavors
J. M. Kanter and K. Veeramachaneni · 2015
Earlier work this paper cites.
Xgboost: A scalable tree boosting system
T. Chen and C. Guestrin · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Tpot: A tree-based pipeline optimization tool for automating machine learning
R. S. Olson and J. H. Moore · 2016
Earlier work this paper cites.
Auto-weka 2.0: Automatic model selection and hyperparameter optimization in weka
L. Kotthoff, C. Thornton, H. H. Hoos, F. Hutter, and K. Leyton-Brown · 2017
Earlier work this paper cites.
Neural architecture search with reinforcement learning
B. Zoph and Q. V. Le · 2017
Earlier work this paper cites.
Efficient neural architecture search via parameters sharing
H. Pham, M. Guan, B. Zoph, Q. Le, and J. Dean · 2018
Earlier work this paper cites.
Catboost: unbiased boosting with categorical features
L. Prokhorenkova, G. Gusev, A. Vorobev, A. V. Dorogush, and A. Gulin · 2018
Earlier work this paper cites.
Neural architecture search: A survey
T. Elsken, J. H. Metzen, and F. Hutter · 2019
Earlier work this paper cites.
Brief review of image denoising techniques
L. Fan, F. Zhang, H. Fan, and C. Zhang · 2019
Earlier work this paper cites.
The autofeat python library for automated feature engineering and selection
F. Horn, R. Pack, and M. Rieger · 2019
Earlier work this paper cites.
Auto-keras: An efficient neural architecture search system
H. Jin, Q. Song, and X. Hu · 2019
Earlier work this paper cites.
Regularized evolution for image classifier architecture search
E. Real, A. Aggarwal, Y. Huang, and Q. V. Le · 2019
Cited alongside, same era.
Efficientnet: Rethinking model scaling for convolutional neural networks
M. Tan and Q. Le · 2019
Cited alongside, same era.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Cited alongside, same era.
Autogluon-tabular: Robust and accurate automl for structured data
N. Erickson, J. Mueller, A. Shirkov, H. Zhang, P. Larroy, M. Li, and A. Smola · 2020
Cited alongside, same era.
H2O AutoML: Scalable automatic machine learning
E. LeDell and S. Poirier · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Better by default: Strong pre-tuned mlps and boosted trees on tabular data
D. Holzmüller, L. Grinsztajn, and I. Steinwart · 2024
Later among the works it cites.
Data interpreter: An llm agent for data science
S. Hong, Y. Lin, B. Liu, B. Liu, B. Wu, C. Zhang, C. Wei, D. Li, J. Chen, J. Zhang, et al · 2024
Later among the works it cites.
Infiagent-dabench: Evaluating agents on data analysis tasks
X. Hu, Z. Zhao, S. Wei, Z. Chai, Q. Ma, G. Wang, X. Wang, J. Su, J. Xu, M. Zhu, et al · 2024
Later among the works it cites.
Swe-bench: Can language models resolve real-world github issues?
C. E. Jimenez, J. Yang, A. Wettig, S. Yao, K. Pei, O. Press, and K. Narasimhan · 2024
Later among the works it cites.
Autokaggle: A multi-agent framework for autonomous data science competitions
Z. Li, Q. Zang, D. Ma, J. Guo, T. Zheng, M. Liu, X. Niu, Y. Wang, J. Yang, J. Liu, et al · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby · 2021
Cited alongside, same era.
Auto-sklearn 2.0: Hands-free automl via meta-learning
M. Feurer, K. Eggensperger, S. Falkner, M. Lindauer, and F. Hutter · 2022
Cited alongside, same era.
Competition-level code generation with alphacode
Y. Li, D. Choi, J. Chung, N. Kushman, J. Schrittwieser, R. Leblond, T. Eccles, J. Keeling, F. Gimeno, A. Dal Lago, et al · 2022
Cited alongside, same era.
Large language models for automated data science: Introducing caafe for context-aware automated feature engineering
N. Hollmann, S. Müller, and F. Hutter · 2023
Cited alongside, same era.
Learning a data-driven policy network for pre-training automated feature engineering
L. Li, H. Wang, L. Zha, Q. Huang, S. Wu, G. Chen, and J. Zhao · 2023
Cited alongside, same era.
Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face
Y. Shen, K. Song, X. Tan, D. Li, W. Lu, and Y. Zhuang · 2023
Cited alongside, same era.
Llama: Open and efficient foundation language models
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, et al · 2023
Cited alongside, same era.
Later among the works it cites.
Optimized feature generation for tabular data via llms with decision tree reasoning
J. Nam, K. Kim, S. Oh, J. Tack, J. Kim, and J. Shin · 2024
Later among the works it cites.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
G. Team, P. Georgiev, V. I. Lei, R. Burnell, L. Bai, A. Gulati, G. Tanzer, D. Vincent, Z. Pan, S. Wang, et al · 2024
Later among the works it cites.
Openhands: An open platform for ai software developers as generalist agents
X. Wang, B. Li, Y. Song, F. F. Xu, X. Tang, M. Zhuge, J. Pan, Y. Song, B. Li, J. Singh, et al · 2024
Later among the works it cites.
Mle-bench: Evaluating machine learning agents on machine learning engineering
J. S. Chan, N. Chowdhury, O. Jaffe, J. Aung, D. Sherburn, E. Mays, G. Starace, K. Liu, L. Maksin, T. Patwardhan, et al · 2025
Closest in time.
Accurate predictions on small data with a tabular foundation model
N. Hollmann, S. Müller, L. Purucker, A. Krishnakumar, M. Körfer, S. B. Hoo, R. T. Schirrmeister, and F. Hutter · 2025
Closest in time.
Evaluation of best-of-n sampling strategies for language model alignment
Y. Ichihara, Y. Jinnai, T. Morimura, K. Abe, K. Ariu, M. Sakamoto, and E. Uchibe · 2025
Closest in time.
Livecodebench: Holistic and contamination free evaluation of large language models for code
N. Jain, K. Han, A. Gu, W.-D. Li, F. Yan, T. Zhang, S. Wang, A. Solar-Lezama, K. Sen, and I. Stoica · 2025
Closest in time.
Aide: Ai-driven exploration in the space of code
Z. Jiang, D. Schmidt, D. Srikanth, D. Xu, I. Kaplan, D. Jacenko, and Y. Wu · 2025
Closest in time.
Dsbench: How far are data science agents to becoming data science experts?
L. Jing, Z. Huang, X. Wang, W. Yao, W. Yu, K. Ma, H. Zhang, X. Du, and D. Yu · 2025
Closest in time.
Agent laboratory: Using llm agents as research assistants
S. Schmidgall, Y. Su, Z. Wang, X. Sun, J. Wu, X. Yu, J. Liu, Z. Liu, and E. Barsoum · 2025
Closest in time.
Datawiseagent: A notebook-centric llm agent framework for automated data science
Z. You, Y. Zhang, D. Xu, Y. Lou, Y. Yan, W. Wang, H. Zhang, and Y. Huang · 2025
Closest in time.