Fetching the paper…
Reading the bibliography…
AI research agents are demonstrating great potential to accelerate scientific progress by automating the design, implementation, and training of machine learning models.
Scientific discovery, 1987
Pat Langley, Herbert A Simon, Gary L Bradshaw, and Jan M Zytkow · 1987
Earlier work this paper cites.
Branching rules for satisfiability
John N Hooker and V Vinay · 1995
Earlier work this paper cites.
Experimental results on the crossover point in random 3-sat
James M Crawford and Larry D Auton · 1996
Earlier work this paper cites.
Bandit based monte-carlo planning
Levente Kocsis and Csaba Szepesvári · 2006
Earlier work this paper cites.
Random search for hyper-parameter optimization
James Bergstra and Yoshua Bengio · 2012
Earlier work this paper cites.
Auto-weka: combined selection and hyperparameter optimization of classification algorithms
Chris Thornton, Frank Hutter, Holger H. Hoos, and Kevin Leyton-Brown · 2013
Earlier work this paper cites.
Exploration and exploitation in evolutionary algorithms: A survey
Matej Črepinšek, Shih-Hsi Liu, and Marjan Mernik · 2013
Earlier work this paper cites.
Evolutionary computation: a unified approach
Kenneth De Jong · 2017
Earlier work this paper cites.
Simple and efficient architecture search for convolutional neural networks
Thomas Elsken, Jan-Hendrik Metzen, and Frank Hutter · 2017
Earlier work this paper cites.
Neural architecture search with reinforcement learning
Barret Zoph and Quoc Le · 2017
Earlier work this paper cites.
Hyperband: A novel bandit-based approach to hyperparameter optimization
Lisha Li, Kevin Jamieson, Giulia DeSalvo, Afshin Rostamizadeh, and Ameet Talwalkar · 2018
Earlier work this paper cites.
DARTS: Differentiable architecture search
Hanxiao Liu, Karen Simonyan, and Yiming Yang · 2019
Earlier work this paper cites.
TPOT: A Tree-Based Pipeline Optimization Tool for Automating Machine Learning , pages 151–160
Randal S. Olson and Jason H. Moore · 2019
Earlier work this paper cites.
Regularized evolution for image classifier architecture search
Esteban Real, Alok Aggarwal, Yanping Huang, and Quoc V. Le · 2019
Earlier work this paper cites.
hpcng/singularity: Singularity 3.7.3, April 2021
Gregory M. Kurtzer, cclerget, Michael Bauer, Ian Kaneshiro, David Trudgian, and David Godlove · 2021
Earlier work this paper cites.
Auto-sklearn 2.0: hands-free automl via meta-learning
Matthias Feurer, Katharina Eggensperger, Stefan Falkner, Marius Lindauer, and Frank Hutter · 2022
Earlier work this paper cites.
Chain of thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V Le, and Denny Zhou · 2022
Earlier work this paper cites.
Autonomous chemical research with large language models
Daniil A. Boiko, Robert MacKnight, Ben Kline, and Gabe Gomes · 2023
Earlier work this paper cites.
Reasoning with language model is planning with world model, December 2023
Shibo Hao, Yi Gu, Haodi Ma, Joshua Hong, Zhen Wang, Daisy Wang, and Zhiting Hu · 2023
Cited alongside, same era.
Mlagentbench: Evaluating language agents on machine learning experimentation
Qian Huang, Jian Vora, Percy Liang, and Jure Leskovec · 2023
Cited alongside, same era.
Self-consistency improves chain of thought reasoning in language models
Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V Le, Ed H. Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou · 2023
Cited alongside, same era.
The llama 3 herd of models, 2024
Llama Team et. al · 2024
Cited alongside, same era.
Large language models orchestrating structured reasoning achieve kaggle grandmaster level, 2024
Antoine Grosnit, Alexandre Maraval, James Doran, Giuseppe Paolo, Albert Thomas, Refinath Shahul Hameed Nabeezath Beevi, Jonas Gonzalez, Khyati Khandelwal, Ignacio Iacobacci, Abdelhakim Benechehab, Hamza Cherkaoui, Youssef Attia El-Hili, Kun Shao, Jianye Hao, Jun Yao, Balazs Kegl, Haitham Bou-Ammar, and Jun Wang · 2024
Swe-agent: Agent-computer interfaces enable automated software engineering
John Yang, Carlos Jimenez, Alexander Wettig, Kilian Lieret, Shunyu Yao, Karthik Narasimhan, and Ofir Press · 2024
Later among the works it cites.
SWE-search: Enhancing software agents with monte carlo tree search and iterative refinement
Antonis Antoniades, Albert Örwall, Kexun Zhang, Yuxi Xie, Anirudh Goyal, and William Yang Wang · 2025
Closest in time.
MLE-bench: Evaluating machine learning agents on machine learning engineering
Jun Shern Chan, Neil Chowdhury, Oliver Jaffe, James Aung, Dane Sherburn, Evan Mays, Giulio Starace, Kevin Liu, Leon Maksin, Tejal Patwardhan, Aleksander Madry, and Lilian Weng · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025
DeepSeek-AI · 2025
Closest in time.
Gemini: A family of highly capable multimodal models, 2025
Gemini Team et. al · 2025
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Saycanpay: Heuristic planning with large language models using learnable domain knowledge
Rishi Hazra, Pedro Zuidberg Dos Martires, and Luc De Raedt · 2024
Cited alongside, same era.
Megascale: scaling large language model training to more than 10,000 gpus
Ziheng Jiang, Haibin Lin, Yinmin Zhong, Qi Huang, Yangrui Chen, Zhi Zhang, Yanghua Peng, Xiang Li, Cong Xie, Shibiao Nong, Yulu Jia, Sun He, Hongmin Chen, Zhihao Bai, Qi Hou, Shipeng Yan, Ding Zhou, Yiyao Sheng, Zhuo Jiang, Haohan Xu, Haoran Wei, Zhang Zhang, Pengfei Nie, Leqi Zou, Sida Zhao, Liang Xiang, Zherui Liu, Zhe Li, Xiaoying Jia, Jianxi Ye, Xin Jin, and Xin Liu · 2024
Cited alongside, same era.
Large language model-based agents for software engineering: A survey
Junwei Liu, Kaixin Wang, Yixuan Chen, Xin Peng, Zhenpeng Chen, Lingming Zhang, and Yiling Lou · 2024
Cited alongside, same era.
Discovering preference optimization algorithms with and for large language models
Chris Lu, Samuel Holt, Claudio Fanconi, Alex James Chan, Jakob Nicolaus Foerster, Mihaela van der Schaar, and Robert Tjarko Lange · 2024
Cited alongside, same era.
Gpt-4o system card, 2024
OpenAI · 2024
Cited alongside, same era.
Introducing o3 and o4 mini, 2024
OpenAI · 2024
Cited alongside, same era.
Tool learning with foundation models
Yujia Qin, Shengding Hu, Yankai Lin, Weize Chen, Ning Ding, Ganqu Cui, Zheni Zeng, Xuanhe Zhou, Yufei Huang, Chaojun Xiao, Chi Han, Yi Ren Fung, Yusheng Su, Huadong Wang, Cheng Qian, Runchu Tian, Kunlun Zhu, Shihao Liang, Xingyu Shen, Bokai Xu, Zhen Zhang, Yining Ye, Bowen Li, Ziwei Tang, Jing Yi, Yuzhang Zhu, Zhenning Dai, Lan Yan, Xin Cong, Yaxi Lu, Weilin Zhao, Yuxiang Huang, Junxi Yan, Xu Han, Xian Sun, Dahai Li, Jason Phang, Cheng Yang, Tongshuang Wu, Heng Ji, Guoliang Li, Zhiyuan Liu, and Maosong Sun · 2024
Cited alongside, same era.
Closest in time.
OMNI-EPIC: Open-endedness via models of human notions of interestingness with environments programmed in code
Maxence Faldor, Jenny Zhang, Antoine Cully, and Jeff Clune · 2025
Closest in time.
Rlef: Grounding code llms in execution feedback with reinforcement learning, 2025
Jonas Gehring, Kunhao Zheng, Jade Copet, Vegard Mella, Quentin Carbonneaux, Taco Cohen, and Gabriel Synnaeve · 2025
Closest in time.
REvolve: Reward evolution with large language models using human feedback
Rishi Hazra, Alkis Sygkounas, Andreas Persson, Amy Loutfi, and Pedro Zuidberg Dos Martires · 2025
Closest in time.
Cursor: The ai code editor, 2025
Anysphere Inc · 2025
Closest in time.
Inspect AI: Framework for Large Language Model Evaluations, May 2025
UK AI Security Institute · 2025
Closest in time.
AIDE: AI-Driven Exploration in the Space of Code
Zhengyao Jiang, Dominik Schmidt, Dhruv Srikanth, Dixing Xu, Ian Kaplan, Deniss Jacenko, and Yuxiang Wu · 2025
Closest in time.
Revisiting Reliability in Large-Scale Machine Learning Research Clusters
Apostolos Kokolis, Michael Kuchnik, John Hoffman, Adithya Kumar, Parth Malani, Faye Ma, Zachary DeVito, Shubho Sengupta, Kalyan Saladi, and Carole-Jean Wu · 2025
Closest in time.
MLGym: A New Framework and Benchmark for Advancing AI Research Agents, 2025
Deepak Nathani, Lovish Madaan, Nicholas Roberts, Nikolay Bashlykov, Ajay Menon, Vincent Moens, Amar Budhiraja, Despoina Magka, Vladislav Vorotilov, Gaurav Chaurasia, Dieuwke Hupkes, Ricardo Silveira Cabral, Tatiana Shavrina, Jakob Foerster, Yoram Bachrach, William Yang Wang, and Roberta Raileanu · 2025
Closest in time.
Gpt-5 system card, 2025
OpenAI · 2025
Closest in time.
Livebench: A challenging, contamination-free LLM benchmark
Colin White, Samuel Dooley, Manley Roberts, Arka Pal, Benjamin Feuer, Siddhartha Jain, Ravid Shwartz-Ziv, Neel Jain, Khalid Saifullah, Sreemanti Dey, Shubh-Agrawal, Sandeep Singh Sandha, Siddartha Venkat Naidu, Chinmay Hegde, Yann LeCun, Tom Goldstein, Willie Neiswanger, and Micah Goldblum · 2025
Closest in time.
The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search, 2025
Yutaro Yamada, Robert Tjarko Lange, Cong Lu, Shengran Hu, Chris Lu, Jakob Foerster, Jeff Clune, and David Ha · 2025
Closest in time.
Xu Yang, Xiao Yang, Shikai Fang, Bowen Xian, Yuante Li, Jian Wang, Minrui Xu, Haoran Pan, Xinpeng Hong, Weiqing Liu, Yelong Shen, Weizhu Chen, and Jiang Bian · 2025
Closest in time.