Fetching the paper…
Reading the bibliography…
We introduce Meta MLGym and MLGym-Bench, a new framework and benchmark for evaluating and developing LLM agents on AI research tasks.
Kenny Young and Tian Tian · 1903
Earlier work this paper cites.
Some experimental games
Merrill M Flood · 1958
Earlier work this paper cites.
The complexity of theorem-proving procedures
Stephen A Cook · 1971
Earlier work this paper cites.
Effective choice in the prisoner’s dilemma
Robert Axelrod · 1980
Earlier work this paper cites.
Artifical intelligence and nuclear power
Artificial Intelligence Task Team · 1985
Earlier work this paper cites.
Scientific discovery: Computational explorations of the creative processes
P Langley · 1987
Earlier work this paper cites.
Potential applications of artificial intelligence to the field of software engineering
M Emrich, A Agarwal, B Jairam, N Murthy, and OAK RIDGE NATIONAL LAB TN · 1988
Earlier work this paper cites.
Communication in the battle of the sexes game: some experimental results
Russell Cooper, Douglas V DeJong, Robert Forsythe, and Thomas W Ross · 1989
Earlier work this paper cites.
Game theory
Drew Fudenberg and Jean Tirole · 1991
Earlier work this paper cites.
Oak Ridge National Laboratory: the first fifty years
Leland Johnson and Daniel Schaffer · 1994
Earlier work this paper cites.
Benchmarking optimization software with performance profiles
Elizabeth D. Dolan and Jorge J. Moré · 2002
Earlier work this paper cites.
Thomas Miconi, Aditya Rawal, Jeff Clune, and Kenneth O. Stanley · 2002
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
The logic of scientific discovery
Karl Popper · 2005
Earlier work this paper cites.
The colonel blotto game
Brian Roberson · 2006
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Games and decisions: Introduction and critical survey
R Duncan Luce and Howard Raiffa · 2012
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Earlier work this paper cites.
House prices - advanced regression techniques
Kaggle · 2016
Earlier work this paper cites.
Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms, 2017
Han Xiao, Kashif Rasul, and Roland Vollgraf · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin · 2018
Earlier work this paper cites.
Artificial intelligence in drug design
Gerhard Hessler and Karl-Heinz Baringhaus · 2018
Earlier work this paper cites.
Learning a SAT solver from single-bit supervision
Daniel Selsam, Matthew Lamm, Benedikt Bünz, Percy Liang, Leonardo de Moura, and David L. Dill · 2018
Earlier work this paper cites.
The multi-genre nli corpus
Adina Williams, Nikita Nangia, and Samuel R Bowman · 2018
Earlier work this paper cites.
Pitfalls and best practices in algorithm configuration
Katharina Eggensperger, Marius Lindauer, and Frank Hutter · 2019
Earlier work this paper cites.
Neural architecture search: A survey
Thomas Elsken, Jan Hendrik Metzen, and Frank Hutter · 2019
Earlier work this paper cites.
Best practices for scientific research on neural architecture search
Marius Lindauer and Frank Hutter · 2020
Earlier work this paper cites.
Rethinking drug design in the artificial intelligence era
Petra Schneider, W Patrick Walters, Alleyn T Plowright, Norman Sieroka, Jennifer Listgarten, Robert A Goodnow Jr, Jasmin Fisher, Johanna M Jansen, José S Duca, Thomas S Rush, et al · 2020
Earlier work this paper cites.
Artificial intelligence and machine learning in design of mechanical materials
Kai Guo, Zhenze Yang, Chi-Hua Yu, and Markus J Buehler · 2021
Cited alongside, same era.
gymnax: A JAX-based reinforcement learning environment library, 2022
Robert Tjarko Lange · 2022
Cited alongside, same era.
Webgpt: Browser-assisted question-answering with human feedback, 2022
Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, Xu Jiang, Karl Cobbe, Tyna Eloundou, Gretchen Krueger, Kevin Button, Matthew Knight, Benjamin Chess, and John Schulman · 2022
Cited alongside, same era.
Automl decathlon: Diverse tasks, modern methods, and efficiency at scale
Nicholas Roberts, Samuel Guo, Cong Xu, Ameet Talwalkar, David Lander, Lvfang Tao, Linhang Cai, Shuaicheng Niu, Jianyu Heng, Hongyang Qin, Minwen Deng, Johannes Hog, Alexander Pfefferle, Sushil Ammanaghatta Shivakumar, Arjun Krishnakumar, Yubo Wang, Rhea Sukthanker, Frank Hutter, Euxhen Hasanaj, Tien-Dung Le, Mikhail Khodak, Yuriy Nevmyvaka, Kashif Rasul, Frederic Sala, Anderson Schneider, Junhong Shen, and Evan Sparks · 2022
Cited alongside, same era.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al · 2024
Later among the works it cites.
Magentic-one: A generalist multi-agent system for solving complex tasks, 2024
Adam Fourney, Gagan Bansal, Hussein Mozannar, Cheng Tan, Eduardo Salinas, Erkang, Zhu, Friederike Niedtner, Grace Proebsting, Griffin Bassman, Jack Gerrits, Jacob Alber, Peter Chang, Ricky Loynd, Robert West, Victor Dibia, Ahmed Awadallah, Ece Kamar, Rafah Hosn, and Saleema Amershi · 2024
Later among the works it cites.
Large language models orchestrating structured reasoning achieve kaggle grandmaster level, 2024
Antoine Grosnit, Alexandre Maraval, James Doran, Giuseppe Paolo, Albert Thomas, Refinath Shahul Hameed Nabeezath Beevi, Jonas Gonzalez, Khyati Khandelwal, Ignacio Iacobacci, Abdelhakim Benechehab, Hamza Cherkaoui, Youssef Attia El-Hili, Kun Shao, Jianye Hao, Jun Yao, Balazs Kegl, Haitham Bou-Ammar, and Jun Wang · 2024
Later among the works it cites.
Mamba: Linear-time sequence modeling with selective state spaces, 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Renbo Tu, Nicholas Roberts, Mikhail Khodak, Junhong Shen, Frederic Sala, and Ameet Talwalkar · 2022
Cited alongside, same era.
The clrs algorithmic reasoning benchmark
Petar Veličković, Adrià Puigdomènech Badia, David Budden, Razvan Pascanu, Andrea Banino, Misha Dashevskiy, Raia Hadsell, and Charles Blundell · 2022
Cited alongside, same era.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Cited alongside, same era.
Benchmarking neural network training algorithms, 06 2023
George Dahl, Frank Schneider, Zachary Nado, Naman Agarwal, Chandramouli Shama Sastry, Philipp Hennig, Sourabh Medapati, Runa Eschenhagen, Priya Kasimbeg, Daniel Suo, Juhan Bae, Justin Gilmer, Abel Peirson, Bilal Khan, Rohan Anil, Mike Rabbat, Shankar Krishnan, Daniel Snider, Ehsan Amid, and Peter Mattson · 2023
Cited alongside, same era.
Mind2web: Towards a generalist agent for the web, 2023
Xiang Deng, Yu Gu, Boyuan Zheng, Shijie Chen, Samuel Stevens, Boshi Wang, Huan Sun, and Yu Su · 2023
Cited alongside, same era.
Swe-bench: Can language models resolve real-world github issues?
Carlos E Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan · 2023
Cited alongside, same era.
Challenges and applications of large language models
Jean Kaddour, Joshua Harris, Maximilian Mozes, Herbie Bradley, Roberta Raileanu, and Robert McHardy · 2023
Cited alongside, same era.
Agentsims: An open-source sandbox for large language model evaluation, 2023
Jiaju Lin, Haoran Zhao, Aochi Zhang, Yiting Wu, Huqiuyue Ping, and Qin Chen · 2023
Cited alongside, same era.
Albert Gu and Tri Dao · 2024
Later among the works it cites.
Data Interpreter: An LLM Agent For Data Science, March 2024
Sirui Hong, Yizhang Lin, Bang Liu, Bangbang Liu, Binhao Wu, Danyang Li, Jiaqi Chen, Jiayi Zhang, Jinlin Wang, Li Zhang, Lingyao Zhang, Min Yang, Mingchen Zhuge, Taicheng Guo, Tuo Zhou, Wei Tao, Wenyi Wang, Xiangru Tang, Xiangtao Lu, Xiawu Zheng, Xinbing Liang, Yaying Fei, Yuheng Cheng, Zongze Xu, and Chenglin Wu · 2024
Later among the works it cites.
MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation, April 2024
Qian Huang, Jian Vora, Percy Liang, and Jure Leskovec · 2024
Later among the works it cites.
Peter Jansen, Marc-Alexandre Côté, Tushar Khot, Erin Bransom, Bhavana Dalvi Mishra, Bodhisattwa Prasad Majumder, Oyvind Tafjord, and Peter Clark · 2024
Later among the works it cites.
modded-nanogpt: Speedrunning the nanogpt baseline, 2024
Keller Jordan, Jeremy Bernstein, Brendan Rappazzo, @fernbear.bsky.social, Boza Vlado, You Jiacheng, Franz Cesista, Braden Koszarsky, and @Grad62304977 · 2024
Later among the works it cites.
Sayash Kapoor, Benedikt Stroebl, Zachary S. Siegel, Nitya Nadgir, and Arvind Narayanan · 2024
Later among the works it cites.
Spider 2.0: Evaluating language models on real-world enterprise text-to-sql workflows, 2024
Fangyu Lei, Jixuan Chen, Yuxiao Ye, Ruisheng Cao, Dongchan Shin, Hongjin Su, Zhaoqing Suo, Hongcheng Gao, Wenjing Hu, Pengcheng Yin, Victor Zhong, Caiming Xiong, Ruoxi Sun, Qian Liu, Sida Wang, and Tao Yu · 2024
Later among the works it cites.
Autokaggle: A multi-agent framework for autonomous data science competitions, 2024
Ziming Li, Qianbo Zang, David Ma, Jiawei Guo, Tuney Zheng, Minghao Liu, Xinyao Niu, Yue Wang, Jian Yang, Jiaheng Liu, Wanjun Zhong, Wangchunshu Zhou, Wenhao Huang, and Ge Zhang · 2024
Later among the works it cites.
The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery, August 2024
Chris Lu, Cong Lu, Robert Tjarko Lange, Jakob Foerster, Jeff Clune, and David Ha · 2024
Later among the works it cites.
Sciagent: Tool-augmented language models for scientific reasoning, 2024
Yubo Ma, Zhibin Gou, Junheng Hao, Ruochen Xu, Shuohang Wang, Liangming Pan, Yujiu Yang, Yixin Cao, Aixin Sun, Hany Awadalla, and Weizhu Chen · 2024
Later among the works it cites.
Evaluating frontier ai r&d capabilities of language model agents against human experts, 11 2024
METR · 2024
Later among the works it cites.
Llmatic: neural architecture search via large language models and quality diversity optimization
Muhammad Umair Nasir, Sam Earle, Julian Togelius, Steven James, and Christopher Cleghorn · 2024
Later among the works it cites.
Balrog: Benchmarking agentic llm and vlm reasoning on games
Davide Paglieri, Bartłomiej Cupiał, Samuel Coward, Ulyana Piterbarg, Maciej Wolczyk, Akbir Khan, Eduardo Pignatelli, Łukasz Kuciński, Lerrel Pinto, Rob Fergus, et al · 2024
Later among the works it cites.
The fineweb datasets: Decanting the web for the finest text data at scale, 2024
Guilherme Penedo, Hynek Kydlíček, Loubna Ben allal, Anton Lozhkov, Margaret Mitchell, Colin Raffel, Leandro Von Werra, and Thomas Wolf · 2024
Later among the works it cites.
Can llms generate novel research ideas? a large-scale human study with 100+ nlp researchers, 2024
Chenglei Si, Diyi Yang, and Tatsunori Hashimoto · 2024
Later among the works it cites.
Xiangru Tang, Yuliang Liu, Zefan Cai, Yanjun Shao, Junjie Lu, Yichi Zhang, Zexuan Deng, Helan Hu, Kaikai An, Ruijun Huang, Shuzheng Si, Sheng Chen, Haozhe Zhao, Liang Chen, Yan Wang, Tianyu Liu, Zhiwei Jiang, Baobao Chang, Yin Fang, Yujia Qin, Wangchunshu Zhou, Yilun Zhao, Arman Cohan, and Mark Gerstein · 2024
Later among the works it cites.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Gemini Team, Petko Georgiev, Ving Ian Lei, Ryan Burnell, Libin Bai, Anmol Gulati, Garrett Tanzer, Damien Vincent, Zhufeng Pan, Shibo Wang, et al · 2024
Later among the works it cites.
Gymnasium: A standard interface for reinforcement learning environments, 2024
Mark Towers, Ariel Kwiatkowski, Jordan Terry, John U. Balis, Gianluca De Cola, Tristan Deleu, Manuel Goulão, Andreas Kallinteris, Markus Krimmel, Arjun KG, Rodrigo Perez-Vicente, Andrea Pierré, Sander Schulhoff, Jun Jet Tai, Hannah Tan, and Omar G. Younis · 2024
Later among the works it cites.
Harsh Trivedi, Tushar Khot, Mareike Hartmann, Ruskin Manku, Vinty Dong, Edward Li, Shashank Gupta, Ashish Sabharwal, and Niranjan Balasubramanian · 2024
Later among the works it cites.
Os-copilot: Towards generalist computer agents with self-improvement
Zhiyong Wu, Chengcheng Han, Zichen Ding, Zhenmin Weng, Zhoumianze Liu, Shunyu Yao, Tao Yu, and Lingpeng Kong · 2024
Later among the works it cites.
Agentless: Demystifying LLM-based Software Engineering Agents, July 2024
Chunqiu Steven Xia, Yinlin Deng, Soren Dunn, and Lingming Zhang · 2024
Later among the works it cites.
SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering, May 2024
John Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret, Shunyu Yao, Karthik Narasimhan, and Ofir Press · 2024
Later among the works it cites.
AssistantBench: Can Web Agents Solve Realistic and Time-Consuming Tasks?, July 2024
Ori Yoran, Samuel Joseph Amouyal, Chaitanya Malaviya, Ben Bogin, Ofir Press, and Jonathan Berant · 2024
Later among the works it cites.
Exact: Teaching ai agents to explore with reflective-mcts and exploratory learning, 2025
Xiao Yu, Baolin Peng, Vineeth Vajipey, Hao Cheng, Michel Galley, Jianfeng Gao, and Zhou Yu · 2025
Closest in time.
A survey on large language model based autonomous agents
Lei Wang, Chen Ma, Xueyang Feng, Zeyu Zhang, Hao Yang, Jingsen Zhang, Zhiyuan Chen, Jiakai Tang, Xu Chen, Yankai Lin, Wayne Xin Zhao, Zhewei Wei, and Jirong Wen · 2095
Closest in time.