Fetching the paper…
Reading the bibliography…
Large language models are quickly becoming the foundation for intelligent agents that are capable of using tools.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J. Williams · 2004
Earlier work this paper cites.
Active learning of inverse models with intrinsically motivated goal exploration in robots
Adrien Baranes and Pierre-Yves Oudeyer · 2012
Earlier work this paper cites.
Proximal policy optimization algorithms, 2017
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Intrinsic motivation and automatic curricula via asymmetric self-play, 2018
Sainbayar Sukhbaatar, Zeming Lin, Ilya Kostrikov, Gabriel Synnaeve, Arthur Szlam, and Rob Fergus · 2018
Earlier work this paper cites.
Asymmetric self-play for automatic goal discovery in robotic manipulation, 2021
OpenAI OpenAI, Matthias Plappert, Raul Sampedro, Tao Xu, Ilge Akkaya, Vineet Kosaraju, Peter Welinder, Ruben D’Sa, Arthur Petron, Henrique P. d. O. Pinto, Alex Paino, Hyeonwoo Noh, Lilian Weng, Qiming Yuan, Casey Chu, and Wojciech Zaremba · 2021
Earlier work this paper cites.
Language models as zero-shot planners: Extracting actionable knowledge for embodied agents, 2022
Wenlong Huang, Pieter Abbeel, Deepak Pathak, and Igor Mordatch · 2022
Earlier work this paper cites.
Augmenting autotelic agents with large language models, 2023
Cédric Colas, Laetitia Teodorescu, Pierre-Yves Oudeyer, Xingdi Yuan, and Marc-Alexandre Côté · 2023
Earlier work this paper cites.
Voxposer: Composable 3d value maps for robotic manipulation with language models, 2023
Wenlong Huang, Chen Wang, Ruohan Zhang, Yunzhu Li, Jiajun Wu, and Li Fei-Fei · 2023
Earlier work this paper cites.
Code as policies: Language model programs for embodied control, 2023
Jacky Liang, Wenlong Huang, Fei Xia, Peng Xu, Karol Hausman, Brian Ichter, Pete Florence, and Andy Zeng · 2023
Earlier work this paper cites.
Agentbench: Evaluating llms as agents, 2023
Xiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu, Xuanyu Lei, Hanyu Lai, Yu Gu, Hangliang Ding, Kaiwen Men, Kejuan Yang, Shudan Zhang, Xiang Deng, Aohan Zeng, Zhengxiao Du, Chenhui Zhang, Sheng Shen, Tianjun Zhang, Yu Su, Huan Sun, Minlie Huang, Yuxiao Dong, and Jie Tang · 2023
Earlier work this paper cites.
Toolllm: Facilitating large language models to master 16000+ real-world apis, 2023
Yujia Qin, Shihao Liang, Yining Ye, Kunlun Zhu, Lan Yan, Yaxi Lu, Yankai Lin, Xin Cong, Xiangru Tang, Bill Qian, Sihan Zhao, Lauren Hong, Runchu Tian, Ruobing Xie, Jie Zhou, Mark Gerstein, Dahai Li, Zhiyuan Liu, and Maosong Sun · 2023
Earlier work this paper cites.
Toolformer: Language models can teach themselves to use tools, 2023
Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom · 2023
Earlier work this paper cites.
Self-instruct: Aligning language models with self-generated instructions, 2023
Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, Daniel Khashabi, and Hannaneh Hajishirzi · 2023
Earlier work this paper cites.
React: Synergizing reasoning and acting in language models, 2023
Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao · 2023
Earlier work this paper cites.
Digirl: Training in-the-wild device-control agents with autonomous reinforcement learning, 2024
Hao Bai, Yifei Zhou, Mert Cemri, Jiayi Pan, Alane Suhr, Sergey Levine, and Aviral Kumar · 2024
Earlier work this paper cites.
Self-play fine-tuning converts weak language models to strong language models
Zixiang Chen, Yihe Deng, Huizhuo Yuan, Kaixuan Ji, and Quanquan Gu · 2024
Cited alongside, same era.
Mengkang Hu, Pu Zhao, Can Xu, Qingfeng Sun, Jianguang Lou, Qingwei Lin, Ping Luo, Saravan Rajmohan, and Dongmei Zhang · 2024
Cited alongside, same era.
Swe-bench: Can language models resolve real-world github issues?, 2024
Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan · 2024
Cited alongside, same era.
Toolverifier: Generalization to new tools via self-verification, 2024
Dheeraj Mekala, Jason Weston, Jack Lanchantin, Roberta Raileanu, Maria Lomeli, Jingbo Shang, and Jane Dwivedi-Yu · 2024
Cited alongside, same era.
Omni: Open-endedness via models of human notions of interestingness, 2024
Jenny Zhang, Joel Lehman, Kenneth Stanley, and Jeff Clune · 2024
Later among the works it cites.
Digi-q: Learning q-value functions for training device-control agents, 2025
Hao Bai, Yifei Zhou, Li Erran Li, Sergey Levine, and Aviral Kumar · 2025
Closest in time.
Stp: Self-play llm theorem provers with iterative conjecturing and proving, 2025
Kefan Dong and Tengyu Ma · 2025
Closest in time.
Maxence Faldor, Jenny Zhang, Antoine Cully, and Jeff Clune · 2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Shikhar Murty, Christopher Manning, Peter Shaw, Mandar Joshi, and Kenton Lee · 2024
Cited alongside, same era.
OpenAI · 2024
Cited alongside, same era.
Direct preference optimization: Your language model is secretly a reward model, 2024
Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D. Manning, and Chelsea Finn · 2024
Cited alongside, same era.
Deepseekmath: Pushing the limits of mathematical reasoning in open language models, 2024
Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, Y. K. Li, Y. Wu, and Daya Guo · 2024
Cited alongside, same era.
The llama 3 herd of models, 2024
Llma3 Team · 2024
Cited alongside, same era.
Appworld: A controllable world of apps and people for benchmarking interactive coding agents, 2024
Harsh Trivedi, Tushar Khot, Mareike Hartmann, Ruskin Manku, Vinty Dong, Edward Li, Shashank Gupta, Ashish Sabharwal, and Niranjan Balasubramanian · 2024
Cited alongside, same era.
Executable code actions elicit better llm agents, 2024
Xingyao Wang, Yangyi Chen, Lifan Yuan, Yizhe Zhang, Yunzhu Li, Hao Peng, and Heng Ji · 2024
Cited alongside, same era.
Theagentcompany: Benchmarking llm agents on consequential real world tasks, 2024
Frank F. Xu, Yufan Song, Boxuan Li, Yuxuan Tang, Kritanjali Jain, Mengxue Bao, Zora Z. Wang, Xuhui Zhou, Zhitong Guo, Murong Cao, Mingyang Yang, Hao Yang Lu, Amaad Martin, Zhe Su, Leander Maben, Raj Mehta, Wayne Chi, Lawrence Jang, Yiqing Xie, Shuyan Zhou, and Graham Neubig · 2024
Cited alongside, same era.
Zhicheng Guo, Sijie Cheng, Hao Wang, Shihao Liang, Yujia Qin, Peng Li, Zhiyuan Liu, Maosong Sun, and Yang Liu · 2025
Closest in time.
Nnetnav: Unsupervised learning of browser agents through environment interaction in the wild, 2025
Shikhar Murty, Hao Zhu, Dzmitry Bahdanau, and Christopher D. Manning · 2025
Closest in time.
Webrl: Training llm web agents via self-evolving online curriculum reinforcement learning, 2025
Zehan Qi, Xiao Liu, Iat Long Iong, Hanyu Lai, Xueqiao Sun, Wenyi Zhao, Yu Yang, Xinyue Yang, Jiadai Sun, Shuntian Yao, Tianjie Zhang, Wei Xu, Jie Tang, and Yuxiao Dong · 2025
Closest in time.
Tool learning in the wild: Empowering language models as automatic tool agents, 2025
Zhengliang Shi, Shen Gao, Lingyong Yan, Yue Feng, Xiuyi Chen, Zhumin Chen, Dawei Yin, Suzan Verberne, and Zhaochun Ren · 2025
Closest in time.
Beyond browsing: Api-based web agents, 2025
Yueqi Song, Frank Xu, Shuyan Zhou, and Graham Neubig · 2025
Closest in time.
Learn-by-interact: A data-centric framework for self-adaptive agents in realistic environments, 2025
Hongjin Su, Ruoxi Sun, Jinsung Yoon, Pengcheng Yin, Tao Yu, and Sercan Ö. Arık · 2025
Closest in time.
Taiyi Wang, Zhihao Wu, Jianheng Liu, Jianye Hao, Jun Wang, and Kun Shao · 2025
Closest in time.
Building math agents with multi-turn iterative preference learning, 2025
Wei Xiong, Chengshuai Shi, Jiaming Shen, Aviv Rosenberg, Zhen Qin, Daniele Calandriello, Misha Khalman, Rishabh Joshi, Bilal Piot, Mohammad Saleh, Chi Jin, Tong Zhang, and Tianqi Liu · 2025
Closest in time.
Absolute zero: Reinforced self-play reasoning with zero data, 2025
Andrew Zhao, Yiran Wu, Yang Yue, Tong Wu, Quentin Xu, Yang Yue, Matthieu Lin, Shenzhi Wang, Qingyun Wu, Zilong Zheng, and Gao Huang · 2025
Closest in time.
Sweet-rl: Training multi-turn llm agents on collaborative reasoning tasks, 2025
Yifei Zhou, Song Jiang, Yuandong Tian, Jason Weston, Sergey Levine, Sainbayar Sukhbaatar, and Xian Li · 2025
Closest in time.