Fetching the paper…
Reading the bibliography…
Training large language models (LLMs) as interactive agents presents unique challenges including long-horizon decision making and interacting with stochastic environment feedback.
Sokoban: Enhancing general single-agent search methods using domain knowledge
A. Junghanns and J. Schaeffer · 2001
Earlier work this paper cites.
Active learning literature survey
B. Settles · 2009
Earlier work this paper cites.
G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba · 2016
Earlier work this paper cites.
Proximal policy optimization algorithms, 2017
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Earlier work this paper cites.
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine · 2018
Earlier work this paper cites.
High-dimensional continuous control using generalized advantage estimation, 2018
J. Schulman, P. Moritz, S. Levine, M. Jordan, and P. Abbeel · 2018
Earlier work this paper cites.
Buy 4 REINFORCE samples, get a baseline for free!, 2019
W. Kool, H. van Hoof, and M. Welling · 2019
Earlier work this paper cites.
Decision transformer: Reinforcement learning via sequence modeling
L. Chen, K. Lu, A. Rajeswaran, K. Lee, A. Grover, M. Laskin, P. Abbeel, A. Srinivas, and I. Mordatch · 2021
Earlier work this paper cites.
The frozen lake problem. an example of optimization policy, 12 2021
P. Dell’Aversana · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models, 2021
E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen · 2021
Earlier work this paper cites.
Training language models to follow instructions with human feedback, 2022
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. L. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, J. Schulman, J. Hilton, F. Kelton, L. Miller, M. Simens, A. Askell, P. Welinder, P. Christiano, J. Leike, and R. Lowe · 2022
Earlier work this paper cites.
Star: Bootstrapping reasoning with reasoning, 2022
E. Zelikman, Y. Wu, J. Mu, and N. D. Goodman · 2022
Earlier work this paper cites.
Agentverse: Facilitating multi-agent collaboration and exploring emergent behaviors in agents
W. Chen, Y. Su, J. Zuo, C. Yang, C. Yuan, C. Qian, C.-M. Chan, Y. Qin, Y. Lu, R. Xie, et al · 2023
Earlier work this paper cites.
Deep reinforcement learning from human preferences, 2023
P. Christiano, J. Leike, T. B. Brown, M. Martic, S. Legg, and D. Amodei · 2023
Earlier work this paper cites.
Improving factuality and reasoning in language models through multiagent debate
Y. Du, S. Li, A. Torralba, J. B. Tenenbaum, and I. Mordatch · 2023
Earlier work this paper cites.
Camel: Communicative agents for" mind" exploration of large language model society
G. Li, H. Hammoud, H. Itani, D. Khizbullin, and B. Ghanem · 2023
Earlier work this paper cites.
H. Lightman, V. Kosaraju, Y. Burda, H. Edwards, B. Baker, T. Lee, J. Leike, J. Schulman, I. Sutskever, and K. Cobbe · 2023
Earlier work this paper cites.
Agentsims: An open-source sandbox for large language model evaluation, 2023
J. Lin, H. Zhao, A. Zhang, Y. Wu, H. Ping, and Q. Chen · 2023
Earlier work this paper cites.
Llm+ p: Empowering large language models with optimal planning proficiency
B. Liu, Y. Jiang, X. Zhang, Q. Liu, S. Zhang, J. Biswas, and P. Stone · 2023
Earlier work this paper cites.
Generative agents: Interactive simulacra of human behavior, 2023
J. S. Park, J. C. O’Brien, C. J. Cai, M. R. Morris, P. Liang, and M. S. Bernstein · 2023
Cited alongside, same era.
Autogen: Enabling next-gen llm applications via multi-agent conversation framework
Q. Wu, G. Bansal, J. Zhang, Y. Wu, S. Zhang, E. Zhu, B. Li, L. Jiang, X. Zhang, and C. Wang · 2023
Cited alongside, same era.
Openagents: An open platform for language agents in the wild
T. Xie, F. Zhou, Z. Cheng, P. Shi, L. Weng, Y. Liu, T. J. Hua, J. Zhao, Q. Liu, C. Liu, et al · 2023
Cited alongside, same era.
Rewoo: Decoupling reasoning from observations for efficient augmented language models
B. Xu, Z. Peng, B. Lei, S. Mukherjee, Y. Liu, and D. Xu · 2023
Cited alongside, same era.
Trial and error: Exploration-based trajectory optimization for llm agents, 2024
Y. Song, D. Yin, X. Yue, J. Huang, S. Li, and B. Y. Lin · 2024
Later among the works it cites.
Teaching embodied reinforcement learning agents: Informativeness and diversity of language use, 2024
J. Xi, Y. He, J. Yang, Y. Dai, and J. Chai · 2024
Later among the works it cites.
A. Yang, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Li, D. Liu, F. Huang, H. Wei, H. Lin, J. Yang, J. Tu, J. Zhang, J. Yang, J. Yang, J. Zhou, J. Lin, K. Dang, K. Lu, K. Bao, K. Yang, L. Yu, M. Li, M. Xue, P. Zhang, Q. Zhu, R. Men, R. Lin, T. Li, T. Xia, X. Ren, X. Ren, Y. Fan, Y. Su, Y. Zhang, Y. Wan, Y. Liu, Z. Cui, Z. Zhang, and Z. Qiu · 2024
Later among the works it cites.
Archer: Training language model agents via hierarchical multi-turn rl, 2024
Y. Zhou, A. Zanette, J. Pan, S. Levine, and A. Kumar · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Zhuang, X. Chen, T. Yu, S. Mitra, V. Bursztyn, R. A. Rossi, S. Sarkhel, and C. Zhang · 2023
Cited alongside, same era.
Deepseek LLM: scaling open-source language models with longtermism
DeepSeek-AI · 2024
Cited alongside, same era.
Deepseek-coder-v2: Breaking the barrier of closed-source models in code intelligence, 2024
DeepSeek-AI, Q. Zhu, D. Guo, Z. Shao, D. Yang, P. Wang, R. Xu, Y. Wu, Y. Li, H. Gao, S. Ma, W. Zeng, X. Bi, Z. Gu, H. Xu, D. Dai, K. Dong, L. Zhang, Y. Piao, Z. Gou, Z. Xie, Z. Hao, B. Wang, J. Song, D. Chen, X. Xie, K. Guan, Y. You, A. Liu, Q. Du, W. Gao, X. Lu, Q. Chen, Y. Wang, C. Deng, J. Li, C. Zhao, C. Ruan, F. Luo, and W. Liang · 2024
Cited alongside, same era.
Regressing the relative future: Efficient policy optimization for multi-turn rlhf, 2024
Z. Gao, W. Zhan, J. D. Chang, G. Swamy, K. Brantley, J. D. Lee, and W. Sun · 2024
Cited alongside, same era.
Think before you speak: Training language models with pause tokens, 2024
S. Goyal, Z. Ji, A. S. Rawat, A. K. Menon, S. Kumar, and V. Nagarajan · 2024
Cited alongside, same era.
Teaching large language models to reason with reinforcement learning, 2024
A. Havrilla, Y. Du, S. C. Raparthy, C. Nalmpantis, J. Dwivedi-Yu, M. Zhuravinskyi, E. Hambro, S. Sukhbaatar, and R. Raileanu · 2024
Cited alongside, same era.
Thinking tokens for language modeling, 2024
D. Herel and T. Mikolov · 2024
Cited alongside, same era.
SWE-bench: Can language models resolve real-world github issues?
C. E. Jimenez, J. Yang, A. Wettig, S. Yao, K. Pei, O. Press, and K. Narasimhan · 2024
Cited alongside, same era.
DeepSeek-AI, D. Guo, D. Yang, H. Zhang, J. Song, R. Zhang, R. Xu, Q. Zhu, S. Ma, P. Wang, X. Bi, X. Zhang, X. Yu, Y. Wu, Z. F. Wu, Z. Gou, Z. Shao, Z. Li, Z. Gao, A. Liu, B. Xue, B. Wang, B. Wu, B. Feng, C. Lu, C. Zhao, C. Deng, C. Zhang, C. Ruan, D. Dai, D. Chen, D. Ji, E. Li, F. Lin, F. Dai, F. Luo, G. Hao, G. Chen, G. Li, H. Zhang, H. Bao, H. Xu, H. Wang, H. Ding, H. Xin, H. Gao, H. Qu, H. Li, J. Guo, J. Li, J. Wang, J. Chen, J. Yuan, J. Qiu, J. Li, J. L. Cai, J. Ni, J. Liang, J. Chen, K. Dong, K. Hu, K. Gao, K. Guan, K. Huang, K. Yu, L. Wang, L. Zhang, L. Zhao, L. Wang, L. Zhang, L. Xu, L. Xia, M. Zhang, M. Zhang, M. Tang, M. Li, M. Wang, M. Li, N. Tian, P. Huang, P. Zhang, Q. Wang, Q. Chen, Q. Du, R. Ge, R. Zhang, R. Pan, R. Wang, R. J. Chen, R. L. Jin, R. Chen, S. Lu, S. Zhou, S. Chen, S. Ye, S. Wang, S. Yu, S. Zhou, S. Pan, S. S. Li, S. Zhou, S. Wu, S. Ye, T. Yun, T. Pei, T. Sun, T. Wang, W. Zeng, W. Zhao, W. Liu, W. Liang, W. Gao, W. Yu, W. Zhang, W. L. Xiao, W. An, X. Liu, X. Wang, X. Chen, X. Nie, X. Cheng, X. Liu, X. Xie, X. Liu, X. Yang, X. Li, X. Su, X. Lin, X. Q. Li, X. Jin, X. Shen, X. Chen, X. Sun, X. Wang, X. Song, X. Zhou, X. Wang, X. Shan, Y. K. Li, Y. Q. Wang, Y. X. Wei, Y. Zhang, Y. Xu, Y. Li, Y. Zhao, Y. Sun, Y. Wang, Y. Yu, Y. Zhang, Y. Shi, Y. Xiong, Y. He, Y. Piao, Y. Wang, Y. Tan, Y. Ma, Y. Liu, Y. Guo, Y. Ou, Y. Wang, Y. Gong, Y. Zou, Y. He, Y. Xiong, Y. Luo, Y. You, Y. Liu, Y. Zhou, Y. X. Zhu, Y. Xu, Y. Huang, Y. Li, Y. Zheng, Y. Zhu, Y. Ma, Y. Tang, Y. Zha, Y. Yan, Z. Z. Ren, Z. Ren, Z. Sha, Z. Fu, Z. Xu, Z. Xie, Z. Zhang, Z. Hao, Z. Ma, Z. Yan, Z. Wu, Z. Gu, Z. Zhu, Z. Liu, Z. Li, Z. Xie, Z. Song, Z. Pan, Z. Huang, Z. Xu, Z. Zhang, and Z. Zhang · 2025
Closest in time.
Group-in-group policy optimization for llm agent training, 2025
L. Feng, Z. Xue, T. Liu, and B. An · 2025
Closest in time.
Enhancing llm reasoning with multi-path collaborative reactive and reflection agents, 2025
C. He, B. Zou, X. Li, J. Chen, and H. M. Junliang Xing · 2025
Closest in time.
Embodied agent interface: Benchmarking llms for embodied decision making, 2025
M. Li, S. Zhao, Q. Wang, K. Wang, Y. Zhou, S. Srivastava, C. Gokmen, T. Lee, L. E. Li, R. Zhang, W. Liu, P. Liang, L. Fei-Fei, J. Mao, and J. Wu · 2025
Closest in time.
Understanding r1-zero-like training: A critical perspective, 2025
Z. Liu, C. Chen, W. Li, P. Qi, C. D. Tianyu Pang, W. S. Lee, and M. Lin · 2025
Closest in time.
Introducing ChatGPT o1, 2024
OpenAI · 2025
Closest in time.
Tinyzero
J. Pan, J. Zhang, X. Wang, L. Yuan, H. Peng, and A. Suhr · 2025
Closest in time.
Webrl: Training llm web agents via self-evolving online curriculum reinforcement learning, 2025
Z. Qi, X. Liu, I. L. Iong, H. Lai, X. Sun, W. Zhao, Y. Yang, X. Yang, J. Sun, S. Yao, T. Zhang, W. Xu, J. Tang, and Y. Dong · 2025
Closest in time.
Ui-tars: Pioneering automated gui interaction with native agents, 2025
Y. Qin, Y. Ye, J. Fang, H. Wang, S. Liang, S. Tian, J. Zhang, J. Li, Y. Li, S. Huang, W. Zhong, K. Li, J. Yang, Y. Miao, W. Lin, L. Liu, X. Jiang, Q. Ma, J. Li, X. Xiao, K. Cai, C. Li, Y. Zheng, C. Jin, C. Li, X. Zhou, M. Wang, H. Chen, Z. Li, H. Yang, H. Liu, F. Lin, T. Peng, X. Liu, and G. Shi · 2025
Closest in time.
Enigmaeval: A benchmark of long multimodal reasoning challenges, 2025
C. J. Wang, D. Lee, C. Menghini, J. Mols, J. Doughty, A. Khoja, J. Lynch, S. Hendryx, S. Yue, and D. Hendrycks · 2025
Closest in time.
Vagen: Training vlm agents with multi-turn reinforcement learning, 2025
K. Wang*, P. Zhang*, Z. Wang*, Q. Wang*, Y. Gao*, L. Li*, Z. Yang, C. Wan, H. Chen, Y. Lu, and M. Li · 2025
Closest in time.
Webagent-r1: Training web agents via end-to-end multi-turn reinforcement learning, 2025
Z. Wei, W. Yao, Y. Liu, W. Zhang, Q. Lu, L. Qiu, C. Yu, P. Xu, C. Zhang, B. Yin, H. Yun, and L. Li · 2025
Closest in time.
Dapo: An open-source llm reinforcement learning system at scale, 2025
Q. Yu, Z. Zhang, R. Zhu, Y. Yuan, X. Zuo, Y. Yue, T. Fan, G. Liu, L. Liu, X. Liu, H. Lin, Z. Lin, B. Ma, G. Sheng, Y. Tong, C. Zhang, M. Zhang, W. Zhang, H. Zhu, J. Zhu, J. Chen, J. Chen, C. Wang, H. Yu, W. Dai, Y. Song, X. Wei, H. Zhou, J. Liu, W.-Y. Ma, Y.-Q. Zhang, L. Yan, M. Qiao, Y. Wu, and M. Wang · 2025
Closest in time.
The lessons of developing process reward models in mathematical reasoning, 2025
Z. Zhang, C. Zheng, Y. Wu, B. Zhang, R. Lin, B. Yu, D. Liu, J. Zhou, and J. Lin · 2025
Closest in time.