Fetching the paper…
Reading the bibliography…
Language agents have become a promising solution to complex interactive tasks.
Learning to predict by the methods of temporal differences
Sutton, R. S · 1988
Earlier work this paper cites.
Q-learning
Watkins, C. J. and Dayan, P · 1992
Earlier work this paper cites.
Applied dynamic programming , volume 2050
Bellman, R. E. and Dreyfus, S. E · 2015
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
Is q-learning provably efficient?
Jin, C., Allen-Zhu, Z., Bubeck, S., and Jordan, M. I · 2018
Earlier work this paper cites.
{ALFW}orld: Aligning text and embodied environments for interactive learning
Shridhar, M., Yuan, X., Cote, M.-A., Bisk, Y., Trischler, A., and Hausknecht, M · 2021
Earlier work this paper cites.
Solving math word problems with process- and outcome-based feedback, 2022
Uesato, J., Kushman, N., Kumar, R., Song, F., Siegel, N., Wang, L., Creswell, A., Irving, G., and Higgins, I · 2022
Earlier work this paper cites.
ScienceWorld: Is your agent smarter than a 5th grader?
Wang, R., Jansen, P., Côté, M.-A., and Ammanabrolu, P · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al · 2022
Earlier work this paper cites.
A universal discriminator for zero-shot generalization
Xu, H., Lin, Z., Zhou, J., Zheng, Y., and Yang, Z · 2022
Earlier work this paper cites.
Webshop: Towards scalable real-world web interaction with grounded language agents
Yao, S., Chen, H., Yang, J., and Narasimhan, K · 2022
Earlier work this paper cites.
Fireact: Toward language agent fine-tuning
Chen, B., Shu, C., Shareghi, E., Collier, N., Narasimhan, K., and Yao, S · 2023
Earlier work this paper cites.
Alphazero-like tree-search can guide large language model decoding and training
Feng, X., Wan, Z., Wen, M., Wen, Y., Zhang, W., and Wang, J · 2023
Earlier work this paper cites.
Reinforced self-training (rest) for language modeling, 2023
Gulcehre, C., Paine, T. L., Srinivasan, S., Konyushkova, K., Weerts, L., Sharma, A., Siddhant, A., Ahern, A., Wang, M., Gu, C., Macherey, W., Doucet, A., Firat, O., and de Freitas, N · 2023
Earlier work this paper cites.
Lightman, H., Kosaraju, V., Burda, Y., Edwards, H., Baker, B., Lee, T., Leike, J., Schulman, J., Sutskever, I., and Cobbe, K · 2023
Cited alongside, same era.
Reflexion: language agents with verbal reinforcement learning
Shinn, N., Cassano, F., Gopinath, A., Narasimhan, K. R., and Yao, S · 2023
Cited alongside, same era.
Beyond human data: Scaling self-training for problem-solving with language models
Singh, A., Co-Reyes, J. D., Agarwal, R., Anand, A., Patil, P., Liu, P. J., Harrison, J., Lee, J., Xu, K., Parisi, A., et al · 2023
Cited alongside, same era.
Restgpt: Connecting large language models with real-world restful apis
Song, Y., Xiong, W., Zhu, D., Wu, W., Qian, H., Song, M., Huang, H., Li, C., Wang, K., Yao, R., et al · 2023
Cited alongside, same era.
Math-shepherd: A label-free step-by-step verifier for llms in mathematical reasoning
Rl on incorrect synthetic data scales the efficiency of llm math reasoning by eight-fold
Setlur, A., Garg, S., Geng, X., Garg, N., Smith, V., and Kumar, A · 2024
Later among the works it cites.
Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face
Shen, Y., Song, K., Tan, X., Li, D., Lu, W., and Zhuang, Y · 2024
Later among the works it cites.
Scaling llm test-time compute optimally can be more effective than scaling model parameters
Snell, C., Lee, J., Xu, K., and Kumar, A · 2024
Later among the works it cites.
Trial and error: Exploration-based trajectory optimization of LLM agents
Song, Y., Yin, D., Yue, X., Huang, J., Li, S., and Lin, B. Y · 2024
Later among the works it cites.
Q*: Improving multi-step reasoning for llms with deliberative planning
Wang, C., Deng, Y., Lv, Z., Yan, S., and Bo, A · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Wang, P., Li, L., Shao, Z., Xu, R., Dai, D., Li, Y., Chen, D., Wu, Y., and Sui, Z · 2023
Cited alongside, same era.
React: Synergizing reasoning and acting in language models
Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K. R., and Cao, Y · 2023
Cited alongside, same era.
Scaling relationship on learning mathematical reasoning with large language models
Yuan, Z., Yuan, H., Li, C., Dong, G., Lu, K., Tan, C., Zhou, C., and Zhou, J · 2023
Cited alongside, same era.
Videophy: Evaluating physical commonsense for video generation
Bansal, H., Lin, Z., Xie, T., Zong, Z., Yarom, M., Bitton, Y., Jiang, C., Sun, Y., Chang, K.-W., and Grover, A · 2024
Cited alongside, same era.
AutoPRM: Automating procedural supervision for multi-step reasoning via controllable question decomposition
Chen, Z., Zhao, Z., Zhu, Z., Zhang, R., Li, X., Raj, B., and Yao, H · 2024
Cited alongside, same era.
Reflection-reinforced self-training for language agents
Dou, Z.-Y., Yang, C.-F., Wu, X., Chang, K.-W., and Peng, N · 2024
Cited alongside, same era.
V-star: Training verifiers for self-taught reasoners
Hosseini, A., Yuan, X., Malkin, N., Courville, A., Sordoni, A., and Agarwal, R · 2024
Cited alongside, same era.
Agent q: Advanced reasoning and learning for autonomous ai agents
Putta, P., Mills, E., Garg, N., Motwani, S., Finn, C., Garg, D., and Rafailov, R · 2024
Cited alongside, same era.
Later among the works it cites.
Vdebugger: Harnessing execution feedback for debugging visual programs
Wu, X., Lin, Z., Zhao, S., Wu, T.-L., Lu, P., Peng, N., and Chang, K.-W · 2024
Later among the works it cites.
Agent lumos: Unified and modular training for open-source language agents
Yin, D., Brahman, F., Ravichander, A., Chandu, K., Chang, K.-W., Choi, Y., and Lin, B. Y · 2024
Later among the works it cites.
Free process rewards without process labels
Yuan, L., Li, W., Chen, H., Cui, G., Ding, N., Zhang, K., Zhou, B., Liu, Z., and Peng, H · 2024
Later among the works it cites.
Enhancing decision-making for llm agents via step-level q-value models
Zhai, Y., Yang, T., Xu, K., Dawei, F., Yang, C., Ding, B., and Wang, H · 2024
Later among the works it cites.
ReST-MCTS*: LLM self-training via process reward guided tree search
Zhang, D., Zhoubian, S., Hu, Z., Yue, Y., Dong, Y., and Tang, J · 2024
Later among the works it cites.
Language agent tree search unifies reasoning, acting, and planning in language models
Zhou, A., Yan, K., Shlapentokh-Rothman, M., Wang, H., and Wang, Y.-X · 2024
Later among the works it cites.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Guo, D., Yang, D., Zhang, H., Song, J., Zhang, R., Xu, R., Zhu, Q., Ma, S., Wang, P., Bi, X., et al · 2025
Closest in time.
Kimi k1. 5: Scaling reinforcement learning with llms
Team, K., Du, A., Gao, B., Xing, B., Jiang, C., Chen, C., Li, C., Xiao, C., Du, C., Liao, C., et al · 2025
Closest in time.