Fetching the paper…
Reading the bibliography…
Interactive digital agents (IDAs) leverage APIs of stateful digital environments to perform tasks in response to user requests.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J · 1992
Earlier work this paper cites.
Planning and acting in partially observable stochastic domains
Kaelbling, L. P., Littman, M. L., and Cassandra, A. R · 1998
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Kakade, S. and Langford, J · 2002
Earlier work this paper cites.
Language understanding for text-based games using deep reinforcement learning
Narasimhan, K., Kulkarni, T. D., and Barzilay, R · 2015
Earlier work this paper cites.
High-dimensional continuous control using generalized advantage estimation
Schulman, J., Moritz, P., Levine, S., Jordan, M. I., and Abbeel, P · 2016
Earlier work this paper cites.
Thinking fast and slow with deep learning and tree search
Anthony, T., Tian, Z., and Barber, D · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
Buy 4 reinforce samples, get a baseline for free!
Kool, W., van Hoof, H., and Welling, M · 2019
Earlier work this paper cites.
Fine-tuning language models from human preferences
Ziegler, D. M., Stiennon, N., Wu, J., Brown, T. B., Radford, A., Amodei, D., Christiano, P., and Irving, G · 2019
Earlier work this paper cites.
Learning to summarize from human feedback
Stiennon, N., Ouyang, L., Wu, J., Ziegler, D. M., Lowe, R., Voss, C., Radford, A., Amodei, D., and Christiano, P · 2020
Earlier work this paper cites.
DD-PPO: Learning near-perfect PointGoal navigators from 2.5 billion frames
Wijmans, E., Kadian, A., Morcos, A., Lee, S., Essa, I., Parikh, D., Savva, M., and Batra, D · 2020
Earlier work this paper cites.
Keep CALM and explore: Language models for action generation in text-based games
Yao, S., Rao, R., Hausknecht, M., and Narasimhan, K · 2020
Earlier work this paper cites.
WebGPT: Browser-assisted question-answering with human feedback
Nakano, R., Hilton, J., Balaji, S., Wu, J., Ouyang, L., Kim, C., Hesse, C., Jain, S., Kosaraju, V., Saunders, W., et al · 2021
Earlier work this paper cites.
LoRA: Low-rank adaptation of large language models
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Earlier work this paper cites.
WebShop: Towards scalable real-world web interaction with grounded language agents
Yao, S., Chen, H., Yang, J., and Narasimhan, K · 2022
Earlier work this paper cites.
STar: Bootstrapping reasoning with reasoning
Zelikman, E., Wu, Y., Mu, J., and Goodman, N · 2022
Earlier work this paper cites.
Grounding large language models in interactive environments with online reinforcement learning
Carta, T., Romac, C., Wolf, T., Lamprier, S., Sigaud, O., and Oudeyer, P.-Y · 2023
Cited alongside, same era.
FireAct: Toward language agent fine-tuning
Chen, B., Shu, C., Shareghi, E., Collier, N., Narasimhan, K., and Yao, S · 2023
Cited alongside, same era.
Efficient memory management for large language model serving with PagedAttention
Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., Gonzalez, J., Zhang, H., and Stoica, I · 2023
Cited alongside, same era.
Toolformer: Language models can teach themselves to use tools
Schick, T., Dwivedi-Yu, J., Dessì, R., Raileanu, R., Lomeli, M., Hambro, E., Zettlemoyer, L., Cancedda, N., and Scialom, T · 2023
Cited alongside, same era.
Intercode: standardizing and benchmarking interactive coding with execution feedback
Yang, J., Prabhakar, A., Narasimhan, K., and Yao, S · 2023
Cited alongside, same era.
Introducing OpenAI o1, 2024
OpenAI · 2024
Later among the works it cites.
Agent Q: Advanced reasoning and learning for autonomous AI agents
Putta, P., Mills, E., Garg, N., Motwani, S., Finn, C., Garg, D., and Rafailov, R · 2024
Later among the works it cites.
ToolLLM: Facilitating large language models to master 16000+ real-world APIs
Qin, Y., Liang, S., Ye, Y., Zhu, K., Yan, L., Lu, Y., Lin, Y., Cong, X., Tang, X., Qian, B., et al · 2024
Later among the works it cites.
DeepSeekMath: Pushing the limits of mathematical reasoning in open language models
Shao, Z., Wang, P., Zhu, Q., Xu, R., Song, J., Bi, X., Zhang, H., Zhang, M., Li, Y., Wu, Y., et al · 2024
Later among the works it cites.
Direct multi-turn preference optimization for language agents
Shi, W., Yuan, M., Wu, J., Wang, Q., and Feng, F · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
ReAct: Synergizing reasoning and acting in language models
Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K. R., and Cao, Y · 2023
Cited alongside, same era.
Scaling relationship on learning mathematical reasoning with large language models
Yuan, Z., Yuan, H., Li, C., Dong, G., Lu, K., Tan, C., Zhou, C., and Zhou, J · 2023
Cited alongside, same era.
PyTorch FSDP: Experiences on scaling fully sharded data parallel
Zhao, Y., Gu, A., Varma, R., Luo, L., Huang, C.-C., Xu, M., Wright, L., Shojanazeri, H., Ott, M., Shleifer, S., Desmaison, A., Balioglu, C., Damania, P., Nguyen, B., Chauhan, G., Hao, Y., Mathews, A., and Li, S · 2023
Cited alongside, same era.
Back to basics: Revisiting REINFORCE-style optimization for learning from human feedback in LLMs
Ahmadian, A., Cremer, C., Gallé, M., Fadaee, M., Kreutzer, J., Pietquin, O., Üstün, A., and Hooker, S · 2024
Cited alongside, same era.
DigiRL: Training in-the-wild device-control agents with autonomous reinforcement learning
Bai, H., Zhou, Y., Cemri, M., Pan, J., Suhr, A., Levine, S., and Kumar, A · 2024
Cited alongside, same era.
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al · 2024
Cited alongside, same era.
Teaching large language models to reason with reinforcement learning
Havrilla, A., Du, Y., Raparthy, S. C., Nalmpantis, C., Dwivedi-Yu, J., Zhuravinskyi, M., Hambro, E., Sukhbaatar, S., and Raileanu, R · 2024
Cited alongside, same era.
Reflexion: Language agents with verbal reinforcement learning
Shinn, N., Cassano, F., Gopinath, A., Narasimhan, K., and Yao, S · 2024
Later among the works it cites.
Beyond human data: Scaling self-training for problem-solving with language models
Singh, A., Co-Reyes, J. D., Agarwal, R., Anand, A., Patil, P., Garcia, X., Liu, P. J., Harrison, J., Lee, J., Xu, K., et al · 2024
Later among the works it cites.
torchtune: PyTorch’s finetuning library, April 2024
torchtune maintainers and contributors · 2024
Later among the works it cites.
AppWorld: A controllable world of apps and people for benchmarking interactive coding agents
Trivedi, H., Khot, T., Hartmann, M., Manku, R., Dong, V., Li, E., Gupta, S., Sabharwal, A., and Balasubramanian, N · 2024
Later among the works it cites.
Executable code actions elicit better LLM agents
Wang, X., Chen, Y., Yuan, L., Zhang, Y., Li, Y., Peng, H., and Ji, H · 2024
Later among the works it cites.
Cut your losses in large-vocabulary language models
Wijmans, E., Huval, B., Hertzberg, A., Koltun, V., and Krähenbühl, P · 2024
Later among the works it cites.
Yang, A., Yang, B., Zhang, B., Hui, B., Zheng, B., Yu, B., Li, C., Liu, D., Huang, F., Wei, H., et al · 2024
Later among the works it cites.
Fine-tuning large vision-language models as decision-making agents via reinforcement learning
Zhai, Y., Bai, H., Lin, Z., Pan, J., Tong, S., Zhou, Y., Suhr, A., Xie, S., LeCun, Y., Ma, Y., and Levine, S · 2024
Later among the works it cites.
ArCHer: training language model agents via hierarchical multi-turn RL
Zhou, Y. and Zanette, A · 2024
Later among the works it cites.
DeepSeek-R1: Incentivizing reasoning capability in LLMs via reinforcement learning
DeepSeek-AI · 2025
Closest in time.
A simple and effective reinforcement learning method for text-to-image diffusion fine-tuning
Gupta, S., Ahuja, C., Lin, T.-Y., Roy, S. D., Oosterhuis, H., de Rijke, M., and Shukla, S. N · 2025
Closest in time.