Fetching the paper…
Reading the bibliography…
Large Language Model (LLM)-based web agents excel at knowledge-intensive tasks but face a fundamental conflict between the need for extensive exploration and the constraints of limited context windows.
Deep reinforcement learning from human preferences
Christiano, P. F., Leike, J., Brown, T. B., Martic, M., Legg, S., and Amodei, D · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., de Las Casas, D., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., Lavaud, L. R., Lachaux, M.-A., Stock, P., Scao, T. L., Lavril, T., Wang, T., Lacroix, T., and Sayed, W. E · 2023
Earlier work this paper cites.
Gaia: a benchmark for general ai assistants
Mialon, G., Fourrier, C., Wolf, T., LeCun, Y., and Scialom, T · 2023
Earlier work this paper cites.
React: Synergizing reasoning and acting in language models
Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K., and Cao, Y · 2023
Earlier work this paper cites.
Gu, J., Jiang, X., Shi, Z., Tan, H., Zhai, X., Xu, C., Li, W., Shen, Y., Ma, S., Liu, H., et al · 2024
Earlier work this paper cites.
Llms-as-judges: A comprehensive survey on llm-based evaluation methods, 2024
Li, H., Dong, Q., Chen, J., Su, H., Zhou, Y., Ai, Q., Ye, Z., and Liu, Y · 2024
Earlier work this paper cites.
Deepseekmath: Pushing the limits of mathematical reasoning in open language models
Shao, Z., Wang, P., Zhu, Q., Xu, R., Song, J., Bi, X., Zhang, H., Zhang, M., Li, Y., Wu, Y., et al · 2024
Earlier work this paper cites.
A survey on large language model based autonomous agents
Wang, L., Ma, C., Feng, X., Zhang, Z., Yang, H., Zhang, J., Chen, Z., Tang, J., Chen, X., Lin, Y., et al · 2024
Earlier work this paper cites.
Measuring short-form factuality in large language models, 2024
Wei, J., Karina, N., Chung, H. W., Jiao, Y. J., Papay, S., Glaese, A., Schulman, J., and Fedus, W · 2024
Earlier work this paper cites.
Yang, A., Yang, B., Zhang, B., Hui, B., Zheng, B., Yu, B., Li, C., Liu, D., Huang, F., Wei, H., et al · 2024
Cited alongside, same era.
Web agents with world models: Learning and leveraging environment dynamics in web navigation
Chae, H., Kim, N., iunn Ong, K. T., Gwak, M., Song, G., Kim, J., Kim, S., Lee, D., and Yeo, J · 2025
Cited alongside, same era.
Agentic reinforced policy optimization
Dong, G., Mao, H., Ma, K., Bao, L., Chen, Y., Wang, Z., Chen, Z., Du, J., Wang, H., Zhang, F., et al · 2025
Cited alongside, same era.
Beyond ten turns: Unlocking long-horizon agentic search with large-scale asynchronous rl, 2025
Gao, J., Fu, W., Xie, M., Xu, S., He, C., Mei, Z., Zhu, B., and Wu, Y · 2025
Cited alongside, same era.
DeepSeek-R1: Incentivizing reasoning capability in LLMs via reinforcement learning
Reasoningbank: Scaling agent self-evolving with reasoning memory, 2025
Ouyang, S., Yan, J., Hsu, I.-H., Chen, Y., Jiang, K., Wang, Z., Han, R., Le, L. T., Daruki, S., Tang, X., Tirumalashetty, V., Lee, G., Rofouei, M., Lin, H., Han, J., Lee, C.-Y., and Pfister, T · 2025
Closest in time.
rllm: A framework for post-training language agents, 2025
Tan, S., Luo, M., Cai, C., Venkat, T., Montgomery, K., Hao, A., Wu, T., Balyan, A., Roongta, M., Wang, C., Li, L. E., Popa, R. A., and Stoica, I · 2025
Closest in time.
Webshaper: Agentically data synthesizing via information-seeking formalization
Tao, Z., Wu, J., Yin, W., Zhang, J., Li, B., Shen, H., Li, K., Zhang, L., Wang, X., Jiang, Y., Xie, P., Huang, F., and Zhou, J · 2025
Closest in time.
Team, Q · 2025
Closest in time.
Browsecomp: A simple yet challenging benchmark for browsing agents
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Guo, D., Yang, D., Zhang, H., Song, J., Zhang, R., Xu, R., Zhu, Q., Ma, S., Wang, P., Bi, X., et al · 2025
Cited alongside, same era.
Search-r1: Training llms to reason and leverage search engines with reinforcement learning
Jin, B., Zeng, H., Yue, Z., Yoon, J., Arik, S., Wang, D., Zamani, H., and Han, J · 2025
Cited alongside, same era.
Jina, 2025
Jina.ai · 2025
Cited alongside, same era.
Webexplorer: Explore and evolve for training long-horizon web agents, 2025
Liu, J., Li, Y., Zhang, C., Li, J., Chen, A., Ji, K., Cheng, W., Wu, Z., Du, C., Xu, Q., Song, J., Zhu, Z., Chen, W., Zhao, P., and He, J · 2025
Cited alongside, same era.
Deepdive: Advancing deep search agents with knowledge graphs and multi-turn rl, 2025
Lu, R., Hou, Z., Wang, Z., Zhang, H., Liu, X., Li, Y., Feng, S., Tang, J., and Dong, Y · 2025
Cited alongside, same era.
Openai deep research
OpenAI · 2025
Cited alongside, same era.
Websailor: Navigating super-human reasoning for web agent
Li, K., Zhang, Z., Yin, H., Zhang, L., Ou, L., Wu, J., Yin, W., Li, B., Tao, Z., Wang, X., Shen, W., Zhang, J., Zhang, D., Wu, X., Jiang, Y., Yan, M., Xie, P., Huang, F., and Zhou, J
Cited in the paper.
Li, W., Lin, J., Jiang, Z., Cao, J., Liu, X., Zhang, J., Huang, Z., Chen, Q., Sun, W., Wang, Q., Lu, H., Qin, T., Zhu, C., Yao, Y., Fan, S., Li, X., Wang, T., Liu, P., Zhu, K., Zhu, H., Shi, D., Wang, P., Guan, Y., Tang, X., Liu, M., Jiang, Y. E., Yang, J., Liu, J., Zhang, G., and Zhou, W
Cited in the paper.
Wei, J., Sun, Z., Papay, S., McKinney, S., Han, J., Fulford, I., Chung, H. W., Passos, A. T., Fedus, W., and Glaese, A · 2025
Closest in time.
Xbench-deepsearch, 2025
Xbench-Team · 2025
Closest in time.
The rise and potential of large language model based agents: A survey
Xi, Z., Chen, W., Guo, X., He, W., Ding, Y., Hong, B., Zhang, M., Wang, J., Jin, S., Zhou, E., et al · 2025
Closest in time.
A-mem: Agentic memory for llm agents
Xu, W., Liang, Z., Mei, K., Gao, H., Tan, J., and Zhang, Y · 2025
Closest in time.
Yuan, Q., Lou, J., Li, Z., Chen, J., Lu, Y., Lin, H., Sun, L., Zhang, D., and Han, X · 2025
Closest in time.
Mem1: Learning to synergize memory and reasoning for efficient long-horizon agents, 2025b
Zhou, Z., Qu, A., Wu, Z., Kim, S., Prakash, A., Rus, D., Zhao, J., Low, B. K. H., and Liang, P. P · 2025
Closest in time.