Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have shown remarkable advancements in enabling language agents to tackle simple tasks.
World of bits: An open-domain platform for web-based agents
Shi, T., Karpathy, A., Fan, L., Hernandez, J., and Liang, P · 2017
Earlier work this paper cites.
Reinforcement learning: An introduction
Sutton, R. S · 2018
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2021
Earlier work this paper cites.
Understanding html with large language models
Gur, I., Nachum, O., Miao, Y., Safdari, M., Huang, A., Chowdhery, A., Narang, S., Fiedel, N., and Faust, A · 2022
Earlier work this paper cites.
Large language models are zero-shot reasoners
Kojima, T., Gu, S. S., Reid, M., Matsuo, Y., and Iwasawa, Y · 2022
Earlier work this paper cites.
Self-instruct: Aligning language models with self-generated instructions
Wang, Y., Kordi, Y., Mishra, S., Liu, A., Smith, N. A., Khashabi, D., and Hajishirzi, H · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al · 2022
Earlier work this paper cites.
Gunasekar, S., Zhang, Y., Aneja, J., Mendes, C. C. T., Del Giorno, A., Gopi, S., Javaheripi, M., Kauffmann, P., de Rosa, G., Saarikivi, O., et al · 2023
Earlier work this paper cites.
A real-world webagent with planning, long context understanding, and program synthesis
Gur, I., Furuta, H., Huang, A., Safdari, M., Matsuo, Y., Eck, D., and Faust, A · 2023
Earlier work this paper cites.
Adapt: As-needed decomposition and planning with language models
Prasad, A., Koller, A., Hartmann, M., Clark, P., Sabharwal, A., Bansal, M., and Khot, T · 2023
Earlier work this paper cites.
Heap: Hierarchical policies for web actions using llms
Sodhi, P., Branavan, S., and McDonald, R · 2023
Earlier work this paper cites.
Llm-planner: Few-shot grounded planning for embodied agents with large language models
Song, C. H., Wu, J., Washington, C., Sadler, B. M., Chao, W.-L., and Su, Y · 2023
Earlier work this paper cites.
Hierarchical prompting assists large language model on web navigation
Sridhar, A., Lo, R., Xu, F. F., Zhu, H., and Zhou, S · 2023
Earlier work this paper cites.
Adaplanner: Adaptive planning from feedback with language models
Sun, H., Zhuang, Y., Kong, L., Dai, B., and Zhang, C · 2023
Earlier work this paper cites.
Stanford alpaca: An instruction-following llama model, 2023
Taori, R., Gulrajani, I., Zhang, T., Dubois, Y., Li, X., Guestrin, C., Liang, P., and Hashimoto, T. B · 2023
Earlier work this paper cites.
Plan-and-solve prompting: Improving zero-shot chain-of-thought reasoning by large language models
Wang, L., Xu, W., Lan, Y., Hu, Z., Lan, Y., Lee, R. K.-W., and Lim, E.-P · 2023
Cited alongside, same era.
Wizardlm: Empowering large language models to follow complex instructions
Xu, C., Sun, Q., Zheng, K., Geng, X., Zhao, P., Feng, J., Tao, C., and Jiang, D · 2023
Cited alongside, same era.
Appagent: Multimodal agents as smartphone users
Zhang, C., Yang, Z., Liu, J., Han, Y., Chen, X., Huang, Z., Fu, B., and Yu, G · 2023
Cited alongside, same era.
Webarena: A realistic web environment for building autonomous agents
Zhou, S., Xu, F. F., Zhu, H., Zhou, X., Lo, R., Sridhar, A., Cheng, X., Ou, T., Bisk, Y., Fried, D., et al · 2023
Cited alongside, same era.
Wilbur: Adaptive in-context learning for robust and accurate web agents
Lutz, M., Bohra, A., Saroyan, M., Harutyunyan, A., and Campagna, G · 2024
Later among the works it cites.
Efficient and scalable estimation of tool representations in vector space
Moon, S., Jha, S., Erdogan, L. E., Kim, S., Lim, W., Keutzer, K., and Gholami, A · 2024
Later among the works it cites.
Nnetscape navigator: Complex demonstrations for web agents without a demonstrator
Murty, S., Bahdanau, D., and Manning, C. D · 2024
Later among the works it cites.
Long-horizon planning for multi-agent robots in partially observable environments
Nayak, S., Morrison Orozco, A., Have, M., Zhang, J., Thirumalai, V., Chen, D., Kapoor, A., Robinson, E., Gopalakrishnan, K., Harrison, J., et al · 2024
Later among the works it cites.
Synatra: Turning indirect knowledge into direct demonstrations for digital agents at scale
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Abuelsaad, T., Akkil, D., Dey, P., Jagmohan, A., Vempaty, A., and Kokku, R · 2024
Cited alongside, same era.
Digirl: Training in-the-wild device-control agents with autonomous reinforcement learning
Bai, H., Zhou, Y., Cemri, M., Pan, J., Suhr, A., Levine, S., and Kumar, A · 2024
Cited alongside, same era.
Web agents with world models: Learning and leveraging environment dynamics in web navigation
Chae, H., Kim, N., Ong, K. T.-i., Gwak, M., Song, G., Kim, J., Kim, S., Lee, D., and Yeo, J · 2024
Cited alongside, same era.
Mind2web: Towards a generalist agent for the web
Deng, X., Gu, Y., Zheng, B., Chen, S., Stevens, S., Wang, B., Sun, H., and Su, Y · 2024
Cited alongside, same era.
Tinyagent: Function calling at the edge
Erdogan, L. E., Lee, N., Jha, S., Kim, S., Tabrizi, R., Moon, S., Hooper, C., Anumanchipalli, G., Keutzer, K., and Gholami, A · 2024
Cited alongside, same era.
Geometric-averaged preference optimization for soft preference labels
Furuta, H., Lee, K.-H., Gu, S. S., Matsuo, Y., Faust, A., Zen, H., and Gur, I · 2024
Cited alongside, same era.
Is your llm secretly a world model of the internet? model-based planning for web agents
Gu, Y., Zheng, B., Gou, B., Zhang, K., Chang, C., Srivastava, S., Xie, Y., Qi, P., Sun, H., and Su, Y · 2024
Cited alongside, same era.
Smart-llm: Smart multi-agent robot task planning using large language models
Kannan, S. S., Venkatesh, V. L., and Min, B.-C · 2024
Cited alongside, same era.
Ou, T., Xu, F. F., Madaan, A., Liu, J., Lo, R., Sridhar, A., Sengupta, S., Roth, D., Neubig, G., and Zhou, S · 2024
Later among the works it cites.
Autonomous evaluation and refinement of digital agents
Pan, J., Zhang, Y., Tomlin, N., Zhou, Y., Levine, S., and Suhr, A · 2024
Later among the works it cites.
Large language models can self-improve at web agent tasks
Patel, A., Hofmarcher, M., Leoveanu-Condrei, C., Dinu, M.-C., Callison-Burch, C., and Hochreiter, S · 2024
Later among the works it cites.
Tinyclick: Single-turn agent for empowering gui automation
Pawlowski, P., Zawistowski, K., Lapacz, W., Skorupa, M., Wiacek, A., Postansque, S., and Hoscilowicz, J · 2024
Later among the works it cites.
Webrl: Training llm web agents via self-evolving online curriculum reinforcement learning
Qi, Z., Liu, X., Iong, I. L., Lai, H., Sun, X., Yang, X., Sun, J., Yang, Y., Yao, S., Zhang, T., et al · 2024
Later among the works it cites.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R., Sharma, A., Mitchell, E., Manning, C. D., Ermon, S., and Finn, C · 2024
Later among the works it cites.
Androidinthewild: A large-scale dataset for android device control
Rawles, C., Li, A., Rodriguez, D., Riva, O., and Lillicrap, T · 2024
Later among the works it cites.
Beyond browsing: Api-based web agents
Song, Y., Xu, F. F., Zhou, S., and Neubig, G · 2024
Later among the works it cites.
Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments
Xie, T., Zhang, D., Chen, J., Li, X., Zhao, S., Cao, R., Hua, T. J., Cheng, Z., Shin, D., Lei, F., et al · 2024
Later among the works it cites.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Guo, D., Yang, D., Zhang, H., Song, J., Zhang, R., Xu, R., Zhu, Q., Ma, S., Wang, P., Bi, X., et al · 2025
Closest in time.
Qwq-32b: Embracing the power of reinforcement learning, March 2025
Team, Q · 2025
Closest in time.