Fetching the paper…
Reading the bibliography…
Smartphones have become indispensable in modern life, yet navigating complex tasks on mobile devices often remains frustrating.
Acquisition of cognitive skill
Anderson, J. R · 1982
Earlier work this paper cites.
Structure and function of declarative and nondeclarative memory systems
Squire, L. R. and Zola, S. M · 1996
Earlier work this paper cites.
Episodic memory: From mind to brain
Tulving, E · 2002
Earlier work this paper cites.
Real-time scene text detection with differentiable binarization
Liao, M., Wan, Z., Yao, C., Chen, K., and Bai, X · 2020
Earlier work this paper cites.
Large language models can self-improve
Huang, J., Gu, S. S., Hou, L., Wu, Y., Wang, X., Yu, H., and Han, J · 2022
Earlier work this paper cites.
Qwen-VL: A Frontier Large Vision-Language Model with Versatile Abilities
Bai, J., Bai, S., Yang, S., Wang, S., Tan, S., Wang, P., Lin, J., Zhou, C., and Zhou, J · 2023
Earlier work this paper cites.
Large language models as tool makers
Cai, T., Wang, X., Ma, T., Chen, X., and Zhou, D · 2023
Earlier work this paper cites.
Mind2web: Towards a generalist agent for the web
Deng, X., Gu, Y., Zheng, B., Chen, S., Stevens, S., Wang, B., Sun, H., and Su, Y · 2023
Earlier work this paper cites.
Cogagent: A visual language model for gui agents, 2023
Hong, W., Wang, W., Lv, Q., Xu, J., Yu, W., Ji, J., Wang, Y., Wang, Z., Dong, Y., Ding, M., and Tang, J · 2023
Earlier work this paper cites.
Grounding DINO: marrying DINO with grounded pre-training for open-set object detection
Liu, S., Zeng, Z., Ren, T., Li, F., Zhang, H., Yang, J., Li, C., Yang, J., Su, H., Zhu, J., and Zhang, L · 2023
Earlier work this paper cites.
Creator: Tool creation for disentangling abstract and concrete reasoning of large language models
Qian, C., Han, C., Fung, Y. R., Qin, Y., Liu, Z., and Ji, H · 2023
Cited alongside, same era.
Wang, Z., Mao, S., Wu, W., Ge, T., Wei, F., and Ji, H · 2023
Cited alongside, same era.
Craft: Customizing llms by creating and retrieving from specialized toolsets
Yuan, L., Chen, Y., Wang, X., Fung, Y. R., Peng, H., and Ji, H · 2023
Cited alongside, same era.
Appagent: Multimodal agents as smartphone users, 2023
Zhang, C., Yang, Z., Liu, J., Han, Y., Chen, X., Huang, Z., Fu, B., and Yu, G · 2023
Cited alongside, same era.
Claude 3.5 Sonnet, 2024
Anthropic · 2024
Cited alongside, same era.
Androidworld: A dynamic benchmarking environment for autonomous agents
Rawles, C., Clinckemaillie, S., Chang, Y., Waltz, J., Lau, G., Fair, M., Li, A., Bishop, W., Li, W., Campbell-Ajala, F., et al · 2024
Later among the works it cites.
Infogent: An agent-based framework for web information aggregation
Reddy, R. G., Mukherjee, S., Kim, J., Wang, Z., Hakkani-Tur, D., and Ji, H · 2024
Later among the works it cites.
Reflexion: Language agents with verbal reinforcement learning
Shinn, N., Cassano, F., Gopinath, A., Narasimhan, K., and Yao, S · 2024
Later among the works it cites.
Towards general computer control: A multimodal agent for red dead redemption ii as a case study
Tan, W., Ding, Z., Zhang, W., Li, B., Zhou, B., Yue, J., Xia, H., Jiang, J., Zheng, L., Xu, X., et al · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Webvoyager: Building an end-to-end web agent with large multimodal models
He, H., Yao, W., Ma, K., Yu, W., Dai, Y., Zhang, H., Lan, Z., and Yu, D · 2024
Cited alongside, same era.
Appagent v2: Advanced agent for flexible mobile interactions
Li, Y., Zhang, C., Yang, W., Fu, B., Cheng, P., Chen, X., Chen, L., and Wei, Y · 2024
Cited alongside, same era.
Self-refine: Iterative refinement with self-feedback
Madaan, A., Tandon, N., Gupta, P., Hallinan, S., Gao, L., Wiegreffe, S., Alon, U., Dziri, N., Prabhumoye, S., Yang, Y., et al · 2024
Cited alongside, same era.
Nguyen, D., Chen, J., Wang, Y., Wu, G., Park, N., Hu, Z., Lyu, H., Wu, J., Aponte, R., Xia, Y., et al · 2024
Cited alongside, same era.
GPT-4o System Card, 2024
OpenAI · 2024
Cited alongside, same era.
Autoglm: Autonomous foundation agents for guis
Liu, X., Qin, B., Liang, D., Dong, G., Lai, H., Zhang, H., Zhao, H., Iong, I. L., Sun, J., Wang, J., et al
Cited in the paper.
Visualagentbench: Towards large multimodal models as visual foundation agents
Liu, X., Zhang, T., Gu, Y., Iong, I. L., Xu, Y., Song, X., Zhang, S., Lai, H., Liu, X., Zhao, H., et al
Cited in the paper.
Tao, Z., Lin, T.-E., Chen, X., Li, H., Wu, Y., Li, Y., Jin, Z., Huang, F., Tao, D., and Zhou, J · 2024
Later among the works it cites.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Team, G., Georgiev, P., Lei, V. I., Burnell, R., Bai, L., Gulati, A., Tanzer, G., Vincent, D., Pan, Z., Wang, S., et al · 2024
Later among the works it cites.
Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments
Xie, T., Zhang, D., Chen, J., Li, X., Zhao, S., Cao, R., Hua, T. J., Cheng, Z., Shin, D., Lei, F., et al · 2024
Later among the works it cites.
Assistantbench: Can web agents solve realistic and time-consuming tasks?, 2024
Yoran, O., Amouyal, S. J., Malaviya, C., Bogin, B., Press, O., and Berant, J · 2024
Later among the works it cites.
UFO: A UI-Focused Agent for Windows OS Interaction
Zhang, C., Li, L., He, S., Zhang, X., Qiao, B., Qin, S., Ma, M., Kang, Y., Lin, Q., Rajmohan, S., Zhang, D., and Zhang, Q · 2024
Later among the works it cites.
Gpt-4v(ision) is a generalist web agent, if grounded
Zheng, B., Gou, B., Kil, J., Sun, H., and Su, Y · 2024
Later among the works it cites.