Fetching the paper…
Reading the bibliography…
Mobile device operation tasks are increasingly becoming a popular multi-modal AI application scenario.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. (2020) · 1901
Earlier work this paper cites.
Webshop: Towards scalable real-world web interaction with grounded language agents
Yao, S., Chen, H., Yang, J., and Narasimhan, K. R. (2022) · 2022
Earlier work this paper cites.
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al. (2023) · 2023
Earlier work this paper cites.
Qwen-vl: A frontier large vision-language model with versatile abilities
Bai, J., Bai, S., Yang, S., Wang, S., Tan, S., Wang, P., Lin, J., Zhou, C., and Zhou, J. (2023) · 2023
Earlier work this paper cites.
Minigpt-v2: Large language model as a unified interface for vision-language multi-task learning
Chen, J., Li, D. Z. X. S. X., Zhang, Z. L. P., Xiong, R. K. V. C. Y., and Elhoseiny, M. (2023) · 2023
Earlier work this paper cites.
Holistic analysis of hallucination in gpt-4v (ision): Bias and interference challenges
Cui, C., Zhou, Y., Yang, X., Wu, S., Zhang, L., Zou, J., and Yao, H. (2023) · 2023
Earlier work this paper cites.
Instructblip: Towards general-purpose vision-language models with instruction tuning
Dai, W., Li, J., Li, D., Tiong, A. M. H., Zhao, J., Wang, W., Li, B., Fung, P., and Hoi, S. (2023) · 2023
Earlier work this paper cites.
Mind2web: Towards a generalist agent for the web
Deng, X., Gu, Y., Zheng, B., Chen, S., Stevens, S., Wang, B., Sun, H., and Su, Y. (2023) · 2023
Earlier work this paper cites.
Cogagent: A visual language model for gui agents
Hong, W., Wang, W., Lv, Q., Xu, J., Yu, W., Ji, J., Wang, Y., Wang, Z., Zhang, Y., Li, J., Xu, B., Dong, Y., Ding, M., and Tang, J. (2023) · 2023
Earlier work this paper cites.
mplug-paperowl: Scientific diagram analysis with the multimodal large language model
Hu, A., Shi, Y., Xu, H., Ye, J., Ye, Q., Yan, M., Li, C., Qian, Q., Zhang, J., and Huang, F. (2023) · 2023
Earlier work this paper cites.
OpenAI (2023) · 2023
Earlier work this paper cites.
Generative agents: Interactive simulacra of human behavior
Park, J. S., O’Brien, J., Cai, C. J., Morris, M. R., Liang, P., and Bernstein, M. S. (2023) · 2023
Earlier work this paper cites.
Multi-agent collaboration: Harnessing the power of intelligent llm agents
Talebirad, Y. and Nadiri, A. (2023) · 2023
Cited alongside, same era.
Autodroid: Llm-powered task automation in android
Wen, H., Li, Y., Liu, G., Zhao, S., Yu, T., Li, T. J.-J., Jiang, S., Liu, Y., Zhang, Y., and Liu, Y. (2023) · 2023
Cited alongside, same era.
Autogen: Enabling next-gen llm applications via multi-agent conversation framework
Wu, Q., Bansal, G., Zhang, J., Wu, Y., Zhang, S., Zhu, E., Li, B., Jiang, L., Zhang, X., and Wang, C. (2023) · 2023
Cited alongside, same era.
Gpt-4v in wonderland: Large multimodal models for zero-shot smartphone gui navigation
Yan, A., Yang, Z., Zhu, W., Lin, K., Li, L., Wang, J., Yang, J., Zhong, Y., McAuley, J., Gao, J., Liu, Z., and Wang, L. (2023) · 2023
Cited alongside, same era.
Analyzing and mitigating object hallucination in large vision-language models
MetaGPT: Meta programming for a multi-agent collaborative framework
Hong, S., Zhuge, M., Chen, J., Zheng, X., Cheng, Y., Wang, J., Zhang, C., Wang, Z., Yau, S. K. S., Lin, Z., Zhou, L., Ran, C., Xiao, L., Wu, C., and Schmidhuber, J. (2024) · 2024
Closest in time.
mplug-docowl 1.5: Unified structure learning for ocr-free document understanding
Hu, A., Xu, H., Ye, J., Yan, M., Zhang, L., Zhang, B., Li, C., Zhang, J., Jin, Q., Huang, F., et al. (2024) · 2024
Closest in time.
Welfare diplomacy: Benchmarking language model cooperation
Mukobi, G., Erlebach, H., Lauffer, N., Hammond, L., Chan, A., and Clifton, J. (2024) · 2024
Closest in time.
Small llms are weak tool learners: A multi-llm agent
Shen, W., Li, C., Chen, H., Yan, M., Quan, X., Chen, H., Zhang, J., and Huang, F. (2024) · 2024
Closest in time.
DebateGPT: Fine-tuning large language models with multi-agent debate supervision
Subramaniam, V., Torralba, A., and Li, S. (2024) · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zhou, Y., Cui, C., Yoon, J., Zhang, L., Deng, Z., Finn, C., Bansal, M., and Yao, H. (2023) · 2023
Cited alongside, same era.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
Zhu, D., Chen, J., Shen, X., Li, X., and Elhoseiny, M. (2023) · 2023
Cited alongside, same era.
LLM-deliberation: Evaluating LLMs with interactive multi-agent negotiation game
Abdelnabi, S., Gomaa, A., Sivaprasad, S., Schönherr, L., and Fritz, M. (2024) · 2024
Cited alongside, same era.
Chateval: Towards better LLM-based evaluators through multi-agent debate
Chan, C.-M., Chen, W., Su, Y., Yu, J., Xue, W., Zhang, S., Fu, J., and Liu, Z. (2024) · 2024
Cited alongside, same era.
Agentverse: Facilitating multi-agent collaboration and exploring emergent behaviors
Chen, W., Su, Y., Zuo, J., Yang, C., Yuan, C., Chan, C.-M., Yu, H., Lu, Y., Hung, Y.-H., Qian, C., Qin, Y., Cong, X., Xie, R., Liu, Z., Sun, M., and Zhou, J. (2024) · 2024
Cited alongside, same era.
Seeclick: Harnessing gui grounding for advanced visual gui agents
Cheng, K., Sun, Q., Chu, Y., Xu, F., Li, Y., Zhang, J., and Wu, Z. (2024) · 2024
Cited alongside, same era.
Detecting and preventing hallucinations in large vision language models
Gunjal, A., Yin, J., and Bas, E. (2024) · 2024
Cited alongside, same era.
A real-world webagent with planning, long context understanding, and program synthesis
Gur, I., Furuta, H., Huang, A. V., Safdari, M., Matsuo, Y., Eck, D., and Faust, A. (2024) · 2024
Cited alongside, same era.
Closest in time.
Chain-of-discussion: A multi-model framework for complex evidence-based question answering
Tao, M., Zhao, D., and Feng, Y. (2024) · 2024
Closest in time.
Mobile-agent: Autonomous multi-modal mobile device agent with visual perception
Wang, J., Xu, H., Ye, J., Yan, M., Shen, W., Zhang, J., Huang, F., and Sang, J. (2024) · 2024
Closest in time.
Autogen: Enabling next-gen LLM applications via multi-agent conversation
Wu, Q., Bansal, G., Zhang, J., Wu, Y., Li, B., Zhu, E., Jiang, L., Zhang, X., Zhang, S., Liu, J., Awadallah, A. H., White, R. W., Burger, D., and Wang, C. (2024) · 2024
Closest in time.
Language agents with reinforcement learning for strategic play in the werewolf game
Xu, Z., Yu, C., Fang, F., Wang, Y., and Wu, Y. (2024) · 2024
Closest in time.
Expel: Llm agents are experiential learners
Zhao, A., Huang, D., Xu, Q., Lin, M., Liu, Y.-J., and Huang, G. (2024) · 2024
Closest in time.
Gpt-4v(ision) is a generalist web agent, if grounded
Zheng, B., Gou, B., Kil, J., Sun, H., and Su, Y. (2024) · 2024
Closest in time.