Fetching the paper…
Reading the bibliography…
The emergence of multimodal LLM-based agents (MLAs) has transformed interaction paradigms by seamlessly integrating vision, language, action and dynamic environments, enabling unprecedented autonomous capabilities across GUI applications ranging from web automation to mobile systems.
W. Johannsen, “Elemente der exakten erblichkeitslehre. 1909,” Gustav Fischer, Jena , 1913
1913
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
T. Sumers, S. Yao, K. Narasimhan, and T. Griffiths, “Cognitive architectures for language agents,” Transactions on Machine Learning Research , 2023
2023
Earlier work this paper cites.
D. Surís, S. Menon, and C. Vondrick, “Vipergpt: Visual inference via python execution for reasoning,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 11 888–11 898
2023
Earlier work this paper cites.
P. Lu, B. Peng, H. Cheng, M. Galley, K.-W. Chang, Y. N. Wu, S.-C. Zhu, and J. Gao, “Chameleon: Plug-and-play compositional reasoning with large language models,” Advances in Neural Information Processing Systems , vol. 36, pp. 43 447–43 478, 2023
2023
Earlier work this paper cites.
Y. Shen, K. Song, X. Tan, D. Li, W. Lu, and Y. Zhuang, “Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face,” Advances in Neural Information Processing Systems , vol. 36, pp. 38 154–38 180, 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
J. Betker, G. Goh, L. Jing, T. Brooks, J. Wang, L. Li, L. Ouyang, J. Zhuang, J. Lee, Y. Guo et al. , “Improving image generation with better captions,” Computer Science. https://cdn. openai. com/papers/dall-e-3. pdf , vol. 2, no. 3, p. 8, 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2024
Earlier work this paper cites.
DeepMind, “Gemini 2.0 flash model card,” Google, Tech. Rep., 2024
2024
Earlier work this paper cites.
Anthropic, “The claude 3 model family: Opus, sonnet, haiku,” 2024. [Online]. Available: https://www-cdn.anthropic.com/de8ba9b01c9ab7cbabf5c33b80b7bbc618857627/Model_Card_Claude_3.pdf
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
S. Liu, J. Chen, S. Ruan, H. Su, and Z. Yin, “Exploring the robustness of decision-level through adversarial attacks on llm-based embodied models,” in Proceedings of the 32nd ACM International Conference on Multimedia , 2024, pp. 8120–8128
2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2024
Cited alongside, same era.
2024
Cited alongside, same era.
C. Ma, J. Zhang, Z. Zhu, C. Yang, Y. Yang, Y. Jin, Z. Lan, L. Kong, and J. He, “Agentboard: An analytical evaluation board of multi-turn llm agents,” NeurIPS 2024 Datasets and Benchmarks Track , 2024
2024
Cited alongside, same era.
2024
Cited alongside, same era.
G. Pappas, A. Robey, and P. Agrawal, “Ai-powered robots can be tricked into acts of violence,” Wired Magazine , 2024
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Later among the works it cites.
2024
Later among the works it cites.
OpenBMB, “Large multi-modal models for strong performance and efficient deployment,” 2024. [Online]. Available: https://github.com/OpenBMB/OmniLMM
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
M. Abdin, J. Aneja, and H. S. Behl, “Phi-4 technical report,” arXiv preprint arXiv: 2412.08905 , 2024
2024
Later among the works it cites.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
P. Kumar, E. Lau, S. Vijayakumar, T. Trinh, E. T. Chang, V. Robinson, S. Zhou, M. Fredrikson, S. M. Hendryx, S. Yue et al. , “Aligned llms are not aligned browser agents,” in The Thirteenth International Conference on Learning Representations , 2025
2025
Closest in time.
Y. Gong, D. Ran, J. Liu, C. Wang, T. Cong, A. Wang, S. Duan, and X. Wang, “Figstep: Jailbreaking large vision-language models via typographic visual prompts,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 39, no. 22, 2025, pp. 23 951–23 959
2025
Closest in time.
2025
Closest in time.
C. H. Wu, R. R. Shah, J. Y. Koh, R. Salakhutdinov, D. Fried, and A. Raghunathan, “Dissecting adversarial robustness of multimodal lm agents,” in The Thirteenth International Conference on Learning Representations , 2025
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.