Fetching the paper…
Reading the bibliography…
Reasoning capabilities have significantly improved the performance of vision-language models (VLMs) in domains such as mathematical problem-solving, coding, and visual question-answering.
META-GUI: towards multi-modal conversational agents on mobile GUI
L. Sun, X. Chen, L. Chen, T. Dai, Z. Zhu, and K. Yu · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
J. Wei, X. Wang, D. Schuurmans, M. Bosma, B. Ichter, F. Xia, E. H. Chi, Q. V. Le, and D. Zhou · 2022
Earlier work this paper cites.
Androidinthewild: A large-scale dataset for android device control
C. Rawles, A. Li, D. Rodriguez, O. Riva, and T. Lillicrap · 2023
Earlier work this paper cites.
Enabling conversational interaction with mobile ui using large language models
B. Wang, G. Li, and Y. Li · 2023
Earlier work this paper cites.
Set-of-mark prompting unleashes extraordinary visual grounding in gpt-4v
J. Yang, H. Zhang, F. Li, X. Zou, C. Li, and J. Gao · 2023
Earlier work this paper cites.
Introducing computer use, a new claude 3.5 sonnet, and claude 3.5 haiku
Anthropic · 2024
Earlier work this paper cites.
Siri - apple
Apple · 2024
Earlier work this paper cites.
Seeclick: Harnessing GUI grounding for advanced visual GUI agents
K. Cheng, Q. Sun, Y. Chu, F. Xu, Y. Li, J. Zhang, and Z. Wu · 2024
Earlier work this paper cites.
Mobile-bench: An evaluation benchmark for llm-based mobile agents
S. Deng, W. Xu, H. Sun, W. Liu, T. Tan, J. Liu, A. Li, J. Luan, B. Wang, R. Yan, and S. Shang · 2024
Earlier work this paper cites.
Mobileviews: A large-scale mobile gui dataset
L. Gao, L. Zhang, S. Wang, S. Wang, Y. Li, and M. Xu · 2024
Earlier work this paper cites.
Google assistant, your own personal google
Google · 2024
Earlier work this paper cites.
Open-rag: Enhanced retrieval-augmented reasoning with open-source large language models
S. B. Islam, M. A. Rahman, K. Hossain, E. Hoque, S. Joty, and M. R. Parvez · 2024
Cited alongside, same era.
On the effects of data scale on ui control agents
W. Li, W. E. Bishop, A. Li, C. Rawles, F. Campbell-Ajala, D. Tyamagundlu, and O. Riva · 2024
Cited alongside, same era.
Small language models: Survey, measurements, and insights
Z. Lu, X. Li, D. Cai, R. Yi, F. Liu, X. Zhang, N. D. Lane, and M. Xu · 2024
Cited alongside, same era.
Androidworld: A dynamic benchmarking environment for autonomous agents
C. Rawles, S. Clinckemaillie, Y. Chang, J. Waltz, G. Lau, M. Fair, A. Li, W. Bishop, W. Li, F. Campbell-Ajala, et al · 2024
Cited alongside, same era.
Scaling llm test-time compute optimally can be more effective than scaling model parameters
New gemini app features, available to try at no cost
Google · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
D. Guo, D. Yang, H. Zhang, J. Song, R. Zhang, R. Xu, Q. Zhu, S. Ma, P. Wang, X. Bi, et al · 2025
Closest in time.
Search-o1: Agentic search-enhanced large reasoning models
X. Li, G. Dong, J. Jin, Y. Zhang, Y. Zhou, Y. Zhu, P. Zhang, and Z. Dou · 2025
Closest in time.
Introducing operator
OpenAI · 2025
Closest in time.
Agent s2
Simular · 2025
Closest in time.
Agentic reasoning: Reasoning llms with tools for the deep research
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
C. Snell, J. Lee, K. Xu, and A. Kumar · 2024
Cited alongside, same era.
Y. Wang, W. Chen, X. Han, X. Lin, H. Zhao, Y. Liu, B. Zhai, J. Yuan, Q. You, and H. Yang · 2024
Cited alongside, same era.
Autodroid: Llm-powered task automation in android
H. Wen, Y. Li, G. Liu, S. Zhao, T. Yu, T. J. Li, S. Jiang, Y. Liu, Y. Zhang, and Y. Liu · 2024
Cited alongside, same era.
Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments
T. Xie, D. Zhang, J. Chen, X. Li, S. Zhao, R. Cao, J. H. Toh, Z. Cheng, D. Shin, F. Lei, et al · 2024
Cited alongside, same era.
Android in the zoo: Chain-of-action-thought for GUI agents
J. Zhang, J. Wu, Y. Teng, M. Liao, N. Xu, X. Xiao, Z. Wei, and D. Tang · 2024
Cited alongside, same era.
You only look at screens: Multimodal chain-of-action agents
Z. Zhang and A. Zhang · 2024
Cited alongside, same era.
Claude’s extended thinking
Anthropic
Cited in the paper.
Claude 3.7 sonnet and claude code
Anthropic
Cited in the paper.
J. Wu, J. Zhu, and Y. Liu · 2025
Closest in time.
Grok 3 beta — the age of reasoning agents
X.ai · 2025
Closest in time.
Every software as an agent: Blueprint and case study
M. Xu · 2025
Closest in time.
Resource-efficient algorithms and systems of foundation models: A survey
M. Xu, D. Cai, W. Yin, S. Wang, X. Jin, and X. Liu · 2025
Closest in time.
Api agents vs. gui agents: Divergence and convergence, 2025
C. Zhang, S. He, L. Li, S. Qin, Y. Kang, Q. Lin, and D. Zhang · 2025
Closest in time.