Fetching the paper…
Reading the bibliography…
This paper investigates multimodal agents, in particular, OpenAI's Computer-User Agent (CUA), trained to control and complete tasks through a standard computer interface, similar to humans.
Visual search: How do we find what we are looking for?
Wolfe, J. M. (2020) · 2020
Earlier work this paper cites.
The sudden rise of wordle
Benveniste, A. (2022) · 2022
Earlier work this paper cites.
Language models (mostly) know what they know
Kadavath, S., Conerly, T., Askell, A., Henighan, T., Drain, D., Perez, E., Schiefer, N., Hatfield-Dodds, Z., DasSarma, N., Tran-Johnson, E., et al. (2022) · 2022
Earlier work this paper cites.
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al. (2023) · 2023
Earlier work this paper cites.
The reversal curse: Llms trained on” a is b” fail to learn” b is a”
Berglund, L., Tong, M., Kaufmann, M., Balesni, M., Stickland, A. C., Korbak, T., and Evans, O. (2023) · 2023
Earlier work this paper cites.
Visual instruction tuning
Liu, H., Li, C., Wu, Q., and Lee, Y. J. (2023) · 2023
Earlier work this paper cites.
Towards understanding sycophancy in language models
Sharma, M., Tong, M., Korbak, T., Duvenaud, D., Askell, A., Bowman, S. R., Cheng, N., Durmus, E., Hatfield-Dodds, Z., Johnston, S. R., et al. (2023) · 2023
Earlier work this paper cites.
What makes for good visual tokenizers for large language models?
Wang, G., Ge, Y., Ding, X., Kankanhalli, M., and Shan, Y. (2023) · 2023
Earlier work this paper cites.
V*: Guided Visual Search as a core mechanism in multimodal LLMs
Wu, P. and Xie, S. (2023) · 2023
Cited alongside, same era.
Webarena: A realistic web environment for building autonomous agents
Zhou, S., Xu, F. F., Zhu, H., Zhou, X., Lo, R., Sridhar, A., Cheng, X., Ou, T., Bisk, Y., Fried, D., et al. (2023) · 2023
Cited alongside, same era.
Agent s: An open agentic framework that uses computers like a human
Agashe, S., Han, J., Gan, S., Yang, J., Li, A., and Wang, X. E. (2024) · 2024
Cited alongside, same era.
Introducing computer use
Anthropic (2024) · 2024
Cited alongside, same era.
Language models do hard arithmetic tasks easily and hardly do easy arithmetic tasks
Gambardella, A., Iwasawa, Y., and Matsuo, Y. (2024) · 2024
Cited alongside, same era.
Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments
Xie, T., Zhang, D., Chen, J., Li, X., Zhao, S., Cao, R., Hua, T. J., Cheng, Z., Shin, D., Lei, F., et al. (2024) · 2024
Later among the works it cites.
Aguvis: Unified pure vision agents for autonomous gui interaction
Xu, Y., Wang, Z., Wang, J., Lu, D., Xie, T., Saha, A., Sahoo, D., Yu, T., and Xiong, C. (2024) · 2024
Later among the works it cites.
Janus-pro: Unified multimodal understanding and generation with data and model scaling
Chen, X., Wu, Z., Liu, X., Pan, Z., Liu, W., Xie, Z., Yu, X., and Ruan, C. (2025) · 2025
Closest in time.
Investigating truthfulness issues in a pre-release o3 model
Chowdhury, N., Johnson, D., Huang, V., Steinhardt, J., and Schwettmann, S. (2025) · 2025
Closest in time.
Llama 4 models
MetaAI (2025) · 2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Webvoyager: Building an end-to-end web agent with large multimodal models
He, H., Yao, W., Ma, K., Yu, W., Dai, Y., Zhang, H., Lan, Z., and Yu, D. (2024) · 2024
Cited alongside, same era.
Taming overconfidence in llms: Reward calibration in rlhf
Leng, J., Huang, C., Zhu, B., and Huang, J. (2024) · 2024
Cited alongside, same era.
Pixtral large
MistralAI (2024) · 2024
Cited alongside, same era.
Browsecomp: a benchmark for browsing agents
OpenAI (2025a)
Cited in the paper.
Computer-using agent
OpenAI (2025b)
Cited in the paper.
Introducing operator: Our first ai agent that can use computers
OpenAI (2025c)
Cited in the paper.
Petrov, I., Dekoninck, J., Baltadzhiev, L., Drencheva, M., Minchev, K., Balunović, M., Jovanović, N., and Vechev, M. (2025) · 2025
Closest in time.
Ui-tars: Pioneering automated gui interaction with native agents
Qin, Y., Ye, Y., Fang, J., Wang, H., Liang, S., Tian, S., Zhang, J., Li, J., Li, Y., Huang, S., et al. (2025) · 2025
Closest in time.
Wang, B. and Sun, H. (2025) · 2025
Closest in time.