Fetching the paper…
Reading the bibliography…
The utilisation of foundation models as smartphone assistants, termed app agents, is a critical research challenge.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Willams, R. J · 1992
Earlier work this paper cites.
Fixing weight decay regularization in adam
Loshchilov, I., Hutter, F., et al · 2017
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Earlier work this paper cites.
Androidenv: A reinforcement learning platform for android
Toyama, D., Hamel, P., Gergely, A., Comanici, G., Glaese, A., Ahmed, Z., Jackson, T., Mourad, S., and Precup, D · 2021
Earlier work this paper cites.
Vip: Towards universal visual reward and representation via value-implicit pre-training
Ma, Y. J., Sodhani, S., Jayaraman, D., Bastani, O., Kumar, V., and Zhang, A · 2022
Earlier work this paper cites.
Vision-language models as a source of rewards
Chan, H., Mnih, V., Behbahani, F., Laskin, M., Wang, L., Pardo, F., Gazeau, M., Sahni, H., Horgan, D., Baumli, K., et al · 2023
Earlier work this paper cites.
Reinforced self-training (rest) for language modeling
Gulcehre, C., Paine, T. L., Srinivasan, S., Konyushkova, K., Weerts, L., Sharma, A., Siddhant, A., Ahern, A., Wang, M., Gu, C., et al · 2023
Earlier work this paper cites.
Androidinthewild: A large-scale dataset for android device control
Rawles, C., Li, A., Rodriguez, D., Riva, O., and Lillicrap, T · 2023
Earlier work this paper cites.
Autodroid: Llm-powered task automation in android
Wen, H., Li, Y., Liu, G., Zhao, S., Yu, T., Li, T. J.-J., Jiang, S., Liu, Y., Zhang, Y., and Liu, Y · 2023
Cited alongside, same era.
Appagent: Multimodal agents as smartphone users
Yang, Z., Liu, J., Han, Y., Chen, X., Huang, Z., Fu, B., and Yu, G · 2023
Cited alongside, same era.
Digirl: Training in-the-wild device-control agents with autonomous reinforcement learning
Bai, H., Zhou, Y., Cemri, M., Pan, J., Suhr, A., Levine, S., and Kumar, A · 2024
Cited alongside, same era.
Paligemma: A versatile 3b vlm for transfer
Beyer, L., Steiner, A., Pinto, A. S., Kolesnikov, A., Wang, X., Salz, D., Neumann, M., Alabdulmohsin, I., Tschannen, M., Bugliarello, E., et al · 2024
Cited alongside, same era.
Spa-bench: A comprehensive benchmark for smartphone agent evaluation
Chen, J., Yuen, D., Xie, B., Yang, Y., Chen, G., Wu, Z., Yixing, L., Zhou, X., Liu, W., Wang, S., et al · 2024
On the effects of data scale on computer control agents
Li, W., Bishop, W., Li, A., Rawles, C., Campbell-Ajala, F., Tyamagundlu, D., and Riva, O · 2024
Later among the works it cites.
Comprehensive cognitive llm agent for smartphone gui automation
Ma, X., Zhang, Z., and Zhao, H · 2024
Later among the works it cites.
Agent q: Advanced reasoning and learning for autonomous ai agents
Putta, P., Mills, E., Garg, N., Motwani, S., Finn, C., Garg, D., and Rafailov, R · 2024
Later among the works it cites.
Androidworld: A dynamic benchmarking environment for autonomous agents
Rawles, C., Clinckemaillie, S., Chang, Y., Waltz, J., Lau, G., Fair, M., Li, A., Bishop, W., Li, W., Campbell-Ajala, F., et al · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Lightweight neural app control
Christianos, F., Papoudakis, G., Coste, T., Hao, J., Wang, J., and Shao, K · 2024
Cited alongside, same era.
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al · 2024
Cited alongside, same era.
Navigating the digital world as humans do: Universal visual grounding for gui agents
Gou, B., Wang, R., Zheng, B., Xie, Y., Chang, C., Shu, Y., Sun, H., and Su, Y · 2024
Cited alongside, same era.
Mobilegpt: Augmenting llm with human-like app memory for mobile task automation
Lee, S., Choi, J., Lee, J., Wasi, M. H., Choi, H., Ko, S., Oh, S., and Shin, I · 2024
Cited alongside, same era.
Wang, J., Xu, H., Jia, H., Zhang, X., Yan, M., Shen, W., Zhang, J., Huang, F., and Sang, J
Cited in the paper.
Mobile-agent: Autonomous multi-modal mobile device agent with visual perception
Wang, J., Xu, H., Ye, J., Yan, M., Shen, W., Zhang, J., Huang, F., and Sang, J
Cited in the paper.
Distrl: An asynchronous distributed reinforcement learning framework for on-device control agents
Wang, T., Wu, Z., Liu, J., Hao, J., Wang, J., and Shao, K
Cited in the paper.
Song, Z., Li, Y., Fang, M., Chen, Z., Shi, Z., Huang, Y., and Chen, L · 2024
Later among the works it cites.
Oscar: Operating system control via state-aware reasoning and re-planning
Wang, X. and Liu, B · 2024
Later among the works it cites.
Gpt-4v (ision) is a generalist web agent, if grounded
Zheng, B., Gou, B., Kil, J., Sun, H., and Su, Y · 2024
Later among the works it cites.
Infiguiagent: A multimodal generalist gui agent with native reasoning and reflection
Liu, Y., Li, P., Wei, Z., Xie, C., Hu, X., Xu, X., Zhang, S., Han, X., Yang, H., and Wu, F · 2025
Closest in time.