Fetching the paper…
Reading the bibliography…
We present an UI agent for user interface (UI) interaction tasks, using Vision-Language Model Florence-2-Base.
J. P. A. Ioannidis, “Why most published research findings are false,” PLOS Medicine , vol. 2, no. 8, p. null, 08 2005. [Online]. Available: https://doi.org/10.1371/journal.pmed.0020124
2005
Earlier work this paper cites.
S. T. Ziliak and D. N. McCloskey, The Cult of Statistical Significance: How the Standard Error Costs Us Jobs, Justice, and Lives . University of Michigan Press, 2008. [Online]. Available: http://www.jstor.org/stable/10.3998/mpub.186351
2008
Earlier work this paper cites.
2018
Earlier work this paper cites.
2022
Earlier work this paper cites.
B. Xiao et al. , “Florence-2: Advancing a unified representation for a variety of vision tasks,” 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pp. 4818–4829, 2023. [Online]. Available: https://api.semanticscholar.org/CorpusID:265128818
2023
Earlier work this paper cites.
G. Li and Y. Li, “Spotlight: Mobile UI understanding using vision-language models with a focus,” in The Eleventh International Conference on Learning Representations , 2023. [Online]. Available: https://openreview.net/forum?id=9yE2xEj0BH7
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
K. Lee et al. , “Pix2struct: screenshot parsing as pretraining for visual language understanding,” in Proceedings of the 40th International Conference on Machine Learning , ser. ICML’23. JMLR.org, 2023
2023
Earlier work this paper cites.
C.-Y. Hsieh et al. , “Distilling step-by-step! outperforming larger language models with less training data and smaller model sizes,” in Findings of the Association for Computational Linguistics: ACL 2023 . Toronto, Canada: Association for Computational Linguistics, Jul. 2023, pp. 8003–8017. [Online]. Available: https://aclanthology.org/2023.findings-acl.507
2023
Earlier work this paper cites.
Z. Chen et al. , “Intern vl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks,” 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pp. 24 185–24 198, 2023. [Online]. Available: https://api.semanticscholar.org/CorpusID:266521410
2023
Earlier work this paper cites.
2024
Earlier work this paper cites.
G. Baechler et al. , “Screenai: A vision-language model for ui and infographics understanding,” 08 2024, pp. 3058–3068
2024
Earlier work this paper cites.
2024
Cited alongside, same era.
W. Hong et al. , “Cogagent: A visual language model for gui agents,” in 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . Los Alamitos, CA, USA: IEEE Computer Society, jun 2024, pp. 14 281–14 290. [Online]. Available: https://doi.ieeecomputersociety.org/10.1109/CVPR52733.2024.01354
2024
Cited alongside, same era.
Z. Zhang and A. Zhang, “You only look at screens: Multimodal chain-of-action agents,” Bangkok, Thailand and virtual meeting, pp. 3132–3149, Aug. 2024. [Online]. Available: https://aclanthology.org/2024.findings-acl.186
2024
Cited alongside, same era.
H. You et al. , “Ferret: Refer and ground anything anywhere at any granularity,” in The Twelfth International Conference on Learning Representations , 2024. [Online]. Available: https://openreview.net/forum?id=2msbbX3ydD
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
D. Jeffries and K. Team, “Wave ui dataset,” 2024. [Online]. Available: https://huggingface.co/datasets/agentsea/wave-ui
2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2024
Cited alongside, same era.
H. Zhang et al. , “Ferret-v2: An improved baseline for referring and grounding with large language models,” in First Conference on Language Modeling , 2024. [Online]. Available: https://openreview.net/forum?id=EEPBOB2Xww
2024
Cited alongside, same era.
X. Deng et al. , “Mind2web: towards a generalist agent for the web,” in Proceedings of the 37th International Conference on Neural Information Processing Systems , ser. NIPS ’23. Red Hook, NY, USA: Curran Associates Inc., 2024
2024
Cited alongside, same era.
I. Gur et al. , “A real-world webagent with planning, long context understanding, and program synthesis,” in The Twelfth International Conference on Learning Representations , 2024. [Online]. Available: https://openreview.net/forum?id=9JQtrumvg8
2024
Cited alongside, same era.
K. Cheng et al. , “Seeclick: Harnessing gui grounding for advanced visual gui agents,” in Annual Meeting of the Association for Computational Linguistics , 2024. [Online]. Available: https://api.semanticscholar.org/CorpusID:267069082
2024
Cited alongside, same era.
Y. Gao et al. , “Enhancing vision-language pre-training with rich supervisions,” 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pp. 13 480–13 491, 2024. [Online]. Available: https://api.semanticscholar.org/CorpusID:268253650
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Closest in time.
Q. Lu et al. , “Gui odyssey: A comprehensive dataset for cross-app gui navigation on mobile devices,” 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.