Fetching the paper…
Reading the bibliography…
GUIs have long been central to human-computer interaction, providing an intuitive and visually-driven way to access and interact with digital systems.
C. E. Shannon, “Prediction and entropy of printed english,” Bell system technical journal , vol. 30, no. 1, pp. 50–64, 1951
1951
Earlier work this paper cites.
R. Koo and S. Toueg, “Checkpointing and rollback-recovery for distributed systems,” IEEE Transactions on Software Engineering , vol. SE-13, pp. 23–31, 1986. [Online]. Available: https://api.semanticscholar.org/CorpusID:206777989
1986
Earlier work this paper cites.
M. L. Puterman, “Markov decision processes,” Handbooks in operations research and management science , vol. 2, pp. 331–434, 1990
1990
Earlier work this paper cites.
W. B. Cavnar, J. M. Trenkle et al. , “N-gram-based text categorization,” in Proceedings of SDAIR-94, 3rd annual symposium on document analysis and information retrieval , vol. 161175. Ann Arbor, Michigan, 1994, p. 14
1994
Earlier work this paper cites.
E. Gamma, “Design patterns: elements of reusable object-oriented software,” Person Education Inc , 1995
1995
Earlier work this paper cites.
L. P. Kaelbling, M. L. Littman, and A. W. Moore, “Reinforcement learning: A survey,” Journal of artificial intelligence research , vol. 4, pp. 237–285, 1996
1996
Earlier work this paper cites.
S. Hochreiter, “Long short-term memory,” Neural Computation MIT-Press , 1997
1997
Earlier work this paper cites.
B. J. Jansen, “The graphical user interface,” ACM SIGCHI Bull. , vol. 30, pp. 22–26, 1998. [Online]. Available: https://api.semanticscholar.org/CorpusID:18416305
1998
Earlier work this paper cites.
S. Berkovits, J. D. Guttman, and V. Swarup, “Authentication for mobile agents,” in Mobile Agents and Security , 1998. [Online]. Available: https://api.semanticscholar.org/CorpusID:13987376
1998
Earlier work this paper cites.
J. Steven, P. Chandra, B. Fleck, and A. Podgurski, “jrapture: A capture/replay tool for observation-based testing,” SIGSOFT Softw. Eng. Notes , vol. 25, no. 5, p. 158–167, Aug. 2000. [Online]. Available: https://doi.org/10.1145/347636.348993
2000
Earlier work this paper cites.
L. R. Medsker, L. Jain et al. , “Recurrent neural networks,” Design and Applications , vol. 5, no. 64-67, p. 2, 2001
2001
Earlier work this paper cites.
A. M. Memon, M. E. Pollack, and M. L. Soffa, “Hierarchical gui test case generation using automated planning,” IEEE transactions on software engineering , vol. 27, no. 2, pp. 144–155, 2001
2001
Earlier work this paper cites.
B. Sierkowski, “Achieving web accessibility,” in Proceedings of the 30th annual ACM SIGUCCS conference on User services , 2002, pp. 288–291
2002
Earlier work this paper cites.
M. A. Boshart and M. J. Kosa, “Growing a gui from an xml tree,” ACM SIGCSE Bulletin , vol. 35, no. 3, pp. 223–223, 2003
2003
Earlier work this paper cites.
A. Memon, I. Banerjee, N. Hashmi, and A. Nagarajan, “Dart: a framework for regression testing "nightly/daily builds" of gui applications,” in International Conference on Software Maintenance, 2003. ICSM 2003. Proceedings. , 2003, pp. 410–419
2003
Earlier work this paper cites.
A. M. Memon, I. Banerjee, and A. Nagarajan, “Gui ripping: reverse engineering of graphical user interfaces for testing.” in WCRE , vol. 3, 2003, p. 260
2003
Earlier work this paper cites.
J. J. Garrett et al. , “Ajax: A new approach to web applications,” 2005
2005
Earlier work this paper cites.
2006
Earlier work this paper cites.
K. Li and M. Wu, Effective GUI testing automation: Developing an automated GUI testing tool . John Wiley & Sons, 2006
2006
Earlier work this paper cites.
X. Xiao and Y. Tao, “Personalized privacy preservation,” in Proceedings of the 2006 ACM SIGMOD international conference on Management of data , 2006, pp. 229–240
2006
Earlier work this paper cites.
S. Mitra and T. Acharya, “Gesture recognition: A survey,” IEEE Transactions on Systems, Man, and Cybernetics, Part C (Applications and Reviews) , vol. 37, no. 3, pp. 311–324, 2007
2007
Earlier work this paper cites.
J. He, I.-L. Yen, T. Peng, J. Dong, and F. Bastani, “An adaptive user interface generation framework for web services,” in 2008 IEEE Congress on Services Part II (services-2 2008) . IEEE, 2008, pp. 175–182
2008
Earlier work this paper cites.
R. Hardy and E. Rukzio, “Touch & interact: touch-based interaction of mobile phones with displays,” in Proceedings of the 10th international conference on Human computer interaction with mobile devices and services , 2008, pp. 245–254
2008
Earlier work this paper cites.
K. Jokinen, “User interaction in mobile navigation applications,” in Map-based Mobile Services: Design, Interaction and Usability . Springer, 2008, pp. 168–197
2008
Earlier work this paper cites.
T. Yeh, T.-H. Chang, and R. C. Miller, “Sikuli: using gui screenshots for search and automation,” in Proceedings of the 22nd annual ACM symposium on User interface software and technology , 2009, pp. 183–192
2009
Earlier work this paper cites.
A. Bruns, A. Kornstadt, and D. Wichmann, “Web application tests with selenium,” IEEE software , vol. 26, no. 5, pp. 88–91, 2009
2009
Earlier work this paper cites.
T.-H. Chang, T. Yeh, and R. C. Miller, “Gui testing using computer vision,” in Proceedings of the SIGCHI Conference on Human Factors in Computing Systems , 2010, pp. 1535–1544
2010
Earlier work this paper cites.
2010
Earlier work this paper cites.
T. D. Hellmann and F. Maurer, “Rule-based exploratory testing of graphical user interfaces,” in 2011 Agile Conference . IEEE, 2011, pp. 107–116
2011
Earlier work this paper cites.
W. Enck, D. Octeau, P. D. McDaniel, and S. Chaudhuri, “A study of android application security.” in USENIX security symposium , vol. 2, no. 2, 2011
2011
Earlier work this paper cites.
M. Egele, C. Kruegel, E. Kirda, and G. Vigna, “Pios: Detecting privacy leaks in ios applications.” in NDSS , vol. 2011, 2011, p. 18th
2011
Earlier work this paper cites.
N. Fernandes, R. Lopes, and L. Carriço, “On web accessibility evaluation environments,” in Proceedings of the International Cross-Disciplinary Conference on Web Accessibility , 2011, pp. 1–10
2011
Earlier work this paper cites.
A. P. Felt, E. Chin, S. Hanna, D. Song, and D. Wagner, “Android permissions demystified,” in Proceedings of the 18th ACM conference on Computer and communications security , 2011, pp. 627–638
2011
Earlier work this paper cites.
R. Gove and J. Faytong, “Machine learning and event-based software testing: classifiers for identifying infeasible gui event sequences,” in Advances in computers . Elsevier, 2012, vol. 86, pp. 109–135
2012
Earlier work this paper cites.
C. B. Browne, E. Powley, D. Whitehouse, S. M. Lucas, P. I. Cowling, P. Rohlfshagen, S. Tavener, D. Perez, S. Samothrakis, and S. Colton, “A survey of monte carlo tree search methods,” IEEE Transactions on Computational Intelligence and AI in games , vol. 4, no. 1, pp. 1–43, 2012
2012
Earlier work this paper cites.
H. Hao, V. Singh, and W. Du, “On the effectiveness of api-level access control using bytecode rewriting in android,” in Proceedings of the 8th ACM SIGSAC symposium on Information, computer and communications security , 2013, pp. 25–36
2013
Earlier work this paper cites.
2014
Earlier work this paper cites.
T. Wetzlmaier, R. Ramler, and W. Putschögl, “A framework for monkey gui testing,” in 2016 IEEE international conference on software testing, verification and validation (ICST) . IEEE, 2016, pp. 416–423
2016
Earlier work this paper cites.
X. Zeng, D. Li, W. Zheng, F. Xia, Y. Deng, W. Lam, W. Yang, and T. Xie, “Automated test input generation for android: are we really there yet in an industrial case?” in Proceedings of the 2016 24th ACM SIGSOFT International Symposium on Foundations of Software Engineering , ser. FSE 2016. New York, NY, USA: Association for Computing Machinery, 2016, p. 987–992. [Online]. Available: https://doi.org/10.1145/2950290.2983958
2016
Earlier work this paper cites.
X. Gu, H. Zhang, D. Zhang, and S. Kim, “Deep api learning,” in Proceedings of the 2016 24th ACM SIGSOFT international symposium on foundations of software engineering , 2016, pp. 631–642
2016
Earlier work this paper cites.
K. Weiss, T. M. Khoshgoftaar, and D. Wang, “A survey of transfer learning,” Journal of Big data , vol. 3, pp. 1–40, 2016
2016
Earlier work this paper cites.
A. Vaswani, “Attention is all you need,” Advances in Neural Information Processing Systems , 2017
2017
Earlier work this paper cites.
T. Shi, A. Karpathy, L. Fan, J. Hernandez, and P. Liang, “World of bits: An open-domain platform for web-based agents,” in Proceedings of the 34th International Conference on Machine Learning , ser. Proceedings of Machine Learning Research, D. Precup and Y. W. Teh, Eds., vol. 70. PMLR, 06–11 Aug 2017, pp. 3135–3144. [Online]. Available: https://proceedings.mlr.press/v70/shi17a.html
2017
Earlier work this paper cites.
B. Deka, Z. Huang, C. Franzen, J. Hibschman, D. Afergan, Y. Li, J. Nichols, and R. Kumar, “Rico: A mobile app dataset for building data-driven design applications,” in Proceedings of the 30th Annual ACM Symposium on User Interface Software and Technology , ser. UIST ’17. New York, NY, USA: Association for Computing Machinery, 2017, p. 845–854. [Online]. Available: https://doi.org/10.1145/3126594.3126651
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
J. M. Bradshaw, P. J. Feltovich, and M. Johnson, “Human–agent interaction,” in The handbook of human-machine interaction . CRC Press, 2017, pp. 283–300
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
A. Radford, “Improving language understanding by generative pre-training,” 2018
2018
Earlier work this paper cites.
O. Gambino, L. Rundo, V. Cannella, S. Vitabile, and R. Pirrone, “A framework for data-driven adaptive gui generation based on dicom,” Journal of biomedical informatics , vol. 88, pp. 37–52, 2018
2018
Earlier work this paper cites.
G. Hu, L. Zhu, and J. Yang, “Appflow: using machine learning to synthesize robust, reusable ui tests,” in Proceedings of the 2018 26th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering , ser. ESEC/FSE 2018. New York, NY, USA: Association for Computing Machinery, 2018, p. 269–282. [Online]. Available: https://doi.org/10.1145/3236024.3236055
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
K. Moran, C. Watson, J. Hoskins, G. Purnell, and D. Poshyvanyk, “Detecting and summarizing gui changes in evolving mobile apps,” in Proceedings of the 33rd ACM/IEEE international conference on automated software engineering , 2018, pp. 543–553
2018
Earlier work this paper cites.
M. Lutaaya, “Rethinking app permissions on ios,” in Extended Abstracts of the 2018 CHI Conference on Human Factors in Computing Systems , 2018, pp. 1–6
2018
Earlier work this paper cites.
L. Ivančić, D. Suša Vugec, and V. Bosilj Vukšić, “Robotic process automation: systematic literature review,” in Business Process Management: Blockchain and Central and Eastern Europe Forum: BPM 2019 Blockchain and CEE Forum, Vienna, Austria, September 1–6, 2019, Proceedings 17 . Springer, 2019, pp. 280–295
2019
Earlier work this paper cites.
S. Agostinelli, A. Marrella, and M. Mecella, “Research challenges for intelligent robotic process automation,” in Business Process Management Workshops: BPM 2019 International Workshops, Vienna, Austria, September 1–6, 2019, Revised Selected Papers 17 . Springer, 2019, pp. 12–18
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever et al. , “Language models are unsupervised multitask learners,” OpenAI blog , vol. 1, no. 8, p. 9, 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
T. D. White, G. Fraser, and G. J. Brown, “Improving random gui testing with image-based widget detection,” in Proceedings of the 28th ACM SIGSOFT international symposium on software testing and analysis , 2019, pp. 307–317
2019
Earlier work this paper cites.
Y. Li, Z. Yang, Y. Guo, and X. Chen, “Humanoid: A deep learning-based approach to automated black-box android app testing,” in 2019 34th IEEE/ACM International Conference on Automated Software Engineering (ASE) . IEEE, 2019, pp. 1070–1073
2019
Earlier work this paper cites.
C. Zhang, P. Patras, and H. Haddadi, “Deep learning in mobile and wireless networking: A survey,” IEEE Communications surveys & tutorials , vol. 21, no. 3, pp. 2224–2287, 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
K. S. Said, L. Nie, A. A. Ajibode, and X. Zhou, “Gui testing for mobile applications: objectives, approaches and challenges,” in Proceedings of the 12th Asia-Pacific Symposium on Internetware , 2020, pp. 51–60
2020
Earlier work this paper cites.
M. Bajammal, A. Stocco, D. Mazinanian, and A. Mesbah, “A survey on the use of computer vision to improve software engineering tasks,” IEEE Transactions on Software Engineering , vol. 48, no. 5, pp. 1722–1742, 2020
2020
Earlier work this paper cites.
R. Syed, S. Suriadi, M. Adams, W. Bandara, S. J. Leemans, C. Ouyang, A. H. Ter Hofstede, I. Van De Weerd, M. T. Wynn, and H. A. Reijers, “Robotic process automation: contemporary themes and challenges,” Computers in Industry , vol. 115, p. 103162, 2020
2020
Earlier work this paper cites.
T. Chakraborti, V. Isahagian, R. Khalaf, Y. Khazaeni, V. Muthusamy, Y. Rizk, and M. Unuvar, “From robotic process automation to intelligent process automation: –emerging trends–,” in Business Process Management: Blockchain and Robotic Process Automation Forum: BPM 2020 Blockchain and RPA Forum, Seville, Spain, September 13–18, 2020, Proceedings 18 . Springer, 2020, pp. 215–228
2020
Earlier work this paper cites.
J. G. Enríquez, A. Jiménez-Ramírez, F. J. Domínguez-Mayo, and J. A. García-García, “Robotic process automation: a scientific and industrial systematic mapping study,” IEEE Access , vol. 8, pp. 39 113–39 129, 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu, “Exploring the limits of transfer learning with a unified text-to-text transformer,” Journal of machine learning research , vol. 21, no. 140, pp. 1–67, 2020
2020
Earlier work this paper cites.
J. Qian, Z. Shang, S. Yan, Y. Wang, and L. Chen, “Roscript: A visual script driven truly non-intrusive robotic testing system for touch screen applications,” in 2020 IEEE/ACM 42nd International Conference on Software Engineering (ICSE) , 2020, pp. 297–308
2020
Earlier work this paper cites.
J. Chen, M. Xie, Z. Xing, C. Chen, X. Xu, L. Zhu, and G. Li, “Object detection for graphical user interface: Old fashioned or deep learning or a combination?” in proceedings of the 28th ACM joint meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering , 2020, pp. 1202–1214
2020
Earlier work this paper cites.
S. Agostinelli, M. Lupia, A. Marrella, and M. Mecella, “Automated generation of executable rpa scripts from user interface logs,” in Business Process Management: Blockchain and Robotic Process Automation Forum: BPM 2020 Blockchain and RPA Forum, Seville, Spain, September 13–18, 2020, Proceedings 18 . Springer, 2020, pp. 116–131
2020
Earlier work this paper cites.
M. Xie, S. Feng, Z. Xing, J. Chen, and C. Chen, “Uied: a hybrid tool for gui element detection,” in Proceedings of the 28th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering , 2020, pp. 1655–1659
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
Y. Li, J. He, X. Zhou, Y. Zhang, and J. Baldridge, “Mapping natural language instructions to mobile ui action sequences,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , 2020, pp. 8198–8210
2020
Earlier work this paper cites.
P. Martins, F. Sá, F. Morgado, and C. Cunha, “Using machine learning for cognitive robotic process automation (rpa),” in 2020 15th Iberian Conference on Information Systems and Technologies (CISTI) . IEEE, 2020, pp. 1–6
2020
Earlier work this paper cites.
P. Lewis, E. Perez, A. Piktus, F. Petroni, V. Karpukhin, N. Goyal, H. Küttler, M. Lewis, W.-t. Yih, T. Rocktäschel et al. , “Retrieval-augmented generation for knowledge-intensive nlp tasks,” Advances in Neural Information Processing Systems , vol. 33, pp. 9459–9474, 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
J. Andreas, J. Bufe, D. Burkett, C. Chen, J. Clausman, J. Crawford, K. Crim, J. DeLoach, L. Dorner, J. Eisner et al. , “Task-oriented dialogue as dataflow synthesis,” Transactions of the Association for Computational Linguistics , vol. 8, pp. 556–571, 2020
2020
Earlier work this paper cites.
H. Sampath, A. Merrick, and A. P. Macvean, “Accessibility of command line interfaces,” Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems , 2021. [Online]. Available: https://api.semanticscholar.org/CorpusID:233987139
2021
Earlier work this paper cites.
O. Rodríguez-Valdés, T. E. Vos, P. Aho, and B. Marín, “30 years of automated gui testing: a bibliometric analysis,” in Quality of Information and Communications Technology: 14th International Conference, QUATIC 2021, Algarve, Portugal, September 8–11, 2021, Proceedings 14 . Springer, 2021, pp. 473–488
2021
Earlier work this paper cites.
J. Ribeiro, R. Lima, T. Eckhardt, and S. Paiva, “Robotic process automation and artificial intelligence in industry 4.0–a literature review,” Procedia Computer Science , vol. 181, pp. 51–58, 2021
2021
Earlier work this paper cites.
M. Nass, E. Alégroth, and R. Feldt, “Why many challenges with gui test automation (will) remain,” Information and Software Technology , vol. 138, p. 106625, 2021
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
Y. Li and O. Hilliges, Artificial intelligence for human computer interaction: a modern approach . Springer, 2021
2021
Earlier work this paper cites.
M. F. Granda, O. Parra, and B. Alba-Sarango, “Towards a model-driven testing framework for gui test cases generation from user stories.” in ENASE , 2021, pp. 453–460
2021
Earlier work this paper cites.
T. J.-J. Li, L. Popowski, T. Mitchell, and B. A. Myers, “Screen2vec: Semantic embedding of gui screens and gui components,” in Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems , 2021, pp. 1–15
2021
Earlier work this paper cites.
J. Ye, K. Chen, X. Xie, L. Ma, R. Huang, Y. Chen, Y. Xue, and J. Zhao, “An empirical study of gui widget detection for industrial mobile games,” in Proceedings of the 29th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering , 2021, pp. 1427–1437
2021
Earlier work this paper cites.
F. YazdaniBanafsheDaragh and S. Malek, “Deep gui: Black-box gui input generation with deep learning,” in 2021 36th IEEE/ACM International Conference on Automated Software Engineering (ASE) . IEEE, 2021, pp. 905–916
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
A. Ramesh, M. Pavlov, G. Goh, S. Gray, C. Voss, A. Radford, M. Chen, and I. Sutskever, “Zero-shot text-to-image generation,” in International conference on machine learning . Pmlr, 2021, pp. 8821–8831
2021
Earlier work this paper cites.
X. Zhan, T. Liu, L. Fan, L. Li, S. Chen, X. Luo, and Y. Liu, “Research on third-party libraries in android apps: A taxonomy and systematic literature review,” IEEE Transactions on Software Engineering , vol. 48, no. 10, pp. 4181–4213, 2021
2021
Earlier work this paper cites.
B. Wang, G. Li, X. Zhou, Z. Chen, T. Grossman, and Y. Li, “Screen2words: Automatic mobile ui summarization with multimodal learning,” in The 34th Annual ACM Symposium on User Interface Software and Technology , 2021, pp. 498–510
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al. , “Learning transferable visual models from natural language supervision,” in International conference on machine learning . PMLR, 2021, pp. 8748–8763
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” in Proceedings of the IEEE/CVF international conference on computer vision , 2021, pp. 10 012–10 022
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray et al. , “Training language models to follow instructions with human feedback,” Advances in neural information processing systems , vol. 35, pp. 27 730–27 744, 2022
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou et al. , “Chain-of-thought prompting elicits reasoning in large language models,” Advances in neural information processing systems , vol. 35, pp. 24 824–24 837, 2022
2022
Earlier work this paper cites.
H. Y. Abuaddous, A. M. Saleh, O. Enaizan, F. Ghabban, and A. B. Al-Badareen, “Automated user experience (ux) testing for mobile application: Strengths and limitations.” International Journal of Interactive Mobile Technologies , vol. 16, no. 4, 2022
2022
Earlier work this paper cites.
N. Rupp, K. Peschke, M. Köppl, D. Drissner, and T. Zuchner, “Establishment of low-cost laboratory automation processes using autoit and 4-axis robots,” SLAS technology , vol. 27, no. 5, pp. 312–318, 2022
2022
Earlier work this paper cites.
J. Qian, Y. Ma, C. Lin, and L. Chen, “Accelerating ocr-based widget localization for test automation of gui applications,” in Proceedings of the 37th IEEE/ACM International Conference on Automated Software Engineering , 2022, pp. 1–13
2022
Earlier work this paper cites.
Z. Stefanidi, G. Margetis, S. Ntoa, and G. Papagiannakis, “Real-time adaptation of context-aware intelligent user interfaces, for enhanced situational awareness,” IEEE Access , vol. 10, pp. 23 367–23 393, 2022
2022
Earlier work this paper cites.
S. Yao, H. Chen, J. Yang, and K. Narasimhan, “Webshop: Towards scalable real-world web interaction with grounded language agents,” Advances in Neural Information Processing Systems , vol. 35, pp. 20 744–20 757, 2022
2022
Earlier work this paper cites.
H. Lee, J. Park, and U. Lee, “A systematic survey on android api usage for data-driven analytics with smartphones,” ACM Computing Surveys , vol. 55, no. 5, pp. 1–38, 2022
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, “High-resolution image synthesis with latent diffusion models,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2022, pp. 10 684–10 695
2022
Earlier work this paper cites.
Z. Liu, H. Mao, C.-Y. Wu, C. Feichtenhofer, T. Darrell, and S. Xie, “A convnet for the 2020s,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2022, pp. 11 976–11 986
2022
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
T. Wu, S. He, J. Liu, S. Sun, K. Liu, Q.-L. Han, and Y. Tang, “A brief overview of chatgpt: The history, status quo and potential future development,” IEEE/CAA Journal of Automatica Sinica , vol. 10, no. 5, pp. 1122–1136, 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
X. Li, “Gui testing for android applications: a survey,” in 2023 7th International Conference on Computer, Software and Modeling (ICCSM) . IEEE, 2023, pp. 6–10
2023
Earlier work this paper cites.
J.-J. Oksanen, “Test automation for windows gui application,” 2023
2023
Earlier work this paper cites.
P. S. Deshmukh, S. S. Date, P. N. Mahalle, and J. Barot, “Automated gui testing for enhancing user experience (ux): A survey of the state of the art,” in International Conference on ICT for Sustainable Development . Springer, 2023, pp. 619–628
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
A. Wali, S. Mahamad, and S. Sulaiman, “Task automation intelligent agents: A review,” Future Internet , vol. 15, no. 6, p. 196, 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
T. S. d. Moura, E. L. Alves, H. F. d. Figueirêdo, and C. d. S. Baptista, “Cytestion: Automated gui testing for web applications,” in Proceedings of the XXXVII Brazilian Symposium on Software Engineering , 2023, pp. 388–397
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
Z. Zou, K. Chen, Z. Shi, Y. Guo, and J. Ye, “Object detection in 20 years: A survey,” Proceedings of the IEEE , vol. 111, no. 3, pp. 257–276, 2023
2023
Earlier work this paper cites.
P. Brie, N. Burny, A. Sluÿters, and J. Vanderdonckt, “Evaluating a large language model on searching for gui layouts,” Proceedings of the ACM on Human-Computer Interaction , vol. 7, no. EICS, pp. 1–37, 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
J. Wang, Z. Liu, L. Zhao, Z. Wu, C. Ma, S. Yu, H. Dai, Q. Yang, Y. Liu, S. Zhang et al. , “Review of large vision models and visual prompt engineering,” Meta-Radiology , p. 100047, 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
X. Wu, J. Ye, K. Chen, X. Xie, Y. Hu, R. Huang, L. Ma, and J. Zhao, “Widget detection-based testing for industrial mobile games,” in 2023 IEEE/ACM 45th International Conference on Software Engineering: Software Engineering in Practice (ICSE-SEIP) . IEEE, 2023, pp. 173–184
2023
Earlier work this paper cites.
A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y. Lo et al. , “Segment anything,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 4015–4026
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
F. P. Ricós, R. Neeft, B. Marín, T. E. Vos, and P. Aho, “Using gui change detection for delta testing,” in International Conference on Research Challenges in Information Science . Springer, 2023, pp. 509–517
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
J. Li, D. Li, S. Savarese, and S. Hoi, “Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,” in International conference on machine learning . PMLR, 2023, pp. 19 730–19 742
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
X. Deng, Y. Gu, B. Zheng, S. Chen, S. Stevens, B. Wang, H. Sun, and Y. Su, “Mind2web: Towards a generalist agent for the web,” Advances in Neural Information Processing Systems , vol. 36, pp. 28 091–28 114, 2023
2023
Earlier work this paper cites.
X. Zhai, B. Mustafa, A. Kolesnikov, and L. Beyer, “Sigmoid loss for language image pre-training,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 11 975–11 986
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
J. Wu, S. Wang, S. Shen, Y.-H. Peng, J. Nichols, and J. P. Bigham, “Webui: A dataset for enhancing visual ui understanding with web semantics,” in Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems , 2023, pp. 1–14
2023
Earlier work this paper cites.
G. Li, H. A. A. K. Hammoud, H. Itani, D. Khizbullin, and B. Ghanem, “Camel: Communicative agents for "mind" exploration of large language model society,” in Thirty-seventh Conference on Neural Information Processing Systems , 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
B. Wang, G. Li, and Y. Li, “Enabling conversational interaction with mobile ui using large language models,” in Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems , 2023, pp. 1–17
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
C. Rawles, A. Li, D. Rodriguez, O. Riva, and T. Lillicrap, “Androidinthewild: A large-scale dataset for android device control,” Advances in Neural Information Processing Systems , vol. 36, pp. 59 708–59 728, 2023
2023
Earlier work this paper cites.
OpenAI, “Gpt-4v(ision) system card,” OpenAI, Tech. Rep., September 2023. [Online]. Available: https://cdn.openai.com/papers/GPTV_System_Card.pdf
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
D. Zimmermann and A. Koziolek, “Gui-based software testing: An automated approach using gpt-4 and selenium webdriver,” in 2023 38th IEEE/ACM International Conference on Automated Software Engineering Workshops (ASEW) . IEEE, 2023, pp. 171–174
2023
Earlier work this paper cites.
Z. Liu, C. Chen, J. Wang, X. Che, Y. Huang, J. Hu, and Q. Wang, “Fill in the blank: Context-aware automated text input generation for mobile gui testing,” in 2023 IEEE/ACM 45th International Conference on Software Engineering (ICSE) . IEEE, 2023, pp. 1355–1367
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
W.-L. Chiang, Z. Li, Z. Lin, Y. Sheng, Z. Wu, H. Zhang, L. Zheng, S. Zhuang, Y. Zhuang, J. E. Gonzalez, I. Stoica, and E. P. Xing, “Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality,” March 2023. [Online]. Available: https://lmsys.org/blog/2023-03-30-vicuna/
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
D. Zimmermann and A. Koziolek, “Automating gui-based software testing with gpt-3,” in 2023 IEEE International Conference on Software Testing, Verification and Validation Workshops (ICSTW) , 2023, pp. 62–65
2023
Earlier work this paper cites.
K. Q. Lin, P. Zhang, J. Chen, S. Pramanick, D. Gao, A. J. Wang, R. Yan, and M. Z. Shou, “Univtg: Towards unified video-language temporal grounding,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 2794–2804
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
X. Huang, W. Ruan, W. Huang, G. Jin, Y. Dong, C. Wu, S. Bensalem, R. Mu, Y. Qi, X. Zhao, K. Cai, Y. Zhang, S. Wu, P. Xu, D. Wu, A. Freitas, and M. A. Mustafa, “A survey of safety and trustworthiness of large language models through the lens of verification and validation,” Artif. Intell. Rev. , vol. 57, p. 175, 2023. [Online]. Available: https://api.semanticscholar.org/CorpusID:258823083
2023
Earlier work this paper cites.
S. Jha, S. K. Jha, P. Lincoln, N. D. Bastian, A. Velasquez, and S. Neema, “Dehallucinating large language models using formal methods guided iterative prompting,” 2023 IEEE International Conference on Assured Autonomy (ICAA) , pp. 149–152, 2023. [Online]. Available: https://api.semanticscholar.org/CorpusID:260810131
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
A. Biswas and W. Talukdar, “Guardrails for trust, safety, and ethical development and deployment of large language models (llm),” Journal of Science & Technology , vol. 4, no. 6, pp. 55–82, 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
A. Piñeiro-Martín, C. García-Mateo, L. Docío-Fernández, and M. D. C. Lopez-Perez, “Ethical challenges in the development of virtual assistants powered by large language models,” Electronics , vol. 12, no. 14, p. 3170, 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
W. contributors, “Large language model — wikipedia, the free encyclopedia,” 2024, accessed: 2024-11-25. [Online]. Available: https://en.wikipedia.org/wiki/Large_language_model
2024
Earlier work this paper cites.
2024
Cited alongside, same era.
Z. Shen, “Llm with tools: A survey,” arXiv preprint arXiv:2409.18807 , 2024
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
S. Feng and C. Chen, “Prompting is all you need: Automated android bug replay with large language models,” in Proceedings of the 46th IEEE/ACM International Conference on Software Engineering , 2024, pp. 1–13
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
D. Ran, H. Wang, Z. Song, M. Wu, Y. Cao, Y. Zhang, W. Yang, and T. Xie, “Guardian: A runtime framework for llm-based ui exploration,” in Proceedings of the 33rd ACM SIGSOFT International Symposium on Software Testing and Analysis , 2024, pp. 958–970
2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2024
Cited alongside, same era.
K. Mei, Z. Li, S. Xu, R. Ye, Y. Ge, and Y. Zhang, “Aios: Llm agent operating system,” arXiv e-prints, pp. arXiv–2403 , 2024
2024
Cited alongside, same era.
W. Aljedaani, A. Habib, A. Aljohani, M. M. Eler, and Y. Feng, “Does chatgpt generate accessible code? investigating accessibility challenges in llm-generated source code,” in International Cross-Disciplinary Conference on Web Accessibility , 2024. [Online]. Available: https://api.semanticscholar.org/CorpusID:273550267
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
M. Zhuge, C. Zhao, D. R. Ashley, W. Wang, D. Khizbullin, Y. Xiong, Z. Liu, E. Chang, R. Krishnamoorthi, Y. Tian, Y. Shi, V. Chandra, and J. Schmidhuber, “Agent-as-a-judge: Evaluate agents with agents,” 2024. [Online]. Available: https://api.semanticscholar.org/CorpusID:273350802
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
Y. Li, Y. Li, and Y. Yang, “Test-agent: A multimodal app automation testing framework based on the large language model,” in 2024 IEEE 4th International Conference on Digital Twins and Parallel Intelligence (DTPI) . IEEE, 2024, pp. 609–614
2024
Closest in time.
J. Gorniak, Y. Kim, D. Wei, and N. W. Kim, “Vizability: Enhancing chart accessibility with llm-based conversational interaction,” in Proceedings of the 37th Annual ACM Symposium on User Interface Software and Technology , 2024, pp. 1–19
2024
Closest in time.
Y. Guan, D. Wang, Z. Chu, S. Wang, F. Ni, R. Song, and C. Zhuang, “Intelligent agents with llm-based process automation,” in Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining , 2024, pp. 5018–5027
2024
Closest in time.
2024
Closest in time.
D. Gao, S. Hu, Z. Bai, Q. Lin, and M. Z. Shou, “Assisteditor: Multi-agent collaboration for gui workflow automation in video creation,” in Proceedings of the 32nd ACM International Conference on Multimedia , 2024, pp. 11 255–11 257
2024
Closest in time.
2024
Closest in time.
W. Gao, K. Du, Y. Luo, W. Shi, C. Yu, and Y. Shi, “Easyask: An in-app contextual tutorial search assistant for older adults with voice and touch inputs,” Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies , vol. 8, no. 3, pp. 1–27, 2024
2024
Closest in time.
OpenAdapt AI, “OpenAdapt: Open Source Generative Process Automation,” 2024, accessed: 2024-10-26. [Online]. Available: https://github.com/OpenAdaptAI/OpenAdapt
2024
Closest in time.
AgentSeaf AI. (2024) Introduction to agentsea platform. Accessed: 2024-10-26. [Online]. Available: https://www.agentsea.ai/
2024
Closest in time.
O. Interpreter, “Open interpreter: A natural language interface for computers,” GitHub repository, 2024, accessed: 2024-10-27. [Online]. Available: https://github.com/OpenInterpreter/open-interpreter
2024
Closest in time.
MultiOn AI. (2024) Multion ai: Ai agents that act on your behalf. Accessed: 2024-10-26. [Online]. Available: https://www.multion.ai/
2024
Closest in time.
HONOR, “Honor introduces magicos 9.0,” 2024, accessed: 2024-11-16. [Online]. Available: https://www.fonearena.com/blog/438680/honor-magicos-9-0-features.html
2024
Closest in time.
H. W. Chung, L. Hou, S. Longpre, B. Zoph, Y. Tay, W. Fedus, Y. Li, X. Wang, M. Dehghani, S. Brahma et al. , “Scaling instruction-finetuned language models,” Journal of Machine Learning Research , vol. 25, no. 70, pp. 1–53, 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
R. Rafailov, A. Sharma, E. Mitchell, C. D. Manning, S. Ermon, and C. Finn, “Direct preference optimization: Your language model is secretly a reward model,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
CogAgent Team, “Cogagent: Cognitive ai agent platform,” https://cogagent.aminer.cn/home , 2024, accessed: 2024-12-17
2024
Closest in time.
S. Yao, D. Yu, J. Zhao, I. Shafran, T. Griffiths, Y. Cao, and K. Narasimhan, “Tree of thoughts: Deliberate problem solving with large language models,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
Baidu Research, “ERNIE Bot: Baidu’s Knowledge-Enhanced Large Language Model Built on Full AI Stack Technology,” 2024, [Online; accessed 9-November-2024]. [Online]. Available: https://research.baidu.com/Blog/index-view?id=183
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
D. Chen, Y. Huang, Z. Ma, H. Chen, X. Pan, C. Ge, D. Gao, Y. Xie, Z. Liu, J. Gao et al. , “Data-juicer: A one-stop data processing system for large language models,” in Companion of the 2024 International Conference on Management of Data , 2024, pp. 120–134
2024
Closest in time.
B. Ding, C. Qin, R. Zhao, T. Luo, X. Li, G. Chen, W. Xia, J. Hu, L. A. Tuan, and S. Joty, “Data augmentation using llms: Data perspectives, learning paradigms and challenges,” in Findings of the Association for Computational Linguistics ACL 2024 , 2024, pp. 1679–1705
2024
Closest in time.
2024
Closest in time.
Z. Guo, S. Cheng, H. Wang, S. Liang, Y. Qin, P. Li, Z. Liu, M. Sun, and Y. Liu, “Stabletoolbench: Towards stable large-scale benchmarking on tool learning of large language models,” 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
MosaicML, “Mosaicml: Mpt-7b,” 2023, accessed: 2024-11-19. [Online]. Available: https://www.mosaicml.com/blog/mpt-7b
2024
Closest in time.
2024
Closest in time.
J. Wang, Y. Huang, C. Chen, Z. Liu, S. Wang, and Q. Wang, “Software testing with large language models: Survey, landscape, and vision,” IEEE Transactions on Software Engineering , 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
X. Zhang, H. Xu, Z. Ba, Z. Wang, Y. Hong, J. Liu, Z. Qin, and K. Ren, “Privacyasst: Safeguarding user privacy in tool-using large language model agents,” IEEE Transactions on Dependable and Secure Computing , 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
J. Lin, J. Tang, H. Tang, S. Yang, W.-M. Chen, W.-C. Wang, G. Xiao, X. Dang, C. Gan, and S. Han, “Awq: Activation-aware weight quantization for on-device llm compression and acceleration,” Proceedings of Machine Learning and Systems , vol. 6, pp. 87–100, 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
W. Kuang, B. Qian, Z. Li, D. Chen, D. Gao, X. Pan, Y. Xie, Y. Li, B. Ding, and J. Zhou, “Federatedscope-llm: A comprehensive package for fine-tuning large language models in federated learning,” in Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining , 2024, pp. 5260–5271
2024
Closest in time.
L. de Castro, A. Polychroniadou, and D. Escudero, “Privacy-preserving large language model inference via gpu-accelerated fully homomorphic encryption,” in Neurips Safe Generative AI Workshop 2024
2024
Closest in time.
J. Wolff, W. Lehr, and C. S. Yoo, “Lessons from gdpr for ai policymaking,” Virginia Journal of Law & Technology , vol. 27, no. 4, p. 2, 2024
2024
Closest in time.
Z. Zhang, M. Jia, H.-P. Lee, B. Yao, S. Das, A. Lerner, D. Wang, and T. Li, ““it’s a fair game”, or is it? examining how users navigate disclosure risks and benefits when using llm-based conversational agents,” in Proceedings of the CHI Conference on Human Factors in Computing Systems , 2024, pp. 1–26
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
W. Lee, J. Lee, J. Seo, and J. Sim, “ { \{ InfiniGen } \} : Efficient generative inference of large language models with dynamic { \{ KV } \} cache management,” in 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24) , 2024, pp. 155–172
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
L. Zhang, Q. Jin, H. Huang, D. Zhang, and F. Wei, “Respond in my language: Mitigating language inconsistency in response generation based on large language models,” in Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , 2024, pp. 4177–4192
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
J. Gao, S. A. Gebreegziabher, K. T. W. Choo, T. J.-J. Li, S. T. Perrault, and T. W. Malone, “A taxonomy for human-llm interaction modes: An initial exploration,” in Extended Abstracts of the CHI Conference on Human Factors in Computing Systems , 2024, pp. 1–11
2024
Closest in time.
2024
Closest in time.
C. Y. Kim, C. P. Lee, and B. Mutlu, “Understanding large-language model (llm)-powered human-robot interaction,” in Proceedings of the 2024 ACM/IEEE International Conference on Human-Robot Interaction , 2024, pp. 371–380
2024
Closest in time.
2024
Closest in time.
J. Wester, T. Schrills, H. Pohl, and N. van Berkel, ““as an ai language model, i cannot”: Investigating llm denials of user requests,” in Proceedings of the CHI Conference on Human Factors in Computing Systems , 2024, pp. 1–14
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
S. Kim, H. Kang, S. Choi, D. Kim, M. Yang, and C. Park, “Large language models meet collaborative filtering: An efficient all-round llm-based recommender system,” in Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining , 2024, pp. 1395–1406
2024
Closest in time.
2024
Closest in time.
I. H. Sarker, “Llm potentiality and awareness: a position paper from the perspective of trustworthy and responsible ai modeling,” Discover Artificial Intelligence , vol. 4, no. 1, p. 40, 2024
2024
Closest in time.
Y. Yu, Y. Zhuang, J. Zhang, Y. Meng, A. J. Ratner, R. Krishna, J. Shen, and C. Zhang, “Large language model as attributed training data generator: A tale of diversity and bias,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
OpenAI, “Computer-using agent: Introducing a universal interface for ai to interact with the digital world,” 2025. [Online]. Available: https://openai.com/index/computer-using-agent
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
K. Li, Z. Meng, H. Lin, Z. Luo, Y. Tian, J. Ma, Z. Huang, and T.-S. Chua, “Screenspot-pro: Gui grounding for professional high-resolution computer use,” 2025
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
P. P. S. Dammu, “Towards ethical and personalized web navigation agents: A framework for user-aligned task execution,” in Proceedings of the Eighteenth ACM International Conference on Web Search and Data Mining , 2025, pp. 1074–1076
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
Y. Yin, Y. Mei, C. Yu, T. J.-J. Li, A. K. Jadoon, S. Cheng, W. Shi, M. Chen, and Y. Shi, “From operation to cognition: Automatic modeling cognitive dependencies from user demonstrations for gui task automation,” in Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems , 2025, pp. 1–24
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
K. You, H. Zhang, E. Schoop, F. Weers, A. Swearngin, J. Nichols, Y. Yang, and Z. Gan, “Ferret-ui: Grounded mobile ui understanding with multimodal llms,” in European Conference on Computer Vision . Springer, 2025, pp. 240–255
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
J. Yang, Y. Dong, S. Liu, B. Li, Z. Wang, H. Tan, C. Jiang, J. Kang, Y. Zhang, K. Zhou et al. , “Octopus: Embodied vision-language programmer from environmental feedback,” in European Conference on Computer Vision . Springer, 2025, pp. 20–38
2025
Closest in time.
Y. Jin, S. Petrangeli, Y. Shen, and G. Wu, “Screenllm: Stateful screen schema for efficient action understanding and prediction,” 2025
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
S. Kara, F. Faisal, and S. Nath, “Waber: Web agent benchmarking for efficiency and reliability,” in ICLR 2025 Workshop on Foundation Models in the Wild
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
B. Wang, X. Wang, J. Deng, T. Xie, R. Li, Y. Zhang, G. Li, T. J. Hua, I. Stoica, W.-L. Chiang, D. Yang, Y. Su, Y. Zhang, Z. Wang, V. Zhong, and T. Yu, “Computer agent arena: Compare & test computer use agents on crowdsourced real-world tasks,” 2025
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
T. Rosenbach, D. Heidrich, and A. Weinert, “Automated testing of the gui of a real-life engineering software using large language models,” in 2025 IEEE International Conference on Software Testing, Verification and Validation Workshops (ICSTW) . IEEE, 2025, pp. 103–110
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
OpenAI, “Operator system card,” Jan. 2025, released on January 23, 2025
2025
Closest in time.
2025
Closest in time.
F. AI, “Eko - build production-ready agentic workflow with natural language,” https://eko.fellou.ai/ , 2025, accessed: 2025-01-15
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
L. Aichberger, A. Paren, Y. Gal, P. Torr, and A. Bibi, “Attacking multimodal os agents with malicious image patches,” in ICLR 2025 Workshop on Foundation Models in the Wild
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
N. Hojo, K. Shinoda, Y. Yamazaki, K. Suzuki, H. Sugiyama, K. Nishida, and K. Saito, “Generativegui: Dynamic gui generation leveraging llms for enhanced user interaction on chat interfaces,” in Proceedings of the Extended Abstracts of the CHI Conference on Human Factors in Computing Systems , 2025, pp. 1–9
2025
Closest in time.
Z. Zhang, E. Schoop, J. Nichols, A. Mahajan, and A. Swearngin, “From interaction to impact: Towards safer ai agent through understanding and evaluating mobile ui operation impacts,” in Proceedings of the 30th International Conference on Intelligent User Interfaces , 2025, pp. 727–744
2025
Closest in time.