Y. Du, C. Li, R. Guo, X. Yin, W. Liu, J. Zhou, Y. Bai, Z. Yu, Y. Yang, Q. Dang, and H. Wang, “Pp-ocr: A practical ultra lightweight ocr system,” 2020. [Online]. Available: https://arxiv.org/abs/2009.09941
Original
2009
Earlier work this paper cites.
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An image is worth 16x16 words: Transformers for image recognition at scale,” 2021. [Online]. Available: https://arxiv.org/abs/2010.11929
Original
2010
Earlier work this paper cites.
S. Ren, K. He, R. Girshick, and J. Sun, “Faster r-cnn: Towards real-time object detection with region proposal networks,” 2016. [Online]. Available: https://arxiv.org/abs/1506.01497
Original
2016
Earlier work this paper cites.
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov, “Proximal policy optimization algorithms,” 2017. [Online]. Available: https://arxiv.org/abs/1707.06347
Original
2017
Earlier work this paper cites.
T. Shi, A. Karpathy, L. Fan, J. Hernandez, and P. Liang, “World of bits: An open-domain platform for web-based agents,” in International Conference on Machine Learning . PMLR, 2017, pp. 3135–3144
2017
Earlier work this paper cites.
B. Deka, Z. Huang, C. Franzen, J. Hibschman, D. Afergan, Y. Li, J. Nichols, and R. Kumar, “Rico: A mobile app dataset for building data-driven design applications,” in Proceedings of the 30th annual ACM symposium on user interface software and technology , 2017, pp. 845–854
2017
Earlier work this paper cites.
E. Z. Liu, K. Guu, P. Pasupat, T. Shi, and P. Liang, “Reinforcement learning on web interfaces using workflow-guided exploration,” arXiv preprint arXiv:1802.08802 , 2018
Original
2018
Earlier work this paper cites.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “Bert: Pre-training of deep bidirectional transformers for language understanding,” 2019. [Online]. Available: https://arxiv.org/abs/1810.04805
Original
2019
Earlier work this paper cites.
Y. Li, J. He, X. Zhou, Y. Zhang, and J. Baldridge, “Mapping natural language instructions to mobile ui action sequences,” arXiv preprint arXiv:2005.03776 , 2020
Original
2020
Earlier work this paper cites.
C. Bai, X. Zang, Y. Xu, S. Sunkara, A. Rastogi, J. Chen, and B. A. y. Arcas, “UIBert: Learning Generic Multimodal Representations for UI Understanding,” Aug. 2021, arXiv:2107.13731. [Online]. Available: http://arxiv.org/abs/2107.13731
Original
2021
Earlier work this paper cites.
OpenAI, “ChatGPT,” https://openai.com/research/chatgpt/ , 2021
2021
Earlier work this paper cites.
X. Chen, Z. Zhao, L. Chen, D. Zhang, J. Ji, A. Luo, Y. Xiong, and K. Yu, “Websrc: A dataset for web-based structural reading comprehension,” arXiv preprint arXiv:2101.09465 , 2021
Original
2021
Earlier work this paper cites.
A. Zeng, M. Attarian, B. Ichter, K. Choromanski, A. Wong, S. Welker, F. Tombari, A. Purohit, M. Ryoo, V. Sindhwani, J. Lee, V. Vanhoucke, and P. Florence, “Socratic models: Composing zero-shot multimodal reasoning with language,” May 2022. [Online]. Available: http://arxiv.org/abs/2204.00598
Original
2022
Earlier work this paper cites.
L. Sun, X. Chen, L. Chen, T. Dai, Z. Zhu, and K. Yu, “META-GUI: Towards Multi-modal Conversational Agents on Mobile GUI,” in Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing , Y. Goldberg, Z. Kozareva, and Y. Zhang, Eds. Abu Dhabi, United Arab Emirates: Association for Computational Linguistics, Dec. 2022, pp. 6699–6712. [Online]. Available: https://aclanthology.org/2022.emnlp-main.449
2022
Earlier work this paper cites.
Y. Huang, T. Lv, L. Cui, Y. Lu, and F. Wei, “Layoutlmv3: Pre-training for document ai with unified text and image masking,” 2022. [Online]. Available: https://arxiv.org/abs/2204.08387
Original
2022
Earlier work this paper cites.
S. Min, M. Lewis, L. Zettlemoyer, and H. Hajishirzi, “Metaicl: Learning to learn in context,” May 2022. [Online]. Available: http://arxiv.org/abs/2110.15943
Original
2022
Earlier work this paper cites.
E. Zelikman, Y. Wu, J. Mu, and N. D. Goodman, “Star: Bootstrapping reasoning with reasoning,” May 2022. [Online]. Available: http://arxiv.org/abs/2203.14465
Original
2022
Earlier work this paper cites.
A. Burns, D. Arsan, S. Agrawal, R. Kumar, K. Saenko, and B. A. Plummer, “A dataset for interactive vision-language navigation with unknown command feasibility,” in European Conference on Computer Vision . Springer, 2022, pp. 312–328
2022
Earlier work this paper cites.
L. Sun, X. Chen, L. Chen, T. Dai, Z. Zhu, and K. Yu, “Meta-gui: Towards multi-modal conversational agents on mobile gui,” arXiv preprint arXiv:2205.11029 , 2022
Original
2022
Earlier work this paper cites.
S. Yao, H. Chen, J. Yang, and K. Narasimhan, “Webshop: Towards scalable real-world web interaction with grounded language agents,” Advances in Neural Information Processing Systems , vol. 35, pp. 20 744–20 757, 2022
2022
Earlier work this paper cites.
C. Zhang, Z. Yang, J. Liu, Y. Han, X. Chen, Z. Huang, B. Fu, and G. Yu, “AppAgent: multimodal agents as smartphone users,” Dec. 2023, arXiv:2312.13771 TLDR: A novel LLM-based multimodal agent framework designed to operate smartphone applications through a simplified action space, mimicking human-like interactions such as tapping and swiping, thereby broadening its applicability across diverse apps. [Online]. Available: http://arxiv.org/abs/2312.13771
Original
2023
Earlier work this paper cites.
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, A. Rodriguez, A. Joulin, E. Grave, and G. Lample, “Llama: Open and efficient foundation language models,” 2023. [Online]. Available: https://arxiv.org/abs/2302.13971
Original
2023
Earlier work this paper cites.
J. Bai, S. Bai, Y. Chu, Z. Cui, K. Dang, X. Deng, Y. Fan, W. Ge, Y. Han, F. Huang, B. Hui, L. Ji, M. Li, J. Lin, R. Lin, D. Liu, G. Liu, C. Lu, K. Lu, J. Ma, R. Men, X. Ren, X. Ren, C. Tan, S. Tan, J. Tu, P. Wang, S. Wang, W. Wang, S. Wu, B. Xu, J. Xu, A. Yang, H. Yang, J. Yang, S. Yang, Y. Yao, B. Yu, H. Yuan, Z. Yuan, J. Zhang, X. Zhang, Y. Zhang, Z. Zhang, C. Zhou, J. Zhou, X. Zhou, and T. Zhu, “Qwen technical report,” 2023. [Online]. Available: https://arxiv.org/abs/2309.16609
Original
2023
Earlier work this paper cites.
H. Liu, C. Li, Q. Wu, and Y. J. Lee, “Visual instruction tuning,” 2023. [Online]. Available: https://arxiv.org/abs/2304.08485
Original
2023
Earlier work this paper cites.
Y. Shen, K. Song, X. Tan, D. Li, W. Lu, and Y. Zhuang, “Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face,” 2023. [Online]. Available: https://arxiv.org/abs/2303.17580
Original
2023
Earlier work this paper cites.
J. Wei, X. Wang, D. Schuurmans, M. Bosma, B. Ichter, F. Xia, E. Chi, Q. Le, and D. Zhou, “Chain-of-thought prompting elicits reasoning in large language models,” Jan. 2023. [Online]. Available: http://arxiv.org/abs/2201.11903
Original
2023
Earlier work this paper cites.
S. Yao, J. Zhao, D. Yu, N. Du, I. Shafran, K. Narasimhan, and Y. Cao, “React: Synergizing reasoning and acting in language models,” Mar. 2023. [Online]. Available: http://arxiv.org/abs/2210.03629
Original
2023
Earlier work this paper cites.
Z. Yang, L. Li, J. Wang, K. Lin, E. Azarnasab, F. Ahmed, Z. Liu, C. Liu, M. Zeng, and L. Wang, “Mm-react: Prompting chatgpt for multimodal reasoning and action,” Mar. 2023. [Online]. Available: http://arxiv.org/abs/2303.11381
Original
2023
Earlier work this paper cites.
S. Zhou, F. F. Xu, H. Zhu, X. Zhou, R. Lo, A. Sridhar, X. Cheng, T. Ou, Y. Bisk, D. Fried et al. , “Webarena: A realistic web environment for building autonomous agents,” arXiv preprint arXiv:2307.13854 , 2023
Original
2023
Earlier work this paper cites.
D. Gao, L. Ji, Z. Bai, M. Ouyang, P. Li, D. Mao, Q. Wu, W. Zhang, P. Wang, X. Guo et al. , “Assistgui: Task-oriented desktop graphical user interface automation,” arXiv preprint arXiv:2312.13108 , 2023
Original
2023
Earlier work this paper cites.
J. Yang, H. Zhang, F. Li, X. Zou, C. Li, and J. Gao, “Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V,” Nov. 2023, arXiv:2310.11441 TLDR: The experiments show that GPT-4V with SoM in zero-shot setting outperforms the state-of-the-art fully-finetuned referring expression comprehension and segmentation model on RefCOCOg, and the effectiveness of SoM on a wide range of fine-grained vision and multimodal tasks is validated. [Online]. Available: http://arxiv.org/abs/2310.11441
Original
2023
Earlier work this paper cites.
W. Hong, W. Wang, Q. Lv, J. Xu, W. Yu, J. Ji, Y. Wang, Z. Wang, Y. Zhang, J. Li, B. Xu, Y. Dong, M. Ding, and J. Tang, “CogAgent: A Visual Language Model for GUI Agents,” Dec. 2023, arXiv:2312.08914 [cs]. [Online]. Available: http://arxiv.org/abs/2312.08914
Original
2023
Earlier work this paper cites.
W. Zhou, Y. E. Jiang, L. Li, J. Wu, T. Wang, S. Qiu, J. Zhang, J. Chen, R. Wu, S. Wang, S. Zhu, J. Chen, W. Zhang, X. Tang, N. Zhang, H. Chen, P. Cui, and M. Sachan, “Agents: An open-source framework for autonomous language agents,” 2023. [Online]. Available: https://arxiv.org/abs/2309.07870
Original
2023
Earlier work this paper cites.
N. Shinn, F. Cassano, E. Berman, A. Gopinath, K. Narasimhan, and S. Yao, “Reflexion: Language agents with verbal reinforcement learning,” Oct. 2023. [Online]. Available: http://arxiv.org/abs/2303.11366
Original
2023
Earlier work this paper cites.
G. Wang, Y. Xie, Y. Jiang, A. Mandlekar, C. Xiao, Y. Zhu, L. Fan, and A. Anandkumar, “Voyager: An open-ended embodied agent with large language models,” Oct. 2023. [Online]. Available: http://arxiv.org/abs/2305.16291
Original
2023
Earlier work this paper cites.
X. Wang, J. Wei, D. Schuurmans, Q. Le, E. Chi, S. Narang, A. Chowdhery, and D. Zhou, “Self-consistency improves chain of thought reasoning in language models,” Mar. 2023. [Online]. Available: http://arxiv.org/abs/2203.11171
Original
2023
Earlier work this paper cites.
S. Yao, D. Yu, J. Zhao, I. Shafran, T. L. Griffiths, Y. Cao, and K. Narasimhan, “Tree of thoughts: Deliberate problem solving with large language models,” 2023. [Online]. Available: https://arxiv.org/abs/2305.10601
Original
2023
Earlier work this paper cites.
K. Ma, H. Zhang, H. Wang, X. Pan, W. Yu, and D. Yu, “LASER: LLM Agent with State-Space Exploration for Web Navigation,” arXiv e-prints , p. arXiv:2309.08172, Sep. 2023
Original
2023
Earlier work this paper cites.
H. Li, J. Su, Y. Chen, Q. Li, and Z. Zhang, “Sheetcopilot: Bringing software productivity to the next level through large language models,” Oct. 2023. [Online]. Available: http://arxiv.org/abs/2305.19308
Original
2023
Earlier work this paper cites.
C. Rawles, A. Li, D. Rodriguez, O. Riva, and T. Lillicrap, “Android in the Wild: A Large-Scale Dataset for Android Device Control,” Oct. 2023, arXiv:2307.10088 TLDR: A dataset for device-control research, Android in the Wild (AITW), which is orders of magnitude larger than current datasets, and contains multi-step tasks that require semantic understanding of language and visual context. [Online]. Available: http://arxiv.org/abs/2307.10088
Original
2023
Earlier work this paper cites.
A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann et al. , “Palm: Scaling language modeling with pathways,” Journal of Machine Learning Research , vol. 24, no. 240, pp. 1–113, 2023
2023
Earlier work this paper cites.
A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y. Lo, P. Dollár, and R. Girshick, “Segment anything,” 2023. [Online]. Available: https://arxiv.org/abs/2304.02643
Original
2023
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” 2023. [Online]. Available: https://arxiv.org/abs/1706.03762
Original
2023
Earlier work this paper cites.
A. Zhao, D. Huang, Q. Xu, M. Lin, Y.-J. Liu, and G. Huang, “Expel: Llm agents are experiential learners,” 2023. [Online]. Available: https://arxiv.org/abs/2308.10144
Original
2023
Earlier work this paper cites.
X. Zhu, Y. Chen, H. Tian, C. Tao, W. Su, C. Yang, G. Huang, B. Li, L. Lu, X. Wang, Y. Qiao, Z. Zhang, and J. Dai, “Ghost in the minecraft: Generally capable agents for open-world environments via large language models with text-based knowledge and memory,” 2023. [Online]. Available: https://arxiv.org/abs/2305.17144
Original
2023
Earlier work this paper cites.
W. Zhou, Y. E. Jiang, P. Cui, T. Wang, Z. Xiao, Y. Hou, R. Cotterell, and M. Sachan, “Recurrentgpt: Interactive generation of (arbitrarily) long text,” 2023. [Online]. Available: https://arxiv.org/abs/2305.13304
Original
2023
Earlier work this paper cites.
H. Su, W. Shi, J. Kasai, Y. Wang, Y. Hu, M. Ostendorf, W. tau Yih, N. A. Smith, L. Zettlemoyer, and T. Yu, “One embedder, any task: Instruction-finetuned text embeddings,” 2023. [Online]. Available: https://arxiv.org/abs/2212.09741
Original
2023
Earlier work this paper cites.
S. Hao, Y. Gu, H. Ma, J. J. Hong, Z. Wang, D. Z. Wang, and Z. Hu, “Reasoning with language model is planning with world model,” Oct. 2023. [Online]. Available: http://arxiv.org/abs/2305.14992
Original
2023
Earlier work this paper cites.
H. Lightman, V. Kosaraju, Y. Burda, H. Edwards, B. Baker, T. Lee, J. Leike, J. Schulman, I. Sutskever, and K. Cobbe, “Let’s verify step by step,” May 2023. [Online]. Available: http://arxiv.org/abs/2305.20050
Original
2023
Earlier work this paper cites.
L. Zheng, W.-L. Chiang, Y. Sheng, S. Zhuang, Z. Wu, Y. Zhuang, Z. Lin, Z. Li, D. Li, E. P. Xing, H. Zhang, J. E. Gonzalez, and I. Stoica, “Judging llm-as-a-judge with mt-bench and chatbot arena,” Dec. 2023. [Online]. Available: http://arxiv.org/abs/2306.05685
Original
2023
Earlier work this paper cites.
D. Zhou, N. Schärli, L. Hou, J. Wei, N. Scales, X. Wang, D. Schuurmans, C. Cui, O. Bousquet, Q. Le, and E. Chi, “Least-to-most prompting enables complex reasoning in large language models,” Apr. 2023. [Online]. Available: http://arxiv.org/abs/2205.10625
Original
2023
Earlier work this paper cites.
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, D. Bikel, L. Blecher, C. C. Ferrer, M. Chen, G. Cucurull, D. Esiobu, J. Fernandes, J. Fu, W. Fu, B. Fuller, C. Gao, V. Goswami, N. Goyal, A. Hartshorn, S. Hosseini, R. Hou, H. Inan, M. Kardas, V. Kerkez, M. Khabsa, I. Kloumann, A. Korenev, P. S. Koura, M.-A. Lachaux, T. Lavril, J. Lee, D. Liskovich, Y. Lu, Y. Mao, X. Martinet, T. Mihaylov, P. Mishra, I. Molybog, Y. Nie, A. Poulton, J. Reizenstein, R. Rungta, K. Saladi, A. Schelten, R. Silva, E. M. Smith, R. Subramanian, X. E. Tan, B. Tang, R. Taylor, A. Williams, J. X. Kuan, P. Xu, Z. Yan, I. Zarov, Y. Zhang, A. Fan, M. Kambadur, S. Narang, A. Rodriguez, R. Stojnic, S. Edunov, and T. Scialom, “Llama 2: Open foundation and fine-tuned chat models,” Jul. 2023. [Online]. Available: http://arxiv.org/abs/2307.09288
Original
2023
Earlier work this paper cites.
A. Madaan, N. Tandon, P. Gupta, S. Hallinan, L. Gao, S. Wiegreffe, U. Alon, N. Dziri, S. Prabhumoye, Y. Yang, S. Gupta, B. P. Majumder, K. Hermann, S. Welleck, A. Yazdanbakhsh, and P. Clark, “Self-refine: Iterative refinement with self-feedback,” May 2023. [Online]. Available: http://arxiv.org/abs/2303.17651
Original
2023
Earlier work this paper cites.
Z. Yuan, H. Yuan, C. Li, G. Dong, K. Lu, C. Tan, C. Zhou, and J. Zhou, “Scaling relationship on learning mathematical reasoning with large language models,” Sep. 2023. [Online]. Available: http://arxiv.org/abs/2308.01825
Original
2023
Earlier work this paper cites.
Y. Wu, P. Zhang, W. Xiong, B. Oguz, J. C. Gee, and Y. Nie, “The role of chain-of-thought in complex vision-language reasoning task,” Nov. 2023. [Online]. Available: http://arxiv.org/abs/2311.09193
Original
2023
Earlier work this paper cites.
G. Zheng, B. Yang, J. Tang, H.-Y. Zhou, and S. Yang, “Ddcot: Duty-distinct chain-of-thought prompting for multimodal reasoning in language models,” Oct. 2023. [Online]. Available: http://arxiv.org/abs/2310.16436
Original
2023
Earlier work this paper cites.
P. Shaw, M. Joshi, J. Cohan, J. Berant, P. Pasupat, H. Hu, U. Khandelwal, K. Lee, and K. Toutanova, “From pixels to ui actions: Learning to follow instructions via graphical user interfaces,” 2023. [Online]. Available: https://arxiv.org/abs/2306.00245
Original
2023
Earlier work this paper cites.
A. Yan, Z. Yang, W. Zhu, K. Lin, L. Li, J. Wang, J. Yang, Y. Zhong, J. McAuley, J. Gao, Z. Liu, and L. Wang, “GPT-4V in Wonderland: Large Multimodal Models for Zero-Shot Smartphone GUI Navigation,” arXiv e-prints , p. arXiv:2311.07562, Nov. 2023
Original
2023
Earlier work this paper cites.
Z. Zhang and A. Zhang, “You Only Look at Screens: Multimodal Chain-of-Action Agents,” arXiv e-prints , p. arXiv:2309.11436, Sep. 2023
Original
2023
Earlier work this paper cites.
D. Zhang, H. Xu, Z. Zhao, L. Chen, R. Cao, and K. Yu, “Mobile-env: an evaluation platform and benchmark for llm-gui interaction,” arXiv preprint arXiv:2305.08144 , 2023
Original
2023
Earlier work this paper cites.
D. Nguyen, J. Chen, Y. Wang, G. Wu, N. Park, Z. Hu, H. Lyu, J. Wu, R. Aponte, Y. Xia, X. Li, J. Shi, H. Chen, V. D. Lai, Z. Xie, S. Kim, R. Zhang, T. Yu, M. Tanjim, N. K. Ahmed, P. Mathur, S. Yoon, L. Yao, B. Kveton, T. H. Nguyen, T. Bui, T. Zhou, R. A. Rossi, and F. Dernoncourt, “Gui agents: A survey,” 2024. [Online]. Available: https://arxiv.org/abs/2412.13501
Original
2024
Earlier work this paper cites.
J. Wang, H. Xu, J. Ye, M. Yan, W. Shen, J. Zhang, F. Huang, and J. Sang, “Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception,” Apr. 2024, arXiv:2401.16158. [Online]. Available: http://arxiv.org/abs/2401.16158
Original
2024
Earlier work this paper cites.
Y. Li, C. Zhang, W. Yang, B. Fu, P. Cheng, X. Chen, L. Chen, and Y. Wei, “AppAgent v2: Advanced Agent for Flexible Mobile Interactions,” Aug. 2024. [Online]. Available: https://arxiv.org/abs/2408.11824v3
Original
2024
Earlier work this paper cites.
K. Cheng, Q. Sun, Y. Chu, F. Xu, Y. Li, J. Zhang, and Z. Wu, “SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents,” Feb. 2024, arXiv:2401.10935 TLDR: This work proposes a novel visual GUI agent – SeeClick, which only relies on screenshots for task automation and creates ScreenSpot, the first realistic GUI grounding benchmark that encompasses mobile, desktop, and web environments. [Online]. Available: http://arxiv.org/abs/2401.10935
Original
2024
Earlier work this paper cites.
C. Zhang, L. Li, S. He, X. Zhang, B. Qiao, S. Qin, M. Ma, Y. Kang, Q. Lin, S. Rajmohan, D. Zhang, and Q. Zhang, “UFO: A UI-Focused Agent for Windows OS Interaction,” May 2024, arXiv:2402.07939 TLDR: UFO, an innovative UI-Focused agent to fulfill user requests tailored to applications on Windows OS, harnessing the capabilities of GPT-Vision is introduced, standing as the first UI agent specifically tailored for task completion within the Windows OS environment. [Online]. Available: http://arxiv.org/abs/2402.07939
Original
2024
Earlier work this paper cites.
K. Q. Lin, L. Li, D. Gao, Z. Yang, S. Wu, Z. Bai, W. Lei, L. Wang, and M. Z. Shou, “Showui: One vision-language-action model for gui visual agent,” 2024. [Online]. Available: https://arxiv.org/abs/2411.17465
Original
2024
Earlier work this paper cites.
Y. He, J. Jin, S. Xia, J. Su, R. Fan, H. Zou, X. Hu, and P. Liu, “Pc agent: While you sleep, ai works – a cognitive journey into digital world,” 2024. [Online]. Available: https://arxiv.org/abs/2412.17589
Original
2024
Earlier work this paper cites.
OpenAI, “Gpt-4 technical report,” 2024. [Online]. Available: https://arxiv.org/abs/2303.08774
Original
2024
Earlier work this paper cites.
T. GLM, “Chatglm: A family of large language models from glm-130b to glm-4 all tools,” 2024. [Online]. Available: https://arxiv.org/abs/2406.12793
Original
2024
Earlier work this paper cites.