Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have achieved superior performance in powering text-based AI agents, endowing them with decision-making and reasoning abilities akin to humans.
Intelligent agents: Theory and practice
Michael Wooldridge and Nicholas R Jennings · 1995
Earlier work this paper cites.
Whisper: Tracing the spatiotemporal process of information diffusion in real time
Nan Cao, Yu-Ru Lin, Xiaohua Sun, David Lazer, Shixia Liu, and Huamin Qu · 2012
Earlier work this paper cites.
Autonomous driving: technical, legal and social aspects
Markus Maurer, J Christian Gerdes, Barbara Lenz, and Hermann Winner · 2016
Earlier work this paper cites.
Policy-focused agent-based modeling using rl behavioral models
Osonde A Osoba, Raffaele Vardavas, Justin Grana, Rushil Zutshi, and Amber Jaycocks · 2020
Earlier work this paper cites.
Completely model-free rl-based consensus of continuous-time multi-agent systems
Xiaoling Wang and Housheng Su · 2020
Earlier work this paper cites.
Paddleseg: A high-efficient development toolkit for image segmentation
Yi Liu, Lutao Chu, Guowei Chen, Zewu Wu, Zeyu Chen, Baohua Lai, and Yuying Hao · 2021
Earlier work this paper cites.
Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi · 2022
Earlier work this paper cites.
A systematic investigation of commonsense knowledge in large language models
Xiang Lorraine Li, Adhiguna Kuncoro, Jordan Hoffmann, Cyprien de Masson d’Autume, Phil Blunsom, and Aida Nematzadeh · 2022
Earlier work this paper cites.
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bjorn Ommer · 2022
Earlier work this paper cites.
Joint feature learning and relation modeling for tracking: A one-stream framework
Botao Ye, Hong Chang, Bingpeng Ma, Shiguang Shan, and Xilin Chen · 2022
Earlier work this paper cites.
Improving image generation with better captions
James Betker, Gabriel Goh, Li Jing, Tim Brooks, Jianfeng Wang, Linjie Li, Long Ouyang, Juntang Zhuang, Joyce Lee, Yufei Guo, et al · 2023
Earlier work this paper cites.
Large language models are visual reasoning coordinators
Liangyu Chen, Bo Li, Sheng Shen, Jingkang Yang, Chunyuan Li, Kurt Keutzer, Trevor Darrell, and Ziwei Liu · 2023
Earlier work this paper cites.
Llava-interactive: An all-in-one demo for image chat, segmentation, generation and editing
Wei-Ge Chen, Irina Spiridonova, Jianwei Yang, Jianfeng Gao, and Chunyuan Li · 2023
Earlier work this paper cites.
Instructblip: Towards general-purpose vision-language models with instruction tuning. arxiv 2023
W Dai, J Li, D Li, AMH Tiong, J Zhao, W Wang, B Li, P Fung, and S Hoi · 2023
Earlier work this paper cites.
Drive like a human: Rethinking autonomous driving with large language models
Daocheng Fu, Xin Li, Licheng Wen, Min Dou, Pinlong Cai, Botian Shi, and Yu Qiao · 2023
Earlier work this paper cites.
Assistgui: Task-oriented desktop graphical user interface automation
Difei Gao, Lei Ji, Zechen Bai, Mingyu Ouyang, Peiran Li, Dongxing Mao, Qinchen Wu, Weichen Zhang, Peiyi Wang, Xiangwu Guo, et al · 2023
Earlier work this paper cites.
Assistgpt: A general multi-modal assistant that can plan, execute, inspect, and learn
Difei Gao, Lei Ji, Luowei Zhou, Kevin Qinghong Lin, Joya Chen, Zihan Fan, and Mike Zheng Shou · 2023
Earlier work this paper cites.
Clova: A closed-loop visual assistant with tool usage and update
Zhi Gao, Yuntao Du, Xintong Zhang, Xiaojian Ma, Wenjuan Han, Song-Chun Zhu, and Qing Li · 2023
Earlier work this paper cites.
Visual programming: Compositional visual reasoning without training
Tanmay Gupta and Aniruddha Kembhavi · 2023
Earlier work this paper cites.
Avis: Autonomous visual information seeking with large language model agent
Ziniu Hu, Ahmet Iscen, Chen Sun, Kai-Wei Chang, Yizhou Sun, David A Ross, Cordelia Schmid, and Alireza Fathi · 2023
Earlier work this paper cites.
Audiogpt: Understanding and generating speech, music, sound, and talking head
Rongjie Huang, Mingze Li, Dongchao Yang, Jiatong Shi, Xuankai Chang, Zhenhui Ye, Yuning Wu, Zhiqing Hong, Jiawei Huang, Jinglin Liu, et al · 2023
Earlier work this paper cites.
Vision-by-language for training-free compositional image retrieval
Shyamgopal Karthik, Karsten Roth, Massimiliano Mancini, and Zeynep Akata · 2023
Earlier work this paper cites.
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C Berg, Wan-Yen Lo, et al · 2023
Earlier work this paper cites.
Sunjae Lee, Junyoung Choi, Jungjae Lee, Hojun Choi, Steven Y Ko, Sangeun Oh, and Insik Shin · 2023
Earlier work this paper cites.
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi · 2023
Earlier work this paper cites.
Llava-plus: Learning to use tools for creating multimodal agents
Shilong Liu, Hao Cheng, Haotian Liu, Hao Zhang, Feng Li, Tianhe Ren, Xueyan Zou, Jianwei Yang, Hang Su, Jun Zhu, et al · 2023
Earlier work this paper cites.
Grounding dino: Marrying dino with grounded pre-training for open-set object detection
Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, Hao Zhang, Jie Yang, Chunyuan Li, Jianwei Yang, Hang Su, Jun Zhu, et al · 2023
Earlier work this paper cites.
Towards robust multi-modal reasoning via model selection
Xiangyan Liu, Rongxue Li, Wei Ji, and Tao Lin · 2023
Cited alongside, same era.
Wavjourney: Compositional audio creation with large language models
Xubo Liu, Zhongkai Zhu, Haohe Liu, Yi Yuan, Meng Cui, Qiushi Huang, Jinhua Liang, Yin Cao, Qiuqiang Kong, Mark D Plumbley, et al · 2023
Cited alongside, same era.
Discuss before moving: Visual language navigation via multi-expert discussions
Yuxing Long, Xiaoqi Li, Wenzhe Cai, and Hao Dong · 2023
Cited alongside, same era.
Chameleon: Plug-and-play compositional reasoning with large language models
Pan Lu, Baolin Peng, Hao Cheng, Michel Galley, Kai-Wei Chang, Ying Nian Wu, Song-Chun Zhu, and Jianfeng Gao · 2023
Cited alongside, same era.
Gpt-4v in wonderland: Large multimodal models for zero-shot smartphone gui navigation
An Yan, Zhengyuan Yang, Wanrong Zhu, Kevin Lin, Linjie Li, Jianfeng Wang, Jianwei Yang, Yiwu Zhong, Julian McAuley, Jianfeng Gao, et al · 2023
Later among the works it cites.
Octopus: Embodied vision-language programmer from environmental feedback
Jingkang Yang, Yuhao Dong, Shuai Liu, Bo Li, Ziyue Wang, Chencheng Jiang, Haoran Tan, Jiamu Kang, Yuanhan Zhang, Kaiyang Zhou, et al · 2023
Later among the works it cites.
Supervised knowledge makes large language models better in-context learners
Linyi Yang, Shuibai Zhang, Zhuohao Yu, Guangsheng Bao, Yidong Wang, Jindong Wang, Ruochen Xu, Wei Ye, Xing Xie, Weizhu Chen, et al · 2023
Later among the works it cites.
Gpt4tools: Teaching large language model to use tools via self-instruction
Rui Yang, Lin Song, Yanwei Li, Sijie Zhao, Yixiao Ge, Xiu Li, and Ying Shan · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jiageng Mao, Yuxi Qian, Hang Zhao, and Yue Wang · 2023
Cited alongside, same era.
Gaia: a benchmark for general ai assistants
Grégoire Mialon, Clémentine Fourrier, Craig Swift, Thomas Wolf, Yann LeCun, and Thomas Scialom · 2023
Cited alongside, same era.
Large language models and knowledge graphs: Opportunities and challenges
Jeff Z Pan, Simon Razniewski, Jan-Christoph Kalo, Sneha Singhania, Jiaoyan Chen, Stefan Dietze, Hajira Jabeen, Janna Omeliyanenko, Wen Zhang, Matteo Lissandrini, et al · 2023
Cited alongside, same era.
Mp5: A multi-modal open-ended embodied system in minecraft via active perception
Yiran Qin, Enshen Zhou, Qichang Liu, Zhenfei Yin, Lu Sheng, Ruimao Zhang, Yu Qiao, and Jing Shao · 2023
Cited alongside, same era.
Toolllm: Facilitating large language models to master 16000+ real-world apis
Yujia Qin, Shihao Liang, Yining Ye, Kunlun Zhu, Lan Yan, Yaxi Lu, Yankai Lin, Xin Cong, Xiangru Tang, Bill Qian, et al · 2023
Cited alongside, same era.
Toolformer: Language models can teach themselves to use tools
Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom · 2023
Cited alongside, same era.
Hugginggpt: Solving ai tasks with chatgpt and its friends in huggingface
Yongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li, Weiming Lu, and Yueting Zhuang · 2023
Cited alongside, same era.
Cognitive architectures for language agents
Theodore Sumers, Shunyu Yao, Karthik Narasimhan, and Thomas L Griffiths · 2023
Cited alongside, same era.
Embodied multi-modal agent trained by an llm from a parallel textworld
Yijun Yang, Tianyi Zhou, Kanxue Li, Dapeng Tao, Lusong Li, Li Shen, Xiaodong He, Jing Jiang, and Yuhui Shi · 2023
Later among the works it cites.
Appagent: Multimodal agents as smartphone users
Zhao Yang, Jiaxuan Liu, Yucheng Han, Xin Chen, Zebiao Huang, Bin Fu, and Gang Yu · 2023
Later among the works it cites.
Mm-react: Prompting chatgpt for multimodal reasoning and action
Zhengyuan Yang, Linjie Li, Jianfeng Wang, Kevin Lin, Ehsan Azarnasab, Faisal Ahmed, Zicheng Liu, Ce Liu, Michael Zeng, and Lijuan Wang · 2023
Later among the works it cites.
Musicagent: An ai agent for music understanding and generation with large language models
Dingyao Yu, Kaitao Song, Peiling Lu, Tianyu He, Xu Tan, Wei Ye, Shikun Zhang, and Jiang Bian · 2023
Later among the works it cites.
Craft: Customizing llms by creating and retrieving from specialized toolsets
Lifan Yuan, Yangyi Chen, Xingyao Wang, Yi R. Fung, Hao Peng, and Heng Ji · 2023
Later among the works it cites.
You only look at screens: Multimodal chain-of-action agents
Zhuosheng Zhan and Aston Zhang · 2023
Later among the works it cites.
Bootstrap your own skills: Learning to solve new tasks with large language model guidance
Jesse Zhang, Jiahui Zhang, Karl Pertsch, Ziyi Liu, Xiang Ren, Minsuk Chang, Shao-Hua Sun, and Joseph J Lim · 2023
Later among the works it cites.
Loop copilot: Conducting ai ensembles for music generation and iterative editing
Yixiao Zhang, Akira Maezawa, Gus Xia, Kazuhiko Yamamoto, and Simon Dixon · 2023
Later among the works it cites.
How do large language models capture the ever-changing world knowledge? a review of recent advances
Zihan Zhang, Meng Fang, Ling Chen, Mohammad-Reza Namazi-Rad, and Jun Wang · 2023
Later among the works it cites.
See and think: Embodied agent in virtual environment
Zhonghan Zhao, Wenhao Chai, Xuan Wang, Li Boyi, Shengyu Hao, Shidong Cao, Tian Ye, Jenq-Neng Hwang, and Gaoang Wang · 2023
Later among the works it cites.
Ddcot: Duty-distinct chain-of-thought prompting for multimodal reasoning in language models
Ge Zheng, Bin Yang, Jiajin Tang, Hong-Yu Zhou, and Sibei Yang · 2023
Later among the works it cites.
Vision language models in autonomous driving and intelligent transportation systems
Xingcheng Zhou, Mingyu Liu, Bare Luka Zagar, Ekim Yurtsever, and Alois C Knoll · 2023
Later among the works it cites.
Visualwebarena: Evaluating multimodal agents on realistic visual web tasks
Yujing Koh, Robert Lo, Lawrence Jang, Vikram Duvvur, Mingchong Lim, Po-yu Huang, Graham Neubig, Shuyan Zhou, Ruslan Salakhutdinov, and Daniel Fried · 2024
Closest in time.
Mulan: Multimodal-llm agent for progressive multi-object diffusion
Sen Li, Ruochen Wang, Cho-Jui Hsieh, et al · 2024
Closest in time.
Styletts 2: Towards human-level text-to-speech through style diffusion and adversarial training with large speech language models
Yinghao Aaron Li, Cong Han, Vinay Raghavan, Gavin Mischler, and Nima Mesgarani · 2024
Closest in time.
Multimodal embodied interactive agent for cafe scene
Yang Liu, Xinshuai Song, Kaixuan Jiang, Weixing Chen, Jingzhou Luo, Guanbin Li, and Liang Lin · 2024
Closest in time.
Weblinx: Real-world website navigation with multi-turn dialogue
Xing Han Lù, Zdeněk Kasner, and Siva Reddy · 2024
Closest in time.
Mllm-tool: A multimodal large language model for tool agent learning
Chenyu Wang, Weixin Luo, Qianyu Chen, Haonan Mai, jindi Guo, Sixun Dong, Xiaohua Xuan, Zhengxin Li, Lin Ma, and Shenghua Gao · 2024
Closest in time.
Mobile-agent: Autonomous multi-modal mobile device agent with visual perception
Junyang Wang, Haiyang Xu, Jiabo Ye, Ming Yan, Weizhou Shen, Ji Zhang, Fei Huang, and Jitao Sang · 2024
Closest in time.
Os-copilot: Towards generalist computer agents with self-improvement
Zhiyong Wu, Chengcheng Han, Zichen Ding, Zhenmin Weng, Zhoumianze Liu, Shunyu Yao, Tao Yu, and Lingpeng Kong · 2024
Closest in time.
Travelplanner: A benchmark for real-world planning with language agents
Jian Xie, Kai Zhang, Jiangjie Chen, Tinghui Zhu, Renze Lou, Yuandong Tian, Yanghua Xiao, and Yu Su · 2024
Closest in time.
Doraemongpt: Toward understanding dynamic scenes with large language models
Zongxin Yang, Guikun Chen, Xiaodi Li, Wenguan Wang, and Yi Yang · 2024
Closest in time.
Gpt-4v-act
ddupont808 · 2026
Closest in time.