Fetching the paper…
Reading the bibliography…
The rapid development of Multimodal Large Language Models (MLLMs) has enabled the integration of multiple modalities, including texts and images, within the large language model (LLM) framework.
Geometric deep learning: Going beyond euclidean data
Michael M. Bronstein, Joan Bruna, Yann LeCun, Arthur Szlam, and Pierre Vandergheynst · 2017
Earlier work this paper cites.
Attention is all you need
A Vaswani · 2017
Earlier work this paper cites.
Visual arts search on mobile devices
Hui Mao, James She, and Ming Cheung · 2019
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al · 2020
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Earlier work this paper cites.
Flamingo: a visual language model for few-shot learning
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al · 2022
Earlier work this paper cites.
Predict then propagate: Graph neural networks meet personalized pagerank, 2022
Johannes Gasteiger, Aleksandar Bojchevski, and Stephan Günnemann · 2022
Earlier work this paper cites.
Edgeformer: A parameter-efficient transformer for on-device seq2seq generation
Tao Ge, Si-Qing Chen, and Furu Wei · 2022
Earlier work this paper cites.
Draw your art dream: Diverse digital art synthesis with multimodal guided diffusion, 2022
Nisha Huang, Fan Tang, Weiming Dong, and Changsheng Xu · 2022
Earlier work this paper cites.
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer · 2022
Cited alongside, same era.
Art and the science of generative ai: A deeper dive
Ziv Epstein, John Kowalski, Laura Thomas, and Steve Zhang · 2023
Cited alongside, same era.
Can llms effectively leverage graph structural information: when and why
Jin Huang, Xingjian Zhang, Qiaozhu Mei, and Jiaqi Ma · 2023
Cited alongside, same era.
Heterformer: Transformer-based deep node representation learning on heterogeneous text-rich networks
Bowen Jin, Yu Zhang, Qi Zhu, and Jiawei Han · 2023
Cited alongside, same era.
Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi · 2023
Cited alongside, same era.
Gnn-rag: Graph neural retrieval for large language model reasoning
Costas Mavromatis and George Karypis · 2024
Later among the works it cites.
Learning on multimodal graphs: A survey, 2024
Ciyuan Peng, Jiayuan He, and Feng Xia · 2024
Later among the works it cites.
A survey of large language models for graphs
Xubin Ren, Jiabin Tang, Dawei Yin, Nitesh Chawla, and Chao Huang · 2024
Later among the works it cites.
Generative multimodal models are in-context learners
Quan Sun, Yufeng Cui, Xiaosong Zhang, Fan Zhang, Qiying Yu, Yueze Wang, Yongming Rao, Jingjing Liu, Tiejun Huang, and Xinlong Wang · 2024
Later among the works it cites.
Graphgpt: Graph instruction tuning for large language models
Jiabin Tang, Yuhao Yang, Wei Wei, Lei Shi, Lixin Su, Suqi Cheng, Dawei Yin, and Chao Huang · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A survey on multimodal large language models
Shukang Yin, Chaoyou Fu, Sirui Zhao, Ke Li, Xing Sun, Tong Xu, and Enhong Chen · 2023
Cited alongside, same era.
Graph-toolformer: To empower llms with graph reasoning ability via prompt augmented by chatgpt
Jiawei Zhang · 2023
Cited alongside, same era.
A review of modern recommender systems using generative models (gen-recsys), 2024
Yashar Deldjoo, Zhankui He, Julian McAuley, Anton Korikov, Scott Sanner, Arnau Ramisa, René Vidal, Maheswaran Sathiamoorthy, Atoosa Kasirzadeh, and Silvia Milano · 2024
Cited alongside, same era.
Dreamllm: Synergistic multimodal comprehension and creation, 2024
Runpei Dong, Chunrui Han, Yuang Peng, Zekun Qi, Zheng Ge, Jinrong Yang, Liang Zhao, Jianjian Sun, Hongyu Zhou, Haoran Wei, Xiangwen Kong, Xiangyu Zhang, Kaisheng Ma, and Li Yi · 2024
Cited alongside, same era.
Graphwiz: An instruction-following language model for graph problems
Nuo Chen, Yuhan Li, Jianheng Tang, and Jia Li
Cited in the paper.
Llaga: Large language and graph assistant, 2024b
Runjin Chen, Tong Zhao, Ajay Jaiswal, Neil Shah, and Zhangyang Wang
Cited in the paper.
Uniglm: Training one unified language model for text-attributed graph embedding, 2024a
Yi Fang, Dongzhe Fan, Sirui Ding, Ninghao Liu, and Qiaoyu Tan
Cited in the paper.
Chameleon Team · 2024
Later among the works it cites.
Jianing Wang, Junda Wu, Yupeng Hou, Yao Liu, Ming Gao, and Julian McAuley · 2024
Later among the works it cites.
Language is all a graph needs, 2024
Ruosong Ye, Caiqi Zhang, Runhui Wang, Shuyuan Xu, and Yongfeng Zhang · 2024
Later among the works it cites.
Mm-llms: Recent advances in multimodal large language models
Duzhen Zhang, Yahan Yu, Jiahua Dong, Chenxing Li, Dan Su, Chenhui Chu, and Dong Yu · 2024
Later among the works it cites.