Fetching the paper…
Reading the bibliography…
Recent advancements in multi-modal large language models (MLLMs) have led to substantial improvements in visual understanding, primarily driven by sophisticated modality alignment strategies.
Distinctive image features from scale-invariant keypoints
D. G. Lowe · 2004
Earlier work this paper cites.
The pascal visual object classes (voc) challenge
M. Everingham, L. Van Gool, C. K. Williams, J. Winn, and A. Zisserman · 2010
Earlier work this paper cites.
Saliency detection via graph-based manifold ranking
C. Yang, L. Zhang, H. Lu, X. Ruan, and M.-H. Yang · 2013
Earlier work this paper cites.
Referitgame: Referring to objects in photographs of natural scenes
S. Kazemzadeh, V. Ordonez, M. Matten, and T. Berg · 2014
Earlier work this paper cites.
Hierarchical image saliency detection on extended cssd
J. Shi, Q. Yan, L. Xu, and J. Jia · 2015
Earlier work this paper cites.
The fast bilateral solver
J. T. Barron and B. Poole · 2016
Earlier work this paper cites.
Tgif: A new dataset and benchmark on animated gif description
Y. Li, Y. Song, L. Cao, J. Tetreault, L. Goldberg, A. Jaimes, and J. Luo · 2016
Earlier work this paper cites.
Generation and comprehension of unambiguous object descriptions
J. Mao, J. Huang, A. Toshev, O. Camburu, A. L. Yuille, and K. Murphy · 2016
Earlier work this paper cites.
Modeling context in referring expressions
L. Yu, P. Poirson, S. Yang, A. C. Berg, and T. L. Berg · 2016
Earlier work this paper cites.
Scene parsing through ade20k dataset
B. Zhou, H. Zhao, X. Puig, S. Fidler, A. Barriuso, and A. Torralba · 2017
Earlier work this paper cites.
Deeply supervised salient object detection with short connections
Q. Hou, M.-M. Cheng, X. Hu, A. Borji, Z. Tu, and P. Torr · 2018
Earlier work this paper cites.
Dvqa: Understanding data visualizations via question answering
K. Kafle, B. Price, S. Cohen, and C. Kanan · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
A. Radford, K. Narasimhan, T. Salimans, I. Sutskever, et al · 2018
Earlier work this paper cites.
Gqa: A new dataset for real-world visual reasoning and compositional question answering
D. A. Hudson and C. D. Manning · 2019
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, A. Desmaison, A. Kopf, E. Yang, Z. DeVito, M. Raison, A. Tejani, S. Chilamkurthy, B. Steiner, L. Fang, J. Bai, and S. Chintala · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever, et al · 2019
Earlier work this paper cites.
Towards vqa models that can read
A. Singh, V. Natarajan, M. Shah, Y. Jiang, X. Chen, D. Batra, D. Parikh, and M. Rohrbach · 2019
Earlier work this paper cites.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models, 2021
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer · 2021
Cited alongside, same era.
Localizing objects with self-supervised transformers and no labels
O. Siméoni, G. Puy, H. V. Vo, S. Roburin, S. Gidaris, A. Bursuc, P. Pérez, R. Marlet, and J. Ponce · 2021
Cited alongside, same era.
Document collection visual question answering
R. Tito, D. Karatzas, and E. Valveny · 2021
Cited alongside, same era.
Video-chatgpt: Towards detailed video understanding via large vision and language models
M. Maaz, H. Rasheed, S. Khan, and F. S. Khan · 2023
Closest in time.
Kosmos-2: Grounding multimodal large language models to the world
Z. Peng, W. Wang, L. Dong, Y. Hao, S. Huang, S. Ma, and F. Wei · 2023
Closest in time.
Paco: Parts and attributes of common objects
V. Ramanathan, A. Kalia, V. Petrovic, Y. Wen, B. Zheng, B. Guo, R. Wang, A. Marquez, R. Kovvuri, A. Kadian, et al · 2023
Closest in time.
Gemini: a family of highly capable multimodal models
G. Team, R. Anil, S. Borgeaud, Y. Wu, J.-B. Alayrac, J. Yu, R. Soricut, J. Schalkwyk, A. M. Dai, A. Hauth, et al · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Flamingo: a visual language model for few-shot learning
J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds, et al · 2022
Cited alongside, same era.
Move: Unsupervised movable object segmentation and detection
A. Bielski and P. Favaro · 2022
Cited alongside, same era.
Learn to explain: Multimodal reasoning via thought chains for science question answering
P. Lu, S. Mishra, T. Xia, L. Qiu, K.-W. Chang, S.-C. Zhu, O. Tafjord, P. Clark, and A. Kalyan · 2022
Cited alongside, same era.
Unsupervised salient object detection with spectral cluster voting
G. Shin, S. Albanie, and W. Xie · 2022
Cited alongside, same era.
Lavt: Language-aware vision transformer for referring image segmentation
Z. Yang, J. Wang, Y. Tang, K. Chen, H. Zhao, and P. H. Torr · 2022
Cited alongside, same era.
Seqtr: A simple yet universal network for visual grounding
C. Zhu, Y. Zhou, Y. Shen, G. Luo, X. Pan, M. Lin, C. Chen, L. Cao, X. Sun, and R. Ji · 2022
Cited alongside, same era.
Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond
J. Bai, S. Bai, S. Yang, S. Wang, S. Tan, P. Wang, J. Lin, C. Zhou, and J. Zhou · 2023
Cited alongside, same era.
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, et al · 2023
Closest in time.
Cogvlm: Visual expert for large language models
W. Wang, Q. Lv, W. Yu, W. Hong, Y. Wang, J. Ji, Z. Yang, L. Zhao, X. Song, J. Xu, X. Bin, H. Li, Y. Dong, M. Ding, and J. Tang · 2023
Closest in time.
Next-gpt: Any-to-any multimodal llm
S. Wu, H. Fei, L. Qu, W. Ji, and T.-S. Chua · 2023
Closest in time.
Universal instance perception as object discovery and retrieval
B. Yan, Y. Jiang, J. Wu, D. Wang, P. Luo, Z. Yuan, and H. Lu · 2023
Closest in time.
mplug-owl: Modularization empowers large language models with multimodality
Q. Ye, H. Xu, G. Xu, J. Ye, M. Yan, Y. Zhou, J. Wang, A. Hu, P. Shi, Y. Shi, et al · 2023
Closest in time.
A survey on multimodal large language models
S. Yin, C. Fu, S. Zhao, K. Li, X. Sun, T. Xu, and E. Chen · 2023
Closest in time.
Osprey: Pixel understanding with visual instruction tuning
Y. Yuan, W. Li, J. Liu, D. Tang, X. Luo, C. Qin, L. Zhang, and J. Zhu · 2023
Closest in time.
Bubogpt: Enabling visual grounding in multi-modal llms
Y. Zhao, Z. Lin, D. Zhou, Z. Huang, J. Feng, and B. Kang · 2023
Closest in time.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
D. Zhu, J. Chen, X. Shen, X. Li, and M. Elhoseiny · 2023
Closest in time.
Allava: Harnessing gpt4v-synthesized data for a lite vision-language model
G. H. Chen, S. Chen, R. Zhang, J. Chen, X. Wu, Z. Zhang, Z. Chen, J. Li, X. Wan, and B. Wang · 2024
Closest in time.
Obelics: An open web-scale filtered dataset of interleaved image-text documents
H. Laurençon, L. Saulnier, L. Tronchon, S. Bekman, A. Singh, A. Lozhkov, T. Wang, S. Karamcheti, A. Rush, D. Kiela, et al · 2024
Closest in time.
Mini-gemini: Mining the potential of multi-modality vision language models
Y. Li, Y. Zhang, C. Wang, Z. Zhong, Y. Chen, R. Chu, S. Liu, and J. Jia · 2024
Closest in time.
Visionllm: Large language model is also an open-ended decoder for vision-centric tasks
W. Wang, Z. Chen, X. Chen, J. Wu, X. Zhu, G. Zeng, P. Luo, T. Lu, J. Zhou, Y. Qiao, et al · 2024
Closest in time.