Fetching the paper…
Reading the bibliography…
Despite significant advancements in Multimodal Large Language Models (MLLMs) for understanding complex human intentions through cross-modal interactions, capturing intricate image details remains challenging.
S. Kazemzadeh, V. Ordonez, M. Matten, and T. Berg, “Referitgame: Referring to objects in photographs of natural scenes,” in Proc. Empirical Methods Natural Lang. Process. , 2014, pp. 787–798
2014
Earlier work this paper cites.
2016
Earlier work this paper cites.
J. Mao, J. Huang, A. Toshev, O. Camburu, A. L. Yuille, and K. Murphy, “Generation and comprehension of unambiguous object descriptions,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR) , 2016, pp. 11–20
2016
Earlier work this paper cites.
T.-Y. Lin, P. Dollár, R. Girshick, K. He, B. Hariharan, and S. Belongie, “Feature pyramid networks for object detection,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR) , 2017, pp. 2117–2125
2017
Earlier work this paper cites.
Y. Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh, “Making the v in vqa matter: Elevating the role of image understanding in visual question answering,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR) , 2017, pp. 6325–6334
2017
Earlier work this paper cites.
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma et al. , “Visual genome: Connecting language and vision using crowdsourced dense image annotations,” Int. J. Comput. Vis. , pp. 32–73, 2017
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in Proc. Adv. Neural Inf. Process. Syst. , 2017
2017
Earlier work this paper cites.
D. Gurari, Q. Li, A. J. Stangl, A. Guo, C. Lin, K. Grauman, J. Luo, and J. P. Bigham, “Vizwiz grand challenge: Answering visual questions from blind people,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR) , 2018, pp. 3608–3617
2018
Earlier work this paper cites.
S. Wang, H. Lu, and Z. Deng, “Fast object detection in compressed video,” in Proc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV) , 2019, pp. 7104–7113
2019
Earlier work this paper cites.
D. A. Hudson and C. D. Manning, “Gqa: A new dataset for real-world visual reasoning and compositional question answering,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR) , 2019, pp. 6700–6709
2019
Earlier work this paper cites.
A. Singh, V. Natarajan, M. Shah, Y. Jiang, X. Chen, D. Batra, D. Parikh, and M. Rohrbach, “Towards vqa models that can read,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR) , 2019, pp. 8317–8326
2019
Earlier work this paper cites.
K. Marino, M. Rastegari, A. Farhadi, and R. Mottaghi, “Ok-vqa: A visual question answering benchmark requiring external knowledge,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR) , 2019, pp. 3195–3204
2019
Earlier work this paper cites.
A. Mishra, S. Shekhar, A. K. Singh, and A. Chakraborty, “Ocr-vqa: Visual question answering by reading text in images,” in Int. J. Doc. Anal. Recog. , 2019, pp. 947–952
2019
Earlier work this paper cites.
X. Zhu, W. Su, L. Lu, B. Li, X. Wang, and J. Dai, “Deformable detr: Deformable transformers for end-to-end object detection,” in Proc. Int. Conf. Learn. Represent. (ICLR) , 2020
2020
Earlier work this paper cites.
O. Sidorov, R. Hu, M. Rohrbach, and A. Singh, “Textcaps: a dataset for image captioning with reading comprehension,” in Proc. Eur. Conf. Comput. Vis. (ECCV) . Springer, 2020, pp. 742–758
2020
Earlier work this paper cites.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al. , “Learning transferable visual models from natural language supervision,” in Proc. Int. Conf. Mach. Learn. (ICML) , 2021
2021
Earlier work this paper cites.
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly et al. , “An image is worth 16x16 words: Transformers for image recognition at scale,” in Proc. Int. Conf. Learn. Represent. (ICLR) , 2021
2021
Earlier work this paper cites.
G. Ilharco, M. Wortsman, R. Wightman, C. Gordon, N. Carlini, R. Taori, A. Dave, V. Shankar, H. Namkoong, J. Miller, H. Hajishirzi, A. Farhadi, and L. Schmidt, “Openclip,” Jul. 2021. [Online]. Available: https://doi.org/10.5281/zenodo.5143773
2021
Earlier work this paper cites.
M. Raghu, T. Unterthiner, S. Kornblith, C. Zhang, and A. Dosovitskiy, “Do vision transformers see like convolutional neural networks?” in Proc. Adv. Neural Inf. Process. Syst. , vol. 34, 2021, pp. 12 116–12 128
2021
Earlier work this paper cites.
OpenAI, “Chatgpt,” 2022. [Online]. Available: https://chat.openai.com/
2022
Earlier work this paper cites.
Z. Liu, H. Mao, C.-Y. Wu, C. Feichtenhofer, T. Darrell, and S. Xie, “A convnet for the 2020s,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR) , 2022, pp. 11 976–11 986
2022
Earlier work this paper cites.
W. Gao, G. Liao, S. Ma, G. Li, Y. Liang, and W. Lin, “Unified information fusion network for multi-modal rgb-d and rgb-t salient object detection,” IEEE Trans. Circuits Syst. Video Technol. , vol. 32, no. 4, pp. 2091–2106, 2022
2022
Earlier work this paper cites.
J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds et al. , “Flamingo: a visual language model for few-shot learning,” in Proc. Adv. Neural Inf. Process. Syst. , vol. 35, 2022, pp. 23 716–23 736
2022
Earlier work this paper cites.
P. Lu, S. Mishra, T. Xia, L. Qiu, K.-W. Chang, S.-C. Zhu, O. Tafjord, P. Clark, and A. Kalyan, “Learn to explain: Multimodal reasoning via thought chains for science question answering,” in Proc. Adv. Neural Inf. Process. Syst. , 2022
2022
Earlier work this paper cites.
D. Schwenk, A. Khandelwal, C. Clark, K. Marino, and R. Mottaghi, “A-okvqa: A benchmark for visual question answering using world knowledge,” in Proc. Eur. Conf. Comput. Vis. (ECCV) , 2022, pp. 146–162
2022
Earlier work this paper cites.
W. Wang, E. Xie, X. Li, D.-P. Fan, K. Song, D. Liang, T. Lu, P. Luo, and L. Shao, “Pvtv2: Improved baselines with pyramid vision transformer,” Computational Visual Media , vol. 8, no. 3, pp. 415–424, 2022
2022
Earlier work this paper cites.
W. Dai, J. Li, D. Li, A. Huat, J. Zhao, W. Wang, B. Li, P. Fung, and S. Hoi, “Instructblip: Towards general-purpose vision-language models with instruction tuning,” in Proc. Adv. Neural Inf. Process. Syst. , vol. 36, 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
OpenAI, “Gpt-4v(ision) system card,” 2023. [Online]. Available: https://api.semanticscholar.org/CorpusID:263218031
2023
Earlier work this paper cites.
W.-L. Chiang, Z. Li, Z. Lin, Y. Sheng, Z. Wu, H. Zhang, L. Zheng, S. Zhuang, Y. Zhuang, J. E. Gonzalez et al. , “Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality,” Mar. 2023. [Online]. Available: https://lmsys.org/blog/2023-03-30-vicuna/
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
2023
Later among the works it cites.
H. Liu, C. Li, Q. Wu, and Y. J. Lee, “Visual instruction tuning,” in Proc. Adv. Neural Inf. Process. Syst. , vol. 36, 2024
2024
Closest in time.
H. Liu, C. Li, Y. Li, and Y. J. Lee, “Improved baselines with visual instruction tuning,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR) , 2024, pp. 26 296–26 306
2024
Closest in time.
Z. Chen, J. Wu, W. Wang, W. Su, G. Chen, S. Xing, M. Zhong, Q. Zhang, X. Zhu, L. Lu et al. , “Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR) , 2024, pp. 24 185–24 198
2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
I. Team, “Internlm: A multilingual language model with progressively enhanced capabilities,” 2023. [Online]. Available: https://github.com/InternLM/InternLM
2023
Cited alongside, same era.
2023
Cited alongside, same era.
X. Zhai, B. Mustafa, A. Kolesnikov, and L. Beyer, “Sigmoid loss for language image pre-training,” in Proc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV) , 2023, pp. 11 975–11 986
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
S. Xu, H. Zhang, X. Xu, X. Hu, Y. Xu, L. Dai, K.-S. Choi, and P.-A. Heng, “Representative feature alignment for adaptive object detection,” IEEE Trans. Circuits Syst. Video Technol. , vol. 33, no. 2, pp. 689–700, 2023
2023
Cited alongside, same era.
J. Li, D. Li, S. Savarese, and S. Hoi, “Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,” in Proc. Int. Conf. Mach. Learn. (ICML) , 2023, pp. 19 730–19 742
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
M. Oquab, T. Darcet, T. Moutakanni, H. Vo, M. Szafraniec, V. Khalidov, P. Fernandez, D. Haziza, F. Massa, A. El-Nouby et al. , “Dinov2: Learning robust visual features without supervision,” Trans. Mach. Learn. Res. , pp. 1–31, 2024
2024
Closest in time.
Y. Fang, Q. Sun, X. Wang, T. Huang, X. Wang, and Y. Cao, “Eva-02: A visual representation for neon genesis,” Image Vision Comput. , vol. 149, p. 105171, 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
S. Tong, Z. Liu, Y. Zhai, Y. Ma, Y. LeCun, and S. Xie, “Eyes wide shut? exploring the visual shortcomings of multimodal llms,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR) , 2024, pp. 9568–9578
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Z. Chen, H. Ji, Y. Zhang, Z. Zhu, and Y. Li, “High-resolution feature pyramid network for small object detection on drone view,” IEEE Trans. Circuits Syst. Video Technol. , vol. 34, no. 1, pp. 475–489, 2024
2024
Closest in time.
2024
Closest in time.
Anthropic, “The claude 3 model family: Opus, sonnet, haiku,” 2024. [Online]. Available: https://www.anthropic.com
2024
Closest in time.
W. Wang, Z. Chen, X. Chen, J. Wu, X. Zhu, G. Zeng, P. Luo, T. Lu, J. Zhou, Y. Qiao et al. , “Visionllm: Large language model is also an open-ended decoder for vision-centric tasks,” in Proc. Adv. Neural Inf. Process. Syst. , vol. 36, 2024
2024
Closest in time.
2024
Closest in time.
W. Wang, M. Shi, Q. Li, W. Wang, Z. Huang, L. Xing, Z. Chen, H. Li, X. Zhu, Z. Cao et al. , “The all-seeing project: Towards panoptic visual recognition and understanding of the open world,” in Proc. Int. Conf. Learn. Represent. (ICLR) , 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
J. Zhu, H. Wang, and M. Shi, “Multi-modal large language model enhanced pseudo 3d perception framework for visual commonsense reasoning,” IEEE Trans. Circuits Syst. Video Technol. , pp. 1–1, 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Z. Li, B. Yang, Q. Liu, Z. Ma, S. Zhang, J. Yang, Y. Sun, Y. Liu, and X. Bai, “Monkey: Image resolution and text label are important things for large multi-modal models,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR) , 2024, pp. 26 763–26 773
2024
Closest in time.
2024
Closest in time.
W. Hu, Y. Xu, Y. Li, W. Li, Z. Chen, and Z. Tu, “Bliva: A simple multimodal llm for better handling of text-rich visual questions,” in Proc. AAAI Conf. Artif. Intell. (AAAI) , vol. 38, no. 3, 2024, pp. 2256–2264
2024
Closest in time.