Fetching the paper…
Reading the bibliography…
Large Multimodal Models (LMMs) often face a modality representation gap during pretraining: while language embeddings remain stable, visual representations are highly sensitive to contextual noise (e.g., background clutter).
The Hungarian method for the assignment problem
Kuhn, H. W. 1955 · 1955
Earlier work this paper cites.
Nearest neighbor pattern classification
Cover, T.; and Hart, P. 1967 · 1967
Earlier work this paper cites.
Combinatorial optimization: algorithms and complexity
Papadimitriou, C. H.; and Steiglitz, K. 1998 · 1998
Earlier work this paper cites.
Taming Transformers for High-Resolution Image Synthesis
Esser, P.; Rombach, R.; and Ommer, B. 2021 · 2012
Earlier work this paper cites.
A diagram is worth a dozen images
Kembhavi, A.; Salvato, M.; Kolve, E.; Seo, M.; Hajishirzi, H.; and Farhadi, A. 2016 · 2016
Earlier work this paper cites.
Decoupled Weight Decay Regularization
Loshchilov, I.; and Hutter, F. 2019 · 2019
Earlier work this paper cites.
End-to-End Object Detection with Transformers
Carion, N.; Massa, F.; Synnaeve, G.; Nicolas Usunier, A. K.; and Zagoruyko, S. 2020 · 2020
Earlier work this paper cites.
Conceptual 12M: Pushing Web-Scale Image-Text Pre-Training To Recognize Long-Tail Visual Concepts
Changpinyo, S.; Sharma, P.; Ding, N.; and Soricut, R. 2021 · 2021
Earlier work this paper cites.
Masked Autoencoders Are Scalable Vision Learners
He, K.; Chen, X.; Xie, S.; Li, Y.; Dollár, P.; and Girshick, R. 2021 · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021 · 2021
Earlier work this paper cites.
Masked autoencoders are scalable vision learners
He, K.; Chen, X.; Xie, S.; Li, Y.; Dollár, P.; and Girshick, R. 2022 · 2022
Earlier work this paper cites.
Milan: Masked image pretraining on language assisted representation
Hou, Z.; Sun, F.; Chen, Y.-K.; Xie, Y.; and Kung, S.-Y. 2022 · 2022
Earlier work this paper cites.
Mvp: Multimodality-guided visual pre-training
Wei, L.; Xie, L.; Zhou, W.; Li, H.; and Tian, Q. 2022 · 2022
Earlier work this paper cites.
Achiam, J.; Adler, S.; Agarwal, S.; Ahmad, L.; Akkaya, I.; Aleman, F. L.; Almeida, D.; Altenschmidt, J.; Altman, S.; Anadkat, S.; et al. 2023 · 2023
Cited alongside, same era.
Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90%* ChatGPT Quality
Chiang, W.-L.; Li, Z.; Lin, Z.; Sheng, Y.; Wu, Z.; Zhang, H.; Zheng, L.; Zhuang, S.; Zhuang, Y.; Gonzalez, J. E.; Stoica, I.; and Xing, E. P. 2023 · 2023
Cited alongside, same era.
On implementing 2D rectangular assignment algorithms
Crouse, D. 2023 · 2023
Cited alongside, same era.
Eva: Exploring the limits of masked visual representation learning at scale
Fang, Y.; Wang, W.; Xie, B.; Sun, Q.; Wu, L.; Wang, X.; Huang, T.; Wang, X.; and Cao, Y. 2023 · 2023
Cited alongside, same era.
Unified language-vision pretraining with dynamic discrete visual tokenization
Jin, Y.; Xu, K.; Chen, L.; Liao, C.; Tan, J.; Chen, B.; Lei, C.; Liu, A.; Song, C.; Lei, X.; et al. 2023 · 2023
Masked Image Modeling: A Survey
Hondru, V.; Croitoru, F. A.; Minaee, S.; Ionescu, R. T.; and Sebe, N. 2024 · 2024
Closest in time.
MANTIS: Interleaved Multi-Image Instruction Tuning
Jiang, D.; He, X.; Zeng, H.; Wei, C.; Ku, M.; Liu, Q.; and Chen, W. 2024 · 2024
Closest in time.
What matters when building vision-language models?
Laurençon, H.; Tronchon, L.; Cord, M.; and Sanh, V. 2024 · 2024
Closest in time.
LLaVA-OneVision: Easy Visual Task Transfer
Li, B.; Zhang, Y.; Guo, D.; Zhang, R.; Li, F.; Zhang, H.; Zhang, K.; Zhang, P.; Li, Y.; Liu, Z.; and Li, C. 2024 · 2024
Closest in time.
Vila: On pre-training for visual language models
Lin, J.; Yin, H.; Ping, W.; Molchanov, P.; Shoeybi, M.; and Han, S. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
4M: Massively Multimodal Masked Modeling
Mizrahi, D.; Bachmann, R.; Kar, O. F.; Yeo, T.; Gao, M.; Dehghan, A.; and Zamir, A. 2023 · 2023
Cited alongside, same era.
Llama: Open and efficient foundation language models
Touvron, H.; Lavril, T.; Izacard, G.; Martinet, X.; Lachaux, M.-A.; Lacroix, T.; Rozière, B.; Goyal, N.; Hambro, E.; Azhar, F.; et al. 2023 · 2023
Cited alongside, same era.
Cogvlm: Visual expert for pretrained language models
Wang, W.; Lv, Q.; Yu, W.; Hong, W.; Qi, J.; Wang, Y.; Ji, J.; Yang, Z.; Zhao, L.; Song, X.; et al. 2023 · 2023
Cited alongside, same era.
Rils: Masked visual reconstruction in language semantic space
Yang, S.; Ge, Y.; Yi, K.; Li, D.; Shan, Y.; Qie, X.; and Wang, X. 2023 · 2023
Cited alongside, same era.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
Zhu, D.; Chen, J.; Shen, X.; Li, X.; and Elhoseiny, M. 2023 · 2023
Cited alongside, same era.
Vlmevalkit: An open-source toolkit for evaluating large multi-modality models
Duan, H.; Yang, J.; Qiao, Y.; Fang, X.; Chen, L.; Liu, Y.; Dong, X.; Zang, Y.; Zhang, P.; Wang, J.; et al. 2024 · 2024
Cited alongside, same era.
MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Fu, C.; Chen, P.; Shen, Y.; Qin, Y.; Zhang, M.; Lin, X.; Yang, J.; Zheng, X.; Li, K.; Sun, X.; Wu, Y.; and Ji, R. 2024 · 2024
Cited alongside, same era.
Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs
Tong, S.; Brown, E.; Wu, P.; Woo, S.; Middepogu, M.; Akula, S. C.; Yang, J.; Yang, S.; Iyer, A.; Pan, X.; Wang, Z.; Fergus, R.; LeCun, Y.; and Xie, S. 2024 · 2024
Closest in time.
Reconstructive Visual Instruction Tuning
Wang, H.; Zheng, A.; Zhao, Y.; Wang, T.; Ge, Z.; Zhang, X.; and Zhang, Z. 2024 · 2024
Closest in time.
Show-o: One Single Transformer to Unify Multimodal Understanding and Generation
Xie, J.; Mao, W.; Bai, Z.; Zhang, D. J.; Wang, W.; Lin, K. Q.; Gu, Y.; Chen, Z.; Yang, Z.; and Shou, M. Z. 2024 · 2024
Closest in time.
MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI
Yue, X.; Ni, Y.; Zhang, K.; Zheng, T.; Liu, R.; Zhang, G.; Stevens, S.; Jiang, D.; Ren, W.; Sun, Y.; Wei, C.; Yu, B.; Yuan, R.; Sun, R.; Yin, M.; Zheng, B.; Yang, Z.; Liu, Y.; Huang, W.; Sun, H.; Su, Y.; and Chen, W. 2024 · 2024
Closest in time.
Qwen; :; Yang, A.; Yang, B.; Zhang, B.; Hui, B.; Zheng, B.; Yu, B.; Li, C.; Liu, D.; Huang, F.; Wei, H.; Lin, H.; Yang, J.; Tu, J.; Zhang, J.; Yang, J.; Yang, J.; Zhou, J.; Lin, J.; Dang, K.; Lu, K.; Bao, K.; Yang, K.; Yu, L.; Li, M.; Xue, M.; Zhang, P.; Zhu, Q.; Men, R.; Lin, R.; Li, T.; Tang, T.; Xia, T.; Ren, X.; Ren, X.; Fan, Y.; Su, Y.; Zhang, Y.; Wan, Y.; Liu, Y.; Cui, Z.; Zhang, Z.; and Qiu, Z. 2025 · 2025
Closest in time.
Tschannen, M.; Gritsenko, A.; Wang, X.; Naeem, M. F.; Alabdulmohsin, I.; Parthasarathy, N.; Evans, T.; Beyer, L.; Xia, Y.; Mustafa, B.; Hénaff, O.; Harmsen, J.; Steiner, A.; and Zhai, X. 2025 · 2025
Closest in time.
Grok-1.5V
xAI. 2024 · 2025
Closest in time.