Fetching the paper…
Reading the bibliography…
Recent Multimodal Large Language Models (MLLMs) excel on benchmark vision-language tasks, yet little is known about how input visual quality shapes their responses.
Microsoft coco captions: Data collection and evaluation server
X. Chen, H. Fang, T.-Y. Lin, R. Vedantam, S. Gupta, P. Dollár, and C. L. Zitnick · 2015
Earlier work this paper cites.
Image super-resolution using deep convolutional networks
C. Dong, C. C. Loy, K. He, and X. Tang · 2016
Earlier work this paper cites.
Photo-realistic single image super-resolution using a generative adversarial network
C. Ledig, L. Theis, F. Huszár, J. Caballero, A. Cunningham, A. Acosta, A. Aitken, A. Tejani, J. Totz, Z. Wang, and W. Shi · 2017
Earlier work this paper cites.
Deep multi-scale convolutional neural network for dynamic scene deblurring
S. Nah, T. Hyun Kim, and K. Mu Lee · 2017
Earlier work this paper cites.
Beyond a gaussian denoiser: Residual learning of deep cnn for image denoising
K. Zhang, W. Zuo, Y. Chen, D. Meng, and L. Zhang · 2017
Earlier work this paper cites.
A high-quality denoising dataset for smartphone cameras
A. Abdelhamed, S. Lin, and M. S. Brown · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
Earlier work this paper cites.
Vizwiz grand challenge: Answering visual questions from blind people
D. Gurari, Q. Li, A. J. Stangl, A. Guo, C. Lin, K. Grauman, J. Luo, and J. P. Bigham · 2018
Earlier work this paper cites.
Deep image prior
D. Ulyanov, A. Vedaldi, and V. Lempitsky · 2018
Earlier work this paper cites.
Nocaps: Novel object captioning at scale
H. Agrawal, K. Desai, Y. Wang, X. Chen, R. Jain, M. Johnson, D. Batra, D. Parikh, S. Lee, and P. Anderson · 2019
Earlier work this paper cites.
Gqa: A new dataset for real-world visual reasoning and compositional question answering
D. A. Hudson and C. D. Manning · 2019
Earlier work this paper cites.
Benchmarking robustness in object detection: Autonomous driving when winter is coming
C. Michaelis, B. Mitzkus, R. Geirhos, E. Rusak, O. Bringmann, A. S. Ecker, M. Bethge, and W. Brendel · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever, et al · 2019
Earlier work this paper cites.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Earlier work this paper cites.
interpreting gpt: the logit lens
nostalgebraist · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu · 2020
Earlier work this paper cites.
Test-time training with self-supervision for generalization under distribution shifts
Q. Sun, A. Tzamarias, and B. Schiele · 2020
Earlier work this paper cites.
Multi-task learning for image super-resolution with auxiliary tasks
C. Yang, Q. Li, M. Luo, F. Wu, and C. Xu · 2020
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Earlier work this paper cites.
Image restoration using very deep convolutional encoder-decoder networks with symmetric skip connections
S. Saha, A. Sinha, and S. Bandyopadhyay · 2021
Earlier work this paper cites.
Tent: Fully test-time adaptation by entropy minimization
D. Wang, J. Bao, X. Dong, J.-Y. Zhu, and J. E. Gonzalez · 2021
Cited alongside, same era.
Restormer: Efficient transformer for high-resolution image restoration
S. W. Zamir, A. Arora, S. H. Khan, M. Hayat, F. Khan, M. Yang, and L. Shao · 2021
Cited alongside, same era.
Simple baselines for image restoration
L. Chen, X. Chu, X. Zhang, and J. Sun · 2022
Cited alongside, same era.
Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
J. Li, D. Li, C. Xiong, and S. Hoi · 2022
Cited alongside, same era.
Learn to explain: Multimodal reasoning via thought chains for science question answering
P. Lu, S. Mishra, T. Xia, L. Qiu, K.-W. Chang, S.-C. Zhu, O. Tafjord, P. Clark, and A. Kalyan · 2022
Cited alongside, same era.
Visual instruction tuning
H. Liu, C. Li, Q. Wu, and Y. J. Lee · 2024
Later among the works it cites.
Llama 3.2: Revolutionizing edge ai and vision with open, customizable models
Meta · 2024
Later among the works it cites.
Lmdrive: Closed-loop end-to-end driving with large language models
H. Shao, Y. Hu, L. Wang, G. Song, S. L. Waslander, Y. Liu, and H. Li · 2024
Later among the works it cites.
Qwen2.5: A party of foundation models, September 2024
Q. Team · 2024
Later among the works it cites.
Drivevlm: The convergence of autonomous driving and large vision-language models
X. Tian, J. Gu, B. Li, Y. Liu, C. Hu, Y. Wang, K. Zhan, P. Jia, X. Lang, and H. Zhao · 2024
Later among the works it cites.
Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M. Yasunaga, A. Aghajanyan, W. Shi, R. James, J. Leskovec, P. Liang, M. Lewis, L. Zettlemoyer, and W.-t. Yih · 2022
Cited alongside, same era.
Bytetrack: Multi-object tracking by associating every detection box
Y. Zhang, P. Sun, Y. Jiang, D. Yu, C. Weng, Z. Yuan, P. Luo, and T. Kong · 2022
Cited alongside, same era.
Mwformer: Multi-weather image restoration using degradation-aware transformers
R. Zhu, Z. Tu, J. Liu, A. C. Bovik, and Y. Fan · 2022
Cited alongside, same era.
MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
C. Fu, P. Chen, Y. Shen, Y. Qin, M. Zhang, X. Lin, J. Yang, X. Zheng, K. Li, X. Sun, Y. Wu, and R. Ji · 2023
Cited alongside, same era.
H. Liu, C. Li, Q. Wu, and Y. J. Lee · 2023
Cited alongside, same era.
Med-flamingo: a multimodal medical few-shot learner
M. Moor, Q. Huang, S. Wu, M. Yasunaga, Y. Dalmia, J. Leskovec, C. Zakka, E. P. Reis, and P. Rajpurkar · 2023
Cited alongside, same era.
Sayplan: Grounding large language models using 3d scene graphs for scalable robot task planning
K. Rana, J. Haviland, S. Garg, J. Abou-Chakra, I. Reid, and N. Suenderhauf · 2023
Cited alongside, same era.
P. Wang, S. Bai, S. Tan, S. Wang, Z. Fan, J. Bai, K. Chen, X. Liu, J. Wang, W. Ge, Y. Fan, K. Dang, M. Du, X. Ren, R. Men, D. Liu, C. Zhou, J. Zhou, and J. Lin · 2024
Later among the works it cites.
Deepseek-vl2: Mixture-of-experts vision-language models for advanced multimodal understanding
Z. Wu, X. Chen, Z. Pan, X. Liu, W. Liu, D. Dai, H. Gao, Y. Ma, C. Wu, B. Wang, et al · 2024
Later among the works it cites.
A. Yang, B. Yang, B. Hui, B. Zheng, B. Yu, C. Zhou, C. Li, C. Li, D. Liu, F. Huang, G. Dong, H. Wei, H. Lin, J. Tang, J. Wang, J. Yang, J. Tu, J. Zhang, J. Ma, J. Xu, J. Zhou, J. Bai, J. He, J. Lin, K. Dang, K. Lu, K. Chen, K. Yang, M. Li, M. Xue, N. Ni, P. Zhang, P. Wang, R. Peng, R. Men, R. Gao, R. Lin, S. Wang, S. Bai, S. Tan, T. Zhu, T. Li, T. Liu, W. Ge, X. Deng, X. Zhou, X. Ren, X. Zhang, X. Wei, X. Ren, Y. Fan, Y. Yao, Y. Zhang, Y. Wan, Y. Chu, Y. Liu, Z. Cui, Z. Zhang, and Z. Fan · 2024
Later among the works it cites.
Scaling up to excellence: Practicing model scaling for photo-realistic image restoration in the wild
F. Yu, J. Gu, Z. Li, J. Hu, X. Kong, X. Wang, J. He, Y. Qiao, and C. Dong · 2024
Later among the works it cites.
Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi
X. Yue, Y. Ni, K. Zhang, T. Zheng, R. Liu, G. Zhang, S. Stevens, D. Jiang, W. Ren, Y. Sun, C. Wei, B. Yu, R. Yuan, R. Sun, M. Yin, B. Zheng, Z. Yang, Y. Liu, W. Huang, H. Sun, Y. Su, and W. Chen · 2024
Later among the works it cites.
Phi-4-mini technical report: Compact yet powerful multimodal language models via mixture-of-loras
A. Abouelenin, A. Ashfaq, A. Atkinson, H. Awadalla, N. Bach, J. Bao, A. Benhaim, M. Cai, V. Chaudhary, C. Chen, et al · 2025
Closest in time.
S. Bai, K. Chen, X. Liu, J. Wang, W. Ge, S. Song, K. Dang, P. Wang, S. Wang, J. Tang, et al · 2025
Closest in time.
Robo2vlm: Visual question answering from large-scale in-the-wild robot manipulation datasets
K. Chen, S. Xie, Z. Ma, and K. Goldberg · 2025
Closest in time.
mrag: Elucidating the design space of multi-modal retrieval-augmented generation
C.-W. Hu, Y. Wang, S. Xing, C.-J. Chen, and Z. Tu · 2025
Closest in time.
Safeflow: A principled protocol for trustworthy and transactional autonomous agent systems
P. Li, X. Zou, Z. Wu, R. Li, S. Xing, H. Zheng, Z. Hu, Y. Wang, H. Li, Q. Yuan, et al · 2025
Closest in time.
V2x-unipool: Unifying multimodal perception and knowledge reasoning for autonomous driving
X. Luo, F. Yang, F. Ding, X. Gao, S. Xing, Y. Zhou, Z. Tu, and C. Liu · 2025
Closest in time.
Position: Prospective of autonomous driving-multimodal llms world models embodied intelligence ai alignment and mamba
Y. Ma, W. Ye, C. Cui, H. Zhang, S. Xing, F. Ke, J. Wang, C. Miao, J. Chen, H. Rezatofighi, et al · 2025
Closest in time.
Generative ai for autonomous driving: Frontiers and opportunities
Y. Wang, S. Xing, C. Can, R. Li, H. Hua, K. Tian, Z. Mo, X. Gao, K. Wu, S. Zhou, et al · 2025
Closest in time.
Mllms know where to look: Training-free perception of small visual details with multimodal llms
J. Zhang, M. Khayatkhoei, P. Chhikara, and F. Ilievski · 2025
Closest in time.