Fetching the paper…
Reading the bibliography…
Multimodal Large Language Models (MLLMs) have achieved notable performance in computer vision tasks that require reasoning across visual and textual modalities, yet their capabilities are limited to their pre-trained data, requiring extensive fine-tuning for updates.
Bag-of-visual-words and spatial extensions for land-use classification
Y. Yang and S. Newsam · 2010
Earlier work this paper cites.
Cats and dogs
O. M. Parkhi, A. Vedaldi, A. Zisserman, and C. Jawahar · 2012
Earlier work this paper cites.
Describing textures in the wild, 2013
M. Cimpoi, S. Maji, I. Kokkinos, S. Mohamed, and A. Vedaldi · 2013
Earlier work this paper cites.
Recognition in terra incognita
S. Beery, G. Van Horn, and P. Perona · 2018
Earlier work this paper cites.
The ham10000 dataset, a large collection of multi-source dermatoscopic images of common pigmented skin lesions
P. Tschandl, C. Rosendahl, and H. Kittler · 2018
Earlier work this paper cites.
Eurosat: A novel dataset and deep learning benchmark for land use and land cover classification
P. Helber, B. Bischke, A. Dengel, and D. Borth · 2019
Earlier work this paper cites.
Chexpert: A large chest radiograph dataset with uncertainty labels and expert comparison
J. Irvin, P. Rajpurkar, M. Ko, Y. Yu, S. Ciurea-Ilcus, C. Chute, H. Marklund, B. Haghgoo, R. Ball, K. Shpanskaya, et al · 2019
Earlier work this paper cites.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
P. Lewis, E. Perez, A. Piktus, F. Petroni, V. Karpukhin, N. Goyal, H. Küttler, M. Lewis, W.-t. Yih, T. Rocktäschel, et al · 2020
Earlier work this paper cites.
Learning transferable visual models from natural language supervision, 2021
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever · 2021
Earlier work this paper cites.
A continual development methodology for large-scale multitask dynamic ml systems
A. Gesmundo · 2022
Earlier work this paper cites.
Atlas: Few-shot learning with retrieval augmented language models, 2022
G. Izacard, P. Lewis, M. Lomeli, L. Hosseini, F. Petroni, T. Schick, J. Dwivedi-Yu, A. Joulin, S. Riedel, and E. Grave · 2022
Earlier work this paper cites.
Fives: A fundus image dataset for artificial intelligence based vessel segmentation
K. Jin, X. Huang, J. Zhou, Y. Li, Y. Yan, Y. Sun, Q. Zhang, Y. Wang, and J. Ye · 2022
Cited alongside, same era.
Fixcaps: An improved capsules network for diagnosis of skin cancer
Z. Lan, S. Cai, X. He, and X. Wen · 2022
Cited alongside, same era.
Learning to diagnose common thorax diseases on chest radiographs from radiology reports in vietnamese
T. Nguyen, T. M. Vo, T. V. Nguyen, H. H. Pham, and H. Q. Nguyen · 2022
Cited alongside, same era.
J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al · 2023
Cited alongside, same era.
Extending context window of large language models via positional interpolation
S. Chen, S. Wong, L. Chen, and Y. Tian · 2023
Cited alongside, same era.
Multimodal autoregressive pre-training of large vision encoders
E. Fini, M. Shukor, X. Li, P. Dufter, M. Klein, D. Haldimann, S. Aitharaju, V. G. T. da Costa, L. Béthune, Z. Gan, et al · 2024
Later among the works it cites.
Retrieval-augmented generation for large language models: A survey, 2024
Y. Gao, Y. Xiong, X. Gao, K. Jia, J. Pan, Y. Bi, Y. Dai, J. Sun, M. Wang, and H. Wang · 2024
Later among the works it cites.
Does fine-tuning llms on new knowledge encourage hallucinations?
Z. Gekhman, G. Yona, R. Aharoni, M. Eyal, A. Feder, R. Reichart, and J. Herzig · 2024
Later among the works it cites.
How well does gpt-4v(ision) adapt to distribution shifts? a preliminary investigation, 2024
Z. Han, G. Zhou, R. He, J. Wang, T. Wu, Y. Yin, S. Khan, L. Yao, T. Liu, and K. Zhang · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Challenges and applications of large language models
J. Kaddour, J. Harris, M. Mozes, H. Bradley, R. Raileanu, and R. McHardy · 2023
Cited alongside, same era.
Lost in the middle: How language models use long contexts, 2023
N. F. Liu, K. Lin, J. Hewitt, A. Paranjape, M. Bevilacqua, F. Petroni, and P. Liang · 2023
Cited alongside, same era.
Multimodal large language models: A survey
J. Wu, W. Gan, Z. Chen, S. Wan, and S. Y. Philip · 2023
Cited alongside, same era.
Towards unified and effective domain generalization
Y. Zhang, K. Gong, X. Ding, K. Zhang, F. Lv, K. Keutzer, and X. Yue · 2023
Cited alongside, same era.
Seven failure points when engineering a retrieval augmented generation system, 2024
S. Barnett, S. Kurniawan, S. Thudumu, Z. Brannelly, and M. Abdelrazek · 2024
Cited alongside, same era.
A survey on in-context learning, 2024
Q. Dong, L. Li, D. Dai, C. Zheng, J. Ma, R. Li, H. Xia, J. Xu, Z. Wu, T. Liu, B. Chang, X. Sun, L. Li, and Z. Sui · 2024
Cited alongside, same era.
M. Douze, A. Guzhva, C. Deng, J. Johnson, G. Szilvasy, P.-E. Mazaré, M. Lomeli, L. Hosseini, and H. Jégou · 2024
Cited alongside, same era.
Y. Jiang, J. Irvin, J. H. Wang, M. A. Chaudhry, J. H. Chen, and A. Y. Ng · 2024
Later among the works it cites.
Improved baselines with visual instruction tuning
H. Liu, C. Li, Y. Li, and Y. J. Lee · 2024
Later among the works it cites.
Task arithmetic in the tangent space: Improved editing of pre-trained models
G. Ortiz-Jimenez, A. Favero, and P. Frossard · 2024
Later among the works it cites.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
G. Team, P. Georgiev, V. I. Lei, R. Burnell, L. Bai, A. Gulati, G. Tanzer, D. Vincent, Z. Pan, S. Wang, et al · 2024
Later among the works it cites.
A survey of large language models
W. X. Zhao, K. Zhou, J. Li, T. Tang, X. Wang, Y. Hou, Y. Min, B. Zhang, J. Zhang, Z. Dong, et al · 2024
Later among the works it cites.
A foundation language-image model of the retina (flair): Encoding expert knowledge in text supervision
J. Silva-Rodriguez, H. Chakor, R. Kobbi, J. Dolz, and I. B. Ayed · 2025
Closest in time.
A survey on multimodal large language models
S. Yin, C. Fu, S. Zhao, K. Li, X. Sun, T. Xu, and E. Chen · 2053
Closest in time.