Fetching the paper…
Reading the bibliography…
Multimodal large language models (MLLMs) represent an evolutionary expansion in the capabilities of traditional large language models, enabling them to tackle challenges that surpass the scope of purely text-based applications.
MIMIC-CXR: A large publicly available database of labeled chest radiographs
Alistair E. W. Johnson, Tom J. Pollard, Seth J. Berkowitz, Nathaniel R. Greenbaum, Matthew P. Lungren, Chih-ying Deng, Roger G. Mark, and Steven Horng. 2019 · 1901
Earlier work this paper cites.
Preparing a collection of radiology examinations for distribution and retrieval
Dina Demner-Fushman, Marc D. Kohli, Marc B. Rosenman, Sonya E. Shooshan, Laritza Rodriguez, Sameer K. Antani, George R. Thoma, and Clement J. McDonald. 2016 · 2016
Earlier work this paper cites.
A dataset of clinically generated visual questions and answers about radiology images
Jason J Lau, et al. 2018 · 2018
Earlier work this paper cites.
Radiology Objects in COntext (ROCO): a multimodal image dataset. In MICCAI . Springer, 180–189
Obioma Pelka, et al. 2018 · 2018
Earlier work this paper cites.
Generating Radiology Reports via Memory-driven Transformer. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, Online, November 16-20, 2020 , Bonnie Webber, Trevor Cohn, Yulan He, and Yang Liu (Eds.). Association for Computational Linguistics, 1439–1449
Zhihong Chen, Yan Song, Tsung-Hui Chang, and Xiang Wan. 2020 · 2020
Earlier work this paper cites.
MedICaT: A Dataset of Medical Images, Captions, and Textual References. In Findings of the Association for Computational Linguistics: EMNLP 2020, Online Event, 16-20 November 2020 (Findings of ACL, Vol. EMNLP 2020) , Trevor Cohn, Yulan He, and Yang Liu (Eds.). Association for Computational Linguistics, 2112–2120
Sanjay Subramanian, Lucy Lu Wang, Ben Bogin, Sachin Mehta, Madeleine van Zuylen, Sravanthi Parasa, Sameer Singh, Matt Gardner, and Hannaneh Hajishirzi. 2020 · 2020
Earlier work this paper cites.
Language Models are Few-Shot Learners. In NeurIPS
Brown Tom B, et al. 2020 · 2020
Earlier work this paper cites.
Multiple meta-model quantifying for medical visual question answering. In MICCAI . 64–74
Tuong Do, et al. 2021 · 2021
Earlier work this paper cites.
Taming Transformers for High-Resolution Image Synthesis. In CVPR 2021 . Computer Vision Foundation / IEEE, 12873–12883
Patrick Esser et al. 2021 · 2021
Earlier work this paper cites.
Towards Visual Question Answering on Pathology Images. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, ACL/IJCNLP 2021, (Volume 2: Short Papers), Virtual Event, August 1-6, 2021 , Chengqing Zong, Fei Xia, Wenjie Li, and Roberto Navigli (Eds.). Association for Computational Linguistics, 708–718
Xuehai He, Zhuo Cai, Wenlan Wei, Yichen Zhang, Luntian Mou, Eric P. Xing, and Pengtao Xie. 2021 · 2021
Earlier work this paper cites.
Slake: A Semantically-Labeled Knowledge-Enhanced Dataset For Medical Visual Question Answering. In 18th IEEE International Symposium on Biomedical Imaging, ISBI 2021, Nice, France, April 13-16, 2021 . IEEE, 1650–1654
Bo Liu, Li-Ming Zhan, Li Xu, Lin Ma, Yan Yang, and Xiao-Ming Wu. 2021b · 2021
Earlier work this paper cites.
Exploring and Distilling Posterior and Prior Knowledge for Radiology Report Generation. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2021, virtual, June 19-25, 2021 . Computer Vision Foundation / IEEE, 13753–13762
Fenglin Liu, Xian Wu, Shen Ge, Wei Fan, and Yuexian Zou. 2021a · 2021
Earlier work this paper cites.
Multi-modal masked autoencoders for medical vision-and-language pre-training. In MICCAI . 679–689
Zhihong Chen, et al. 2022 · 2022
Earlier work this paper cites.
Caption-Aware Medical VQA via Semantic Focusing and Progressive Cross-Modality Comprehension. In ACM MM . 3569–3577
Fuze Cong, et al. 2022 · 2022
Earlier work this paper cites.
Flamingo: a Visual Language Model for Few-Shot Learning. In NeurIPS
Jean-Baptiste Alayrac et al. 2022b · 2022
Earlier work this paper cites.
Introducing ChatGPT
OpenAI. 2022 · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Long Ouyang, et al. 2022 · 2022
Cited alongside, same era.
Overview of ImageCLEFmedical 2022 - Caption Prediction and Concept Detection. In Proceedings of the Working Notes of CLEF 2022 - Conference and Labs of the Evaluation Forum, Bologna, Italy, September 5th - to - 8th, 2022 (CEUR Workshop Proceedings, Vol. 3180) , Guglielmo Faggioli, Nicola Ferro, Allan Hanbury, and Martin Potthast (Eds.). CEUR-WS.org, 1294–1307
Johannes Rückert, Asma Ben Abacha, Alba Garcia Seco de Herrera, Louise Bloch, Raphael Brüngel, Ahmad Idrissi-Yaghir, Henning Schäfer, Henning Müller, and Christoph M. Friedrich. 2022 · 2022
Cited alongside, same era.
Minigpt-v2: large language model as a unified interface for vision-language multi-task learning
Jun Chen, et al. 2023 · 2023
Cited alongside, same era.
PaLM: Scaling Language Modeling with Pathways
Aakanksha Chowdhery et al. 2023a · 2023
Cited alongside, same era.
Harnessing the Power of Pre-trained Vision-Language Models for Efficient Medical Report Generation. In CIKM 2023 . ACM, 1308–1317
Qi Li. 2023d · 2023
Later among the works it cites.
Medical visual question answering: A survey
Zhihong Lin, et al. 2023 · 2023
Later among the works it cites.
Parameter-Efficient Transfer Learning for Medical Visual Question Answering
Jiaxiang Liu, et al. 2023a · 2023
Later among the works it cites.
Llava-plus: Learning to use tools for creating multimodal agents
Shilong Liu, et al. 2023b · 2023
Later among the works it cites.
OpenAI. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Augustin Toma et al. 2023b · 2023
Cited alongside, same era.
LLaVA-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One Day
Chunyuan Li et al. 2023c · 2023
Cited alongside, same era.
Visual Med-Alpaca: A Parameter-Efficient Biomedical LLM with Visual Capabilities
Chang Shu et al. 2023d · 2023
Cited alongside, same era.
PMC-LLaMA: Further Finetuning LLaMA on Medical Papers
Chaoyi Wu et al. 2023e · 2023
Cited alongside, same era.
HuaTuo: Tuning LLaMA Model with Chinese Medical Knowledge
Haochun Wang et al. 2023f · 2023
Cited alongside, same era.
DoctorGLM: Fine-tuning your Chinese Doctor is not a Herculean Task
Honglin Xiong et al. 2023g · 2023
Cited alongside, same era.
Med-Flamingo: A Multimodal Medical Few-shot Learner
Moor et al. 2023h · 2023
Cited alongside, same era.
AnyMAL: An Efficient and Scalable Any-Modality Augmented Language Model
Seungwhan Moon et al. 2023j · 2023
Cited alongside, same era.
Hugo Touvron, et al. 2023a · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Hugo Touvron, et al. 2023b · 2023
Later among the works it cites.
Open-Ended Medical Visual Question Answering Through Prefix Tuning of Language Models. In MICCAI , Vol. 14224. Springer, 726–736
Tom van Sonsbeek et al. 2023 · 2023
Later among the works it cites.
Automatic Radiology Report Generation by Learning with Increasingly Hard Negatives. In ECAI 2023 - 26th European Conference on Artificial Intelligence, September 30 - October 4, 2023, Kraków, Poland - Including 12th Conference on Prestigious Applications of Intelligent Systems (PAIS 2023) (Frontiers in Artificial Intelligence and Applications, Vol. 372) , Kobi Gal, Ann Nowé, Grzegorz J. Nalepa, Roy Fairstein, and Roxana Radulescu (Eds.). IOS Press, 2427–2434
Bhanu Prakash Voutharoja, Lei Wang, and Luping Zhou. 2023 · 2023
Later among the works it cites.
MedXChat: Bridging CXR Modalities with a Unified Multimodal Large Model
Ling Yang, Zhanyu Wang, and Luping Zhou. 2023 · 2023
Later among the works it cites.
Kai Zhang, Jun Yu, Zhiling Yan, Yixin Liu, Eashan Adhikarla, Sunyang Fu, Xun Chen, Chen Chen, Yuyin Zhou, Xiang Li, Lifang He, Brian D. Davison, Quanzheng Li, Yong Chen, Hongfang Liu, and Lichao Sun. 2023 · 2023
Later among the works it cites.
A survey of large language models
Wayne Xin Zhao, et al. 2023 · 2023
Later among the works it cites.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
Deyao Zhu, et al. 2023 · 2023
Later among the works it cites.
Cross-Modal Causal Intervention for Medical Report Generation
Weixing Chen, Yang Liu, Ce Wang, Jiarui Zhu, Shen Zhao, Guanbin Li, Cheng-Lin Liu, and Liang Lin. 2024 · 2024
Closest in time.
PECR: Parameter-Efficient Transfer Learning with Cross-Modal Representation Learning for Remote Sensing Visual Question Answering. In ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 6740–6744
Pengfei Li, Jinlong He, Gang Liu, and Shenjun Zhong. 2024 · 2024
Closest in time.