Fetching the paper…
Reading the bibliography…
We present HealthGPT, a powerful Medical Large Vision-Language Model (Med-LVLM) that integrates medical visual comprehension and generation capabilities within a unified autoregressive paradigm.
MIMIC-CXR-JPG, a large publicly available database of labeled chest radiographs
Johnson, A. E.; Pollard, T. J.; Greenbaum, N. R.; Lungren, M. P.; Deng, C.-y.; Peng, Y.; Lu, Z.; Mark, R. G.; Berkowitz, S. J.; and Horng, S. 2019 · 1901
Earlier work this paper cites.
A multiscale visualization of attention in the transformer model
Vig, J. 2019 · 1906
Earlier work this paper cites.
Bbdm: Image-to-image translation with brownian bridge diffusion models
Li, B.; Xue, K.; Liu, B.; and Lai, Y.-K. 2023a · 1961
Earlier work this paper cites.
Pathvqa: 30000+ questions for medical visual question answering
He, X.; Zhang, Y.; Mou, L.; Xing, E.; and Xie, P. 2020 · 2003
Earlier work this paper cites.
Mixture of experts: a literature survey
Masoudnia, S.; and Ebrahimpour, R. 2014 · 2014
Earlier work this paper cites.
Image-to-image translation with conditional adversarial networks
Isola, P.; Zhu, J.-Y.; Zhou, T.; and Efros, A. A. 2017 · 2017
Earlier work this paper cites.
Unpaired image-to-image translation using cycle-consistent adversarial networks
Zhu, J.-Y.; Park, T.; Isola, P.; and Efros, A. A. 2017 · 2017
Earlier work this paper cites.
A dataset of clinically generated visual questions and answers about radiology images
Lau, J. J.; Gayen, S.; Ben Abacha, A.; and Demner-Fushman, D. 2018 · 2018
Earlier work this paper cites.
Taming transformers for high-resolution image synthesis
Esser, P.; Rombach, R.; and Ommer, B. 2021 · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Hu, E. J.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; and Chen, W. 2021 · 2021
Earlier work this paper cites.
Slake: A semantically-labeled knowledge-enhanced dataset for medical visual question answering
Liu, B.; Zhan, L.-M.; Xu, L.; Ma, L.; Yang, Y.; and Wu, X.-M. 2021 · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021 · 2021
Earlier work this paper cites.
Unified-io: A unified model for vision, language, and multi-modal tasks
Lu, J.; Clark, C.; Zellers, R.; Mottaghi, R.; and Kembhavi, A. 2022 · 2022
Earlier work this paper cites.
Medclip: Contrastive learning from unpaired medical images and text
Wang, Z.; Wu, Z.; Agarwal, D.; and Sun, J. 2022 · 2022
Earlier work this paper cites.
InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
Dai, W.; Li, J.; Li, D.; Tiong, A. M. H.; Zhao, J.; Wang, W.; Li, B.; Fung, P.; and Hoi, S. 2023 · 2023
Earlier work this paper cites.
Parameter-efficient fine-tuning of large-scale pre-trained language models
Ding, N.; Qin, Y.; Yang, G.; Wei, F.; Yang, Z.; Su, Y.; Hu, S.; Chen, Y.; Chan, C.-M.; Chen, W.; et al. 2023 · 2023
Cited alongside, same era.
Dreamllm: Synergistic multimodal comprehension and creation
Dong, R.; Han, C.; Peng, Y.; Qi, Z.; Ge, Z.; Yang, J.; Zhao, L.; Sun, J.; Zhou, H.; Wei, H.; et al. 2023 · 2023
Cited alongside, same era.
Planting a seed of vision in large language model
Ge, Y.; Ge, Y.; Zeng, Z.; Wang, X.; and Shan, Y. 2023 · 2023
Cited alongside, same era.
Visual Instruction Tuning
Liu, H.; Li, C.; Wu, Q.; and Lee, Y. J. 2023 · 2023
Cited alongside, same era.
Med-flamingo: a multimodal medical few-shot learner
Moor, M.; Huang, Q.; Wu, S.; Yasunaga, M.; Dalmia, Y.; Leskovec, J.; Zakka, C.; Reis, E. P.; and Rajpurkar, P. 2023 · 2023
Cited alongside, same era.
Dubey, A.; Jauhri, A.; Pandey, A.; Kadian, A.; Al-Dahle, A.; Letman, A.; Mathur, A.; Schelten, A.; Yang, A.; Fan, A.; et al. 2024 · 2024
Later among the works it cites.
Seed-x: Multimodal models with unified multi-granularity comprehension and generation
Ge, Y.; Zhao, S.; Zhu, J.; Ge, Y.; Yi, K.; Song, L.; Li, C.; Ding, X.; and Shan, Y. 2024 · 2024
Later among the works it cites.
Omnimedvqa: A new large-scale comprehensive evaluation benchmark for medical lvlm
Hu, Y.; Li, T.; Lu, Q.; Shao, W.; He, J.; Qiao, Y.; and Luo, P. 2024 · 2024
Later among the works it cites.
Teamlora: Boosting low-rank adaptation with expert collaboration and competition
Lin, T.; Liu, J.; Zhang, W.; Li, Z.; Dai, Y.; Li, H.; Yu, Z.; He, W.; Li, J.; Jiang, H.; et al. 2024 · 2024
Later among the works it cites.
LLaVA-NeXT: Improved reasoning, OCR, and world knowledge
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
GPT-4V(ision) System Card
OpenAI. 2023 · 2023
Cited alongside, same era.
Large language models encode clinical knowledge
Singhal, K.; Azizi, S.; Tu, T.; Mahdavi, S. S.; Wei, J.; Chung, H. W.; Scales, N.; Tanwani, A.; Cole-Lewis, H.; Pfohl, S.; et al. 2023 · 2023
Cited alongside, same era.
Xraygpt: Chest radiographs summarization using medical vision-language models
Thawkar, O.; Shaker, A.; Mullappilly, S. S.; Cholakkal, H.; Anwer, R. M.; Khan, S.; Laaksonen, J.; and Khan, F. S. 2023 · 2023
Cited alongside, same era.
SynthRAD2023 Grand Challenge dataset: Generating synthetic CT for radiotherapy
Thummerer, A.; van der Bijl, E.; Galapon Jr, A.; Verhoeff, J. J.; Langendijk, J. A.; Both, S.; van den Berg, C. N. A.; and Maspero, M. 2023 · 2023
Cited alongside, same era.
The role of large language models in medical image processing: a narrative review
Tian, D.; Jiang, S.; Zhang, L.; Lu, X.; and Xu, Y. 2023 · 2023
Cited alongside, same era.
Next-gpt: Any-to-any multimodal llm
Wu, S.; Fei, H.; Qu, L.; Ji, W.; and Chua, T.-S. 2023 · 2023
Cited alongside, same era.
A survey of large language models in medicine: Progress, application, and challenge
Zhou, H.; Liu, F.; Gu, B.; Zou, X.; Huang, J.; Wu, J.; Li, Y.; Chen, S. S.; Zhou, P.; Liu, J.; et al. 2023 · 2023
Cited alongside, same era.
Liu, H.; Li, C.; Li, Y.; Li, B.; Zhang, Y.; Shen, S.; and Lee, Y. J. 2024c · 2024
Later among the works it cites.
Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision Language Audio and Action
Lu, J.; Clark, C.; Lee, S.; Zhang, Z.; Khosla, S.; Marten, R.; Hoiem, D.; and Kembhavi, A. 2024 · 2024
Later among the works it cites.
Vila-m3: Enhancing vision-language models with medical expert knowledge
Nath, V.; Li, W.; Yang, D.; Myronenko, A.; Zheng, M.; Lu, Y.; Liu, Z.; Yin, H.; Law, Y. M.; Tang, Y.; et al. 2024 · 2024
Later among the works it cites.
Auto-Encoding Morph-Tokens for Multimodal LLM
Pan, K.; Tang, S.; Li, J.; Fan, Z.; Chow, W.; Yan, S.; Chua, T.-S.; Zhuang, Y.; and Zhang, H. 2024 · 2024
Later among the works it cites.
Chameleon: Mixed-modal early-fusion foundation models
Team, C. 2024 · 2024
Later among the works it cites.
MetaMorph: Multimodal Understanding and Generation via Instruction Tuning
Tong, S.; Fan, D.; Zhu, J.; Xiong, Y.; Chen, X.; Sinha, K.; Rabbat, M.; LeCun, Y.; Xie, S.; and Liu, Z. 2024 · 2024
Later among the works it cites.
Towards generalist biomedical AI
Tu, T.; Azizi, S.; Driess, D.; Schaekermann, M.; Amin, M.; Chang, P.-C.; Carroll, A.; Lau, C.; Tanno, R.; Ktena, I.; et al. 2024 · 2024
Later among the works it cites.
Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation
Wu, C.; Chen, X.; Wu, Z.; Ma, Y.; Liu, X.; Pan, Z.; Liu, W.; Xie, Z.; Yu, X.; Ruan, C.; and Luo, P. 2024 · 2024
Later among the works it cites.
Show-o: One single transformer to unify multimodal understanding and generation
Xie, J.; Mao, W.; Bai, Z.; Zhang, D. J.; Wang, W.; Lin, K. Q.; Gu, Y.; Chen, Z.; Yang, Z.; and Shou, M. Z. 2024 · 2024
Later among the works it cites.
Yi: Open foundation models by 01. ai
Young, A.; Chen, B.; Li, C.; Huang, C.; Zhang, G.; Zhang, G.; Li, H.; Zhu, J.; Chen, J.; Chang, J.; et al. 2024 · 2024
Later among the works it cites.
The IXI Dataset
Davies, R. L.; Royston, P. A.; Leung, M. S.; Haider, M. E. A. M. J.; Barkhof, S. G. A. L.; and B., P. E. T. M. 2014 · 2025
Closest in time.