Fetching the paper…
Reading the bibliography…
Recently, Multimodal Large Language Models (MLLMs) have demonstrated their immense potential in computer-aided diagnosis and decision-making.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J.D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al., 2020 · 1901
Earlier work this paper cites.
2017 robotic instrument segmentation challenge
Allan, M., Shvets, A., Kurmann, T., Zhang, Z., Duggal, R., Su, Y.H., Rieke, N., Laina, I., Kalavakonda, N., et al., 2019 · 1902
Earlier work this paper cites.
Visualbert: A simple and performant baseline for vision and language
Li, L.H., Yatskar, M., Yin, D., Hsieh, C.J., Chang, K.W., 2019 · 1908
Earlier work this paper cites.
Byte pair encoding: A text compression scheme that accelerates pattern matching
Shibata, Y., Kida, T., Fukamachi, S., Takeda, M., Shinohara, A., Shinohara, T., Arikawa, S., 1999 · 1999
Earlier work this paper cites.
2018 robotic scene segmentation challenge
Allan, M., Kondo, S., Bodenstedt, S., Leger, S., Kadkhodamohammadi, R., Luengo, I., Fuentes, F., Flouty, E., Mohammed, A., Pedersen, M., et al., 2020 · 2001
Earlier work this paper cites.
Song, X., Salcianu, A., Song, Y., Dopson, D., Zhou, D., 2020 · 2012
Earlier work this paper cites.
Endonet: a deep architecture for recognition tasks on laparoscopic videos
Twinanda, A.P., Shehata, S., Mutter, D., Marescaux, J., De Mathelin, M., Padoy, N., 2016 · 2016
Earlier work this paper cites.
Mutan: Multimodal tucker fusion for visual question answering, in: Proceedings of the IEEE international conference on computer vision, pp. 2612–2620
Ben-Younes, H., Cadene, R., Cord, M., Thome, N., 2017 · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I., Hutter, F., 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, Ł., Polosukhin, I., 2017 · 2017
Earlier work this paper cites.
Overview of imageclef 2018 medical domain visual question answering task
Hasan, S.A., Ling, Y., Farri, O., Liu, J., Müller, H., Lungren, M., 2018 · 2018
Earlier work this paper cites.
Beyond bilinear: Generalized multimodal factorized high-order pooling for visual question answering
Yu, Z., Yu, J., Xiang, C., Fan, J., Tao, D., 2018 · 2018
Earlier work this paper cites.
Vqa-med: Overview of the medical visual question answering task at imageclef 2019, in: Proceedings of CLEF (Conference and Labs of the Evaluation Forum) 2019 Working Notes, 9-12 September 2019
Ben Abacha, A., Hasan, S.A., Datla, V.V., Demner-Fushman, D., Müller, H., 2019 · 2019
Earlier work this paper cites.
Block: Bilinear superdiagonal fusion for visual question answering and visual relationship detection, in: Proceedings of the AAAI conference on artificial intelligence, pp. 8102–8109
Ben-Younes, H., Cadene, R., Thome, N., Cord, M., 2019 · 2019
Earlier work this paper cites.
A comprehensive review of robotic surgery curriculum and training for residents, fellows, and postgraduate surgical education
Chen, R., Rodrigues Armijo, P., Krause, C., Force, S.R.T., Siu, K.C., Oleynikov, D., 2020 · 2020
Earlier work this paper cites.
Skill-oriented and performance-driven adaptive curricula for training in robot-assisted surgery using simulators: A feasibility study
Mariani, A., Pellegrini, E., De Momi, E., 2020 · 2020
Earlier work this paper cites.
Effect of COVID-19 on surgical training across the United States: a national survey of general surgery residents
Aziz, H., James, T., Remulla, D., Sher, L., Genyk, Y., Sullivan, M.E., Sheikh, M.R., 2021 · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Hu, E.J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., Chen, W., 2021 · 2021
Earlier work this paper cites.
St-mtl: Spatio-temporal multitask learning model to predict scanpath while tracking instruments in robotic surgery
Islam, M., Vibashan, V., Lim, C.M., Ren, H., 2021 · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision, in: International conference on machine learning, PMLR. pp. 8748–8763
Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al., 2021 · 2021
Earlier work this paper cites.
MedFuseNet: An attention-based multimodal deep learning model for visual question answering in the medical domain
Sharma, D., Purushotham, S., Reddy, C.K., 2021 · 2021
Earlier work this paper cites.
Training data-efficient image transformers & distillation through attention, in: International conference on machine learning, PMLR. pp. 10347–10357
Touvron, H., Cord, M., Douze, M., Massa, F., Sablayrolles, A., Jégou, H., 2021 · 2021
Earlier work this paper cites.
Flamingo: a visual language model for few-shot learning
Alayrac, J.B., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., et al., 2022 · 2022
Earlier work this paper cites.
Contrastive decoding: Open-ended text generation as optimization
Li, X.L., Holtzman, A., Fried, D., Liang, P., Eisner, J., Hashimoto, T., Zettlemoyer, L., Lewis, M., 2022 · 2022
Cited alongside, same era.
Surgical-VQA: Visual Question Answering in Surgical Scenes using Transformer, in: International Conference on Medical Image Computing and Computer-Assisted Intervention, Springer. pp. 33–43
Seenivasan, L., Islam, M., Krishna, A., Ren, H., 2022 · 2022
Cited alongside, same era.
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F.L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al., 2023 · 2023
Cited alongside, same era.
ShennongGPT: A Tuning Chinese LLM for Medication Guidance, in: 2023 IEEE International Conference on Medical Artificial Intelligence (MedAI), IEEE. pp. 67–72
Dou, Y., Zhao, X., Zou, H., Xiao, J., Xi, P., Peng, S., 2023 · 2023
Cited alongside, same era.
Surgical-LLaVA: Toward Surgical Scenario Understanding via Large Language and Vision Models
Jin, J., Jeong, C.W., 2024 · 2024
Later among the works it cites.
Yolov11: An overview of the key architectural enhancements
Khanam, R., Hussain, M., 2024 · 2024
Later among the works it cites.
Natural language understanding and inference with mllm in visual question answering: A survey
Kuang, J., Shen, Y., Xie, J., Luo, H., Xu, Z., Li, R., Li, Y., Cheng, X., Lin, X., Han, Y., 2024 · 2024
Later among the works it cites.
Mitigating object hallucinations in large vision-language models through visual contrastive decoding, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 13872–13882
Leng, S., Zhang, H., Chen, G., Li, X., Lu, S., Miao, C., Bing, L., 2024 · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Gozalo-Brizuela, R., Garrido-Merchan, E.C., 2023 · 2023
Cited alongside, same era.
Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models, in: International conference on machine learning, PMLR. pp. 19730–19742
Li, J., Li, D., Savarese, S., Hoi, S., 2023 · 2023
Cited alongside, same era.
A medical multimodal large language model for future pandemics
Liu, F., Zhu, T., Wu, X., Yang, B., You, C., Wang, C., Lu, L., Liu, Z., Zheng, Y., Sun, X., et al., 2023 · 2023
Cited alongside, same era.
Cholectriplet2021: A benchmark challenge for surgical action triplet recognition
Nwoye, C.I., Alapatt, D., Yu, T., Vardazaryan, A., Xia, F., Zhao, Z., Xia, T., Jia, F., Yang, Y., Wang, H., et al., 2023 · 2023
Cited alongside, same era.
Dinov2: Learning robust visual features without supervision
Oquab, M., Darcet, T., Moutakanni, T., Vo, H., Szafraniec, M., Khalidov, V., Fernandez, P., Haziza, D., Massa, F., El-Nouby, A., et al., 2023 · 2023
Cited alongside, same era.
Rad-restruct: A novel vqa benchmark and method for structured radiology reporting, in: International Conference on Medical Image Computing and Computer-Assisted Intervention, Springer. pp. 409–419
Pellegrini, C., Keicher, M., Özsoy, E., Navab, N., 2023 · 2023
Cited alongside, same era.
SurgicalGPT: end-to-end language-vision GPT for visual question answering in surgery, in: International conference on medical image computing and computer-assisted intervention, Springer. pp. 281–290
Seenivasan, L., Islam, M., Kannan, G., Ren, H., 2023 · 2023
Cited alongside, same era.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al., 2023 · 2023
Cited alongside, same era.
Prior-Posterior Knowledge Prompting-and-Reasoning for Surgical Visual Question Localized-Answering, in: 2024 International Joint Conference on Neural Networks (IJCNN), IEEE. pp. 1–9
Peng, P., Fan, W., Liu, W., Yang, X., Zhou, D., 2024 · 2024
Later among the works it cites.
Capabilities of gemini models in medicine
Saab, K., Tu, T., Weng, W.H., Tanno, R., Stutz, D., Wulczyn, E., Zhang, F., Strother, T., Park, C., Vedadi, E., et al., 2024 · 2024
Later among the works it cites.
GP-VLS: A general-purpose vision language model for surgery
Schmidgall, S., Cho, J., Zakka, C., Hiesinger, W., 2024 · 2024
Later among the works it cites.
PneumoLLM: Harnessing the power of large language model for pneumoconiosis diagnosis
Song, M., Wang, J., Yu, Z., Wang, J., Yang, L., Lu, Y., Li, B., Wang, X., Wang, X., Huang, Q., et al., 2024 · 2024
Later among the works it cites.
Xraygpt: Chest radiographs summarization using large medical vision-language models, in: Proceedings of the 23rd workshop on biomedical natural language processing, pp. 440–448
Thawakar, O.C., Shaker, A.M., Mullappilly, S.S., Cholakkal, H., Anwer, R.M., Khan, S., Laaksonen, J., Khan, F., 2024 · 2024
Later among the works it cites.
Towards generalist biomedical ai
Tu, T., Azizi, S., Driess, D., Schaekermann, M., Amin, M., Chang, P.C., Carroll, A., Lau, C., Tanno, R., Ktena, I., et al., 2024 · 2024
Later among the works it cites.
xgen-mm (blip-3): A family of open large multimodal models
Xue, L., Shu, M., Awadalla, A., Wang, J., Yan, A., Purushwalkam, S., Zhou, H., Prabhu, V., Dai, Y., Ryoo, M.S., et al., 2024 · 2024
Later among the works it cites.
Minicpm-v: A gpt-4v level mllm on your phone
Yao, Y., Yu, T., Zhang, A., Wang, C., Cui, J., Zhu, H., Cai, T., Li, H., Zhao, W., He, Z., et al., 2024 · 2024
Later among the works it cites.
Dual modality prompt learning for visual question-grounded answering in robotic surgery
Zhang, Y., Fan, W., Peng, P., Yang, X., Zhou, D., Wei, X., 2024 · 2024
Later among the works it cites.
Pre-trained multimodal large language model enhances dermatological diagnosis using skingpt-4
Zhou, J., He, X., Sun, L., Xu, J., Chen, X., Chu, Y., Zhou, L., Liao, X., Zhang, B., Afvari, S., et al., 2024 · 2024
Later among the works it cites.
Zhu, Z., Zhang, Y., Cheng, X., Huang, Z., Xu, D., Wu, X., Zheng, Y., 2024 · 2024
Later among the works it cites.
Multitask learning in minimally invasive surgical vision: A review
Alabi, O., Vercauteren, T., Shi, M., 2025 · 2025
Closest in time.
Surgical-VQLA++: Adversarial contrastive learning for calibrated robust visual question-localized answering in robotic surgery
Bai, L., Wang, G., Islam, M., Seenivasan, L., Wang, A., Ren, H., 2025 · 2025
Closest in time.
Enhancing chest x-ray datasets with privacy-preserving large language models and multi-type annotations: a data-driven approach for improved classification
Lanfredi, R.B., Mukherjee, P., Summers, R.M., 2025 · 2025
Closest in time.
Integrating language into medical visual recognition and reasoning: A survey
Lu, Y., Wang, A., 2025 · 2025
Closest in time.
Surgitrack: Fine-grained multi-class multi-tool tracking in surgical videos
Nwoye, C.I., Padoy, N., 2025 · 2025
Closest in time.
Yang, A., Yu, B., Li, C., Liu, D., Huang, F., Huang, H., Jiang, J., Tu, J., Zhang, J., Zhou, J., et al., 2025 · 2025
Closest in time.
Procedure-aware surgical video-language pretraining with hierarchical knowledge augmentation
Yuan, K., Navab, N., Padoy, N., et al., 2025 · 2025
Closest in time.