Fetching the paper…
Reading the bibliography…
Endoscopic procedures are essential for diagnosing and treating internal diseases, and multi-modal large language models (MLLMs) are increasingly applied to assist in endoscopy analysis.
Quality indicators for colonoscopy
INDICATORS, S. Q · 2006
Earlier work this paper cites.
Towards automatic polyp detection with a polyp appearance model
Bernal, J., J. Sánchez, F. Vilarino · 2012
Earlier work this paper cites.
Gastric cancer: prevention, screening and early diagnosis
Pasechnikov, V., S. Chukov, E. Fedorov, et al · 2014
Earlier work this paper cites.
Toward embedded detection of polyps in wce images for early diagnosis of colorectal cancer
Silva, J., A. Histace, O. Romain, et al · 2014
Earlier work this paper cites.
The role of endoscopy in inflammatory bowel disease
Shergill, A. K., J. R. Lightdale, D. H. Bruining, et al · 2015
Earlier work this paper cites.
Intraductal biliopancreatic imaging: European society of gastrointestinal endoscopy (esge) technology review
Tringali, A., A. Lemmers, V. Meves, et al · 2015
Earlier work this paper cites.
Wm-dova maps for accurate polyp highlighting in colonoscopy: Validation vs. saliency maps from physicians
Bernal, J., F. J. Sánchez, G. Fernández-Esparrach, et al · 2015
Earlier work this paper cites.
Endonet: a deep architecture for recognition tasks on laparoscopic videos
Twinanda, A. P., S. Shehata, D. Mutter, et al · 2016
Earlier work this paper cites.
Performance measures for lower gastrointestinal endoscopy: a european society of gastrointestinal endoscopy (esge) quality improvement initiative
Kaminski, M. F., S. Thomas-Gibson, M. Bugajski, et al · 2017
Earlier work this paper cites.
Kvasir: A multi-class image dataset for computer aided gastrointestinal disease detection
Pogorelov, K., K. R. Randel, C. Griwodz, et al · 2017
Earlier work this paper cites.
KID project: An internet-based digital video atlas of capsule endoscopy for research purposes
Koulaouzidis, A., D. K. Iakovidis, D. E. Yung, et al · 2017
Earlier work this paper cites.
A benchmark for endoluminal scene segmentation of colonoscopy images
Vázquez, D., J. Bernal, F. J. Sánchez, et al · 2017
Earlier work this paper cites.
A guide to deep learning in healthcare
Esteva, A., A. Robicquet, B. Ramsundar, et al · 2019
Earlier work this paper cites.
2017 robotic instrument segmentation challenge
Allan, M., A. Shvets, T. Kurmann, et al · 2019
Earlier work this paper cites.
Global burden of 369 diseases and injuries in 204 countries and territories, 1990–2019: a systematic analysis for the global burden of disease study 2019
Vos, T., S. S. Lim, C. Abbafati, et al · 2020
Earlier work this paper cites.
2018 robotic scene segmentation challenge
Allan, M., S. Kondo, S. Bodenstedt, et al · 2020
Earlier work this paper cites.
Hyperkvasir, a comprehensive multi-class image and video dataset for gastrointestinal endoscopy
Borgli, H., V. Thambawita, P. H. Smedsrud, et al · 2020
Earlier work this paper cites.
Kvasir-SEG: A segmented polyp dataset
Jha, D., P. H. Smedsrud, M. A. Riegler, et al · 2020
Earlier work this paper cites.
Endoscopy disease detection challenge 2020
Ali, S., N. Ghatwary, B. Braden, et al · 2020
Earlier work this paper cites.
Artificial intelligence in gastroenterology: A state-of-the-art review
Kröner, P. T., M. M. Engels, B. S. Glicksberg, et al · 2021
Earlier work this paper cites.
Kvasir-instrument: Diagnostic and therapeutic tool segmentation dataset in gastrointestinal endoscopy
Jha, D., S. Ali, K. Emanuelsen, et al · 2021
Earlier work this paper cites.
Kvasir-Capsule, a video capsule endoscopy dataset
Smedsrud, P. H., V. Thambawita, S. A. Hicks, et al · 2021
Earlier work this paper cites.
Development of a computer-aided detection system for colonoscopy and a publicly accessible large colonoscopy video database (with video)
Misawa, M., S.-e. Kudo, Y. Mori, et al · 2021
Earlier work this paper cites.
LDPolypVideo benchmark: A large-scale colonoscopy video dataset of diverse polyps
Ma, Y., X. Chen, K. Cheng, et al · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A., J. W. Kim, C. Hallacy, et al · 2021
Earlier work this paper cites.
Impact of artificial intelligence on miss rate of colorectal neoplasia
Wallace, M. B., P. Sharma, P. Bhandari, et al · 2022
Earlier work this paper cites.
Surgical-vqa: Visual question answering in surgical scenes using transformer
Seenivasan, L., M. Islam, A. K. Krishna, et al · 2022
Earlier work this paper cites.
Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Li, J., D. Li, C. Xiong, et al · 2022
Earlier work this paper cites.
Flamingo: a visual language model for few-shot learning
Alayrac, J.-B., J. Donahue, P. Luc, et al · 2022
Earlier work this paper cites.
Towards holistic surgical scene understanding
Valderrama, N., P. Ruiz Puentes, I. Hernández, et al · 2022
Earlier work this paper cites.
Learn to explain: Multimodal reasoning via thought chains for science question answering
Lu, P., S. Mishra, T. Xia, et al · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J., X. Wang, D. Schuurmans, et al · 2022
Earlier work this paper cites.
Self-consistency improves chain of thought reasoning in language models, 2022
Wang, X., J. Wei, D. Schuurmans, et al · 2022
Earlier work this paper cites.
Global burden of digestive diseases: a systematic analysis of the global burden of diseases study, 1990 to 2019
Wang, Y., Y. Huang, R. C. Chase, et al · 2023
Earlier work this paper cites.
Surgicalgpt: end-to-end language-vision gpt for visual question answering in surgery
Seenivasan, L., M. Islam, G. Kannan, et al · 2023
Cited alongside, same era.
Surgical-vqla: Transformer with gated vision-language embedding for visual question localized-answering in robotic surgery
Bai, L., M. Islam, L. Seenivasan, et al · 2023
Cited alongside, same era.
Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Li, J., D. Li, S. Savarese, et al · 2023
Cited alongside, same era.
Instructblip: Towards general-purpose vision-language models with instruction tuning, 2023
Dai, W., J. Li, D. Li, et al · 2023
Cited alongside, same era.
Visual instruction tuning
Liu, H., C. Li, Q. Wu, et al · 2023
Cited alongside, same era.
Med-flamingo: a multimodal medical few-shot learner
Huatuogpt-vision, towards injecting medical visual knowledge into multimodal llms at scale
Chen, J., C. Gui, R. Ouyang, et al · 2024
Later among the works it cites.
Meddr: Diagnosis-guided bootstrapping for large-scale medical vision-language learning
He, S., Y. Nie, Z. Chen, et al · 2024
Later among the works it cites.
Gp-vls: A general-purpose vision language model for surgery
Schmidgall, S., J. Cho, C. Zakka, et al · 2024
Later among the works it cites.
A spectrum evaluation benchmark for medical multi-modal large language models
Liu, J., W. Wang, Y. Su, et al · 2024
Later among the works it cites.
Sepehri, M. S., Z. Fabian, M. Soltanolkotabi, et al · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Moor, M., Q. Huang, S. Wu, et al · 2023
Cited alongside, same era.
Qilin-med: Multi-stage knowledge injection advanced medical large language model
Ye, Q., J. Liu, D. Chong, et al · 2023
Cited alongside, same era.
Zhang, S., Y. Xu, N. Usuyama, et al · 2023
Cited alongside, same era.
Pmc-vqa: Visual instruction tuning for medical visual question answering
Zhang, X., C. Wu, Z. Zhao, et al · 2023
Cited alongside, same era.
GastroVision: A Multi-Class Endoscopy Image Dataset for Computer Aided Gastrointestinal Disease Detection
Jha, D., V. Sharma, N. Dasu, et al · 2023
Cited alongside, same era.
A multi-centre polyp detection and segmentation dataset for generalisability assessment
Ali, S., D. Jha, N. Ghatwary, et al · 2023
Cited alongside, same era.
Judging llm-as-a-judge with mt-bench and chatbot arena
Zheng, L., W.-L. Chiang, Y. Sheng, et al · 2023
Cited alongside, same era.
Later among the works it cites.
Medtrinity-25m: A large-scale multimodal dataset with multigranular annotations for medicine
Xie, Y., C. Zhou, L. Gao, et al · 2024
Later among the works it cites.
Cares: A comprehensive benchmark of trustworthiness in medical vision language models
Xia, P., Z. Chen, J. Tian, et al · 2024
Later among the works it cites.
Handa, P., M. Dhir, A. Mahbod, et al · 2024
Later among the works it cites.
Dinov2: Learning robust visual features without supervision
Oquab, M., T. Darcet, T. Moutakanni, et al · 2024
Later among the works it cites.
Llava-next: Improved reasoning, ocr, and world knowledge, 2024
Liu, H., C. Li, Y. Li, et al · 2024
Later among the works it cites.
Qvq: To see the world with wisdom, 2024
Team, Q · 2024
Later among the works it cites.
Hurst, A., A. Lerer, A. P. Goucher, et al · 2024
Later among the works it cites.
Grattafiori, A., A. Dubey, A. Jauhri, et al · 2024
Later among the works it cites.
Sharegpt4v: Improving large multi-modal models with better captions
Chen, L., J. Li, X. Dong, et al · 2024
Later among the works it cites.
Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi
Yue, X., Y. Ni, K. Zhang, et al · 2024
Later among the works it cites.
Liu, A., B. Feng, B. Xue, et al · 2024
Later among the works it cites.
Medical graph rag: Towards safe medical large language model via graph retrieval-augmented generation, 2024
Wu, J., J. Zhu, Y. Qi · 2024
Later among the works it cites.
Surgical-vqla++: Adversarial contrastive learning for calibrated robust visual question-localized answering in robotic surgery
Bai, L., G. Wang, M. Islam, et al · 2025
Closest in time.
Endochat: Grounded multimodal large language model for endoscopic surgery
Wang, G., L. Bai, J. Wang, et al · 2025
Closest in time.
Janus-pro: Unified multimodal understanding and generation with data and model scaling
Chen, X., Z. Wu, X. Liu, et al · 2025
Closest in time.
Bai, S., K. Chen, X. Liu, et al · 2025
Closest in time.
Vila-m3: Enhancing vision-language models with medical expert knowledge
Nath, V., W. Li, D. Yang, et al · 2025
Closest in time.
A clinically accessible small multimodal radiology model and evaluation metric for chest x-ray findings
Zambrano Chaves, J. M., S.-C. Huang, Y. Xu, et al · 2025
Closest in time.
Towards a multimodal large language model with pixel-level insight for biomedicine
Huang, X., L. Shen, J. Liu, et al · 2025
Closest in time.
Surgical-lvlm: Learning to adapt large vision-language model for grounded visual question answering in robotic surgery
Wang, G., L. Bai, W. J. Nah, et al · 2025
Closest in time.
Medxpertqa: Benchmarking expert-level medical reasoning and understanding
Zuo, Y., S. Qu, Y. Li, et al · 2025
Closest in time.
KyuCapsule Dataset
CapsuleYolo · 2025
Closest in time.
Grok 3 beta — the age of reasoning agents, 2025
xAI · 2025
Closest in time.
Mind your step (by step): Chain-of-thought can reduce performance on tasks where thinking makes humans worse, 2025
Liu, R., Z. Wu, Y. Jurals, et al · 2025
Closest in time.
To cot or not to cot? chain-of-thought helps mainly on math and symbolic reasoning
Sprague, Z., Z. Pu, H. Peng, et al · 2025
Closest in time.
Instruction tuning and cot prompting for contextual medical qa with llms
Le, C., Z. Gong, C. Wang, et al · 2025
Closest in time.
Mmed-rag: Versatile multimodal rag system for medical vision language models
Xia, P., K. Zhu, H. Li, et al · 2025
Closest in time.
Clinical entity augmented retrieval for clinical information extraction
Lopez, I., J. Min, Z.-Y. Zhao, et al · 2025
Closest in time.