Fetching the paper…
Reading the bibliography…
Medical vision-language models (VLMs) combine computer vision (CV) and natural language processing (NLP) to analyze visual and textual medical data.
“Language Models are Few-Shot Learners”
Tom. Brown et al · 1901
Earlier work this paper cites.
“MIMIC-CXR-JPG, a Large Publicly Available Database of Labeled Chest Radiographs”, 2019
Alistair.. Johnson et al · 1901
Earlier work this paper cites.
“2017 Robotic Instrument Segmentation Challenge”, 2019
Max Allan et al · 1902
Earlier work this paper cites.
“EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks”, 2020
Mingxing Tan and Quoc. Le · 1905
Earlier work this paper cites.
“VisualBERT: A Simple and Performant Baseline for Vision and Language”, 2019
Liunian Li et al · 1908
Earlier work this paper cites.
“Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism”, 2019
Mohammad Shoeybi et al · 1909
Earlier work this paper cites.
“Fine-Tuning Language Models from Human Preferences”, 2020
Daniel. Ziegler et al · 1909
Earlier work this paper cites.
“Recurrent Neural Networks (RNNs): A gentle Introduction and Overview”, 2019
Robin. Schmidt · 1912
Earlier work this paper cites.
“Unifying Vision-and-Language Tasks via Text Generation”
Jaemin Cho, Jie Lei, Hao Tan and Mohit Bansal · 1942
Earlier work this paper cites.
“A Stochastic Approximation Method”
Herbert. Robbins · 1951
Earlier work this paper cites.
““Cloze Procedure”: A New Tool for Measuring Readability”
Wilson. Taylor · 1953
Earlier work this paper cites.
“Long Short-Term Memory”
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
“Reinforcement Learning: An Introduction”
R.S. Sutton and A.G. Barto · 1998
Earlier work this paper cites.
“2018 Robotic Scene Segmentation Challenge”, 2020
Max Allan et al · 2001
Earlier work this paper cites.
“A Simple Framework for Contrastive Learning of Visual Representations”, 2020
Ting Chen, Simon Kornblith, Mohammad Norouzi and Geoffrey Hinton · 2002
Earlier work this paper cites.
“Bleu: a Method for Automatic Evaluation of Machine Translation”
Kishore Papineni, Salim Roukos, Todd Ward and Wei-Jing Zhu · 2002
Earlier work this paper cites.
“PathVQA: 30000+ Questions for Medical Visual Question Answering”, 2020
Xuehai He et al · 2003
Earlier work this paper cites.
“ROUGE: A Package for Automatic Evaluation of Summaries”
Chin-Yew Lin · 2004
Earlier work this paper cites.
Akshay Smit et al · 2004
Earlier work this paper cites.
“METEOR: An Automatic Metric for MT Evaluation with Improved Correlation with Human Judgments”
Satanjeev Banerjee and Alon Lavie · 2005
Earlier work this paper cites.
“ImageNet: A Large-Scale Hierarchical Image Database”
Jia Deng et al · 2009
Earlier work this paper cites.
Yiding Hao et al · 2009
Earlier work this paper cites.
“Stacked Convolutional Auto-Encoders for Hierarchical Feature Extraction”
Jonathan Masci, Ueli Meier, Dan. Ciresan and Jürgen Schmidhuber · 2011
Earlier work this paper cites.
“Distributed Representations of Words and Phrases and Their Compositionality”
Tomas Mikolov et al · 2013
Earlier work this paper cites.
“Efficient Estimation of Word Representations in Vector Space”, 2013
Tomas Mikolov, Kai Chen, Greg Corrado and Jeffrey Dean · 2013
Earlier work this paper cites.
“Encyclopedia of Systems Biology”
Karin Verspoor and Kevin Cohen · 2013
Earlier work this paper cites.
“Learning Phrase Representations using RNN Encoder–Decoder for Statistical Machine Translation”
Kyunghyun Cho et al · 2014
Earlier work this paper cites.
“Adam: A Method for Stochastic Optimization”
Diederik. Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
“Glove: Global Vectors for Word Representation”
Jeffrey Pennington, Richard Socher and Christopher Manning · 2014
Earlier work this paper cites.
“VQA: Visual Question Answering”
Stanislaw Antol et al · 2015
Earlier work this paper cites.
“Preparing a Collection of Radiology Examinations for Distribution and Retrieval”
Dina Demner-Fushman et al · 2015
Earlier work this paper cites.
“Deep Residual Learning for Image Recognition”
Kaiming He, Xiangyu Zhang, Shaoqing Ren and Jian Sun · 2016
Earlier work this paper cites.
“Neural Machine Translation of Rare Words with Subword Units”
Rico Sennrich, Barry Haddow and Alexandra Birch · 2016
Earlier work this paper cites.
Yonghui Wu et al · 2016
Earlier work this paper cites.
“Hierarchical Attention Networks for Document Classification”
Zichao Yang et al · 2016
Earlier work this paper cites.
“Enriching Word Vectors with Subword Information”
Piotr Bojanowski, Edouard Grave, Armand Joulin and Tomas Mikolov · 2017
Earlier work this paper cites.
“NegBio: A High-Performance Tool for Negation and Uncertainty Detection in Radiology Reports”
Yifan Peng et al · 2017
Earlier work this paper cites.
“Attention Is All You Need”
Ashish Vaswani et al · 2017
Earlier work this paper cites.
“Aggregated Residual Transformations for Deep Neural Networks”
Saining Xie et al · 2017
Earlier work this paper cites.
“Bilinear Attention Networks”
Jin-Hwa Kim, Jaehyun Jun and Byoung-Tak Zhang · 2018
Earlier work this paper cites.
“A Dataset of Clinically Generated Visual Questions and Answers about Radiology Images”
Jason Lau, Soumya Gayen, Asma Ben and Dina Demner-Fushman · 2018
Earlier work this paper cites.
“Radiology Objects in COntext (ROCO): A Multimodal Image Dataset”
Obioma Pelka et al · 2018
Earlier work this paper cites.
“Lessons from Natural Language Inference in the Clinical Domain”
Alexey Romanov and Chaitanya Shivade · 2018
Earlier work this paper cites.
“Convolutional Neural Networks: an Overview and Application in Radiology”
Rikiya Yamashita, Mizuho Nishio, Richard Do and Kaori Togashi · 2018
Earlier work this paper cites.
“Grounding Referring Expressions in Images by Variational Context”
Hanwang Zhang, Yulei Niu and Shih-Fu Chang · 2018
Earlier work this paper cites.
“VQA-Med: Overview of the Medical Visual Question Answering Task at ImageCLEF 2019”
Asma Abacha et al · 2019
Earlier work this paper cites.
“Overview of the MEDIQA 2019 Shared Task on Textual Inference, Question Entailment and Question Answering”
Asma Ben, Chaitanya Shivade and Dina Demner-Fushman · 2019
Earlier work this paper cites.
“BERT: Pre-Training of Deep Bidirectional Transformers for Language Understanding”
Jacob Devlin, Ming-Wei Chang, Kenton Lee and Kristina Toutanova · 2019
Earlier work this paper cites.
“Convolutional Networks with Dense Connectivity”
Gao Huang et al · 2019
Earlier work this paper cites.
“CheXpert: A Large Chest Radiograph Dataset with Uncertainty Labels and Expert Comparison”
Jeremy. Irvin et al · 2019
Earlier work this paper cites.
“PubMedQA: A Dataset for Biomedical Research Question Answering”
Qiao Jin et al · 2019
Earlier work this paper cites.
“MIMIC-CXR, a De-Identified Publicly Available Database of Chest Radiographs with Free-Text Reports”
Alistair Johnson et al · 2019
Earlier work this paper cites.
“ALBERT: A Lite BERT for Self-Supervised Learning of Language Representations”
Zhenzhong Lan et al · 2019
Earlier work this paper cites.
“ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks”
Jiasen Lu, Dhruv Batra, Devi Parikh and Stefan Lee · 2019
Earlier work this paper cites.
“Representation Learning with Contrastive Predictive Coding”, 2019
Aaron van Oord, Yazhe Li and Oriol Vinyals · 2019
Earlier work this paper cites.
“Generalized Intersection Over Union: A Metric and a Loss for Bounding Box Regression”
Seyed Rezatofighi et al · 2019
Earlier work this paper cites.
“From Recognition to Cognition: Visual Commonsense Reasoning”
Rowan Zellers, Yonatan Bisk, Ali Farhadi and Yejin Choi · 2019
Earlier work this paper cites.
“BioWordVec, Improving Biomedical Word Embeddings with Subword Information and MeSH”
Yijia Zhang et al · 2019
Earlier work this paper cites.
“Deep Supervised Cross-Modal Retrieval”
Liangli Zhen, Peng Hu, Xu Wang and Dezhong Peng · 2019
Earlier work this paper cites.
“Overview of the VQA-Med Task at ImageCLEF 2020: Visual Question Answering and Generation in the Medical Domain”
Asma Abacha et al · 2020
Earlier work this paper cites.
“Clinical Concept Embeddings Learned from Massive Sources of Multimodal Medical Data”
Andrew Beam et al · 2020
Earlier work this paper cites.
“End-to-End Object Detection with Transformers”
Nicolas Carion et al · 2020
Earlier work this paper cites.
“UNITER: UNiversal Image-TExt Representation Learning”
Yen-Chun Chen et al · 2020
Earlier work this paper cites.
“Reinforcement Learning for Intelligent Healthcare Applications: A Survey”
Antonio Coronato, Muddasar Naeem, Giuseppe De Pietro and Giovanni Paragliola · 2020
Cited alongside, same era.
“5 - Computer vision applications”
Qiang Ji · 2020
Cited alongside, same era.
“Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks”
Patrick Lewis et al · 2020
Cited alongside, same era.
“S2ORC: The Semantic Scholar Open Research Corpus”
Kyle Lo et al · 2020
Cited alongside, same era.
“Framework for Extracting Critical Findings in Radiology Reports”
Thusitha Mabotuwana, Christopher Hall and Nathan Cross · 2020
Cited alongside, same era.
“Deep Learning in Generating Radiology Reports: A Survey”
Maram. Monshi, Josiah Poon and Vera Chung · 2020
Cited alongside, same era.
“MedCLIP: Contrastive Learning from Unpaired Medical Images and Text”, 2022
Zifeng Wang, Zhenbang Wu, Dinesh Agarwal and Jimeng Sun · 2022
Later among the works it cites.
“SimVLM: Simple Visual Language Model Pretraining with Weak Supervision”
Zirui Wang et al · 2022
Later among the works it cites.
“SimMIM: A Simple Framework for Masked Image Modeling”
Zhenda Xie et al · 2022
Later among the works it cites.
“An Improved Transformer Network for Skin Cancer Classification”
Chao Xin et al · 2022
Later among the works it cites.
“A Large Language Model for Electronic Health Records”
Xi Yang et al · 2022
Later among the works it cites.
“Transformers in time-series analysis: A tutorial”
Sabeen Ahmed et al · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“MedICaT: A Dataset of Medical Images, Captions, and Textual References”
Sanjay Subramanian et al · 2020
Cited alongside, same era.
“Neural Machine Translation with Byte-Level Subwords”
Changhan Wang, Kyunghyun Cho and Jiatao Gu · 2020
Cited alongside, same era.
“MedDialog: Large-scale Medical Dialogue Datasets”
Guangtao Zeng et al · 2020
Cited alongside, same era.
“Medical Visual Question Answering via Conditional Reasoning”
Li-Ming Zhan et al · 2020
Cited alongside, same era.
“BERTScore: Evaluating Text Generation with BERT”
Tianyi Zhang et al · 2020
Cited alongside, same era.
“Artificial Intelligence in Healthcare: Transforming the Practice of Medicine”
Junaid Bajwa, Usman Munir, Aditya Nori and Bryan Williams · 2021
Cited alongside, same era.
Later among the works it cites.
Jinze Bai et al · 2023
Later among the works it cites.
“CAT-ViL: Co-attention Gated Vision-Language Embedding for Visual Question Localized-Answering in Robotic Surgery”
Long Bai, Mobarakol Islam and Hongliang Ren · 2023
Later among the works it cites.
“Learning to Exploit Temporal Structure for Biomedical Vision-Language Processing”, 2023
Shruthi Bannur et al · 2023
Later among the works it cites.
“Vision–Language Model for Visual Question Answering in Medical Imagery”
Yakoub Bazi, Mohamad Rahhal, Laila Bashmal and Mansour Zuair · 2023
Later among the works it cites.
“VLP: A Survey on Vision-Language Pre-Training”
Feilong Chen et al · 2023
Later among the works it cites.
“Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90%* ChatGPT Quality”, 2023
Wei-Lin Chiang et al · 2023
Later among the works it cites.
“InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning”, 2023
Wenliang Dai et al · 2023
Later among the works it cites.
“FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning”, 2023
Tri Dao · 2023
Later among the works it cites.
“PubMedCLIP: How Much Does CLIP Benefit Visual Question Answering in the Medical Domain?”
Sedigheh Eslami, Christoph Meinel and Gerard De · 2023
Later among the works it cites.
“A Systematic Survey of Prompt Engineering on Vision-Language Foundation Models”, 2023
Jindong Gu et al · 2023
Later among the works it cites.
“MedAlpaca – An Open-Source Collection of Medical Conversational AI Models and Training Data”, 2023
Tianyu Han et al · 2023
Later among the works it cites.
Kai He et al · 2023
Later among the works it cites.
“Multimodal Image-Text Matching Improves Retrieval-based Chest X-Ray Report Generation”, 2023
Jaehwan Jeong et al · 2023
Later among the works it cites.
Albert. Jiang et al · 2023
Later among the works it cites.
“The Importance of Robust Features in Mitigating Catastrophic Forgetting”
Hikmat Khan, Nidhal Bouaynaya and Ghulam Rasool · 2023
Later among the works it cites.
“Masked Vision and Language Modeling for Multi-modal Representation Learning”, 2023
Gukyeong Kwon et al · 2023
Later among the works it cites.
Hyungyung Lee et al · 2023
Later among the works it cites.
“LLaVA-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One Day”, 2023
Chunyuan Li et al · 2023
Later among the works it cites.
“Masked Vision and Language Pre-Training with Unimodal and Multimodal Contrastive Losses for Medical Visual Question Answering”
Pengfei Li et al · 2023
Later among the works it cites.
“ChatDoctor: A Medical Chat Model Fine-Tuned on a Large Language Model Meta-AI (LLaMA) Using Medical Domain Knowledge”
Yunxiang Li et al · 2023
Later among the works it cites.
“PMC-CLIP: Contrastive Language-Image Pre-Training using Biomedical Documents”, 2023
Weixiong Lin et al · 2023
Later among the works it cites.
“Medical Visual Question Answering: A Survey”
Zhihong Lin et al · 2023
Later among the works it cites.
“A Systematic Review of Deep Learning-based Research on Radiology Report Generation”, 2023
Chang Liu, Yuanhe Tian and Yan Song · 2023
Later among the works it cites.
“Visual Instruction Tuning”, 2023
Haotian Liu, Chunyuan Li, Qingyang Wu and Yong Lee · 2023
Later among the works it cites.
“MedViT: A Robust Vision Transformer for Generalized Medical Image Classification”
Omid Manzari et al · 2023
Later among the works it cites.
“Med-Flamingo: A Multimodal Medical Few-Shot Learner”, 2023
Michael Moor et al · 2023
Later among the works it cites.
Chantal Pellegrini et al · 2023
Later among the works it cites.
“Self-supervised Learning: A Succinct Review”
Veenu Rani et al · 2023
Later among the works it cites.
“Retrieval Augmented Chest X-Ray Report Generation using OpenAI GPT models”, 2023
Mercy Ranjit, Gopinath Ganapathy, Ranjit Manuel and Tanuja Ganu · 2023
Later among the works it cites.
Saurav Sengupta and Donald. Brown · 2023
Later among the works it cites.
“Evolution of Visual Data Captioning Methods, Datasets, and Evaluation Metrics: a Comprehensive Survey”
Dhruv Sharma, Chhavi Dhiman and Dinesh Kumar · 2023
Later among the works it cites.
“Medical Vision Language Pretraining: A survey”, 2023
Prashant Shrestha et al · 2023
Later among the works it cites.
“Visual Med-Alpaca: A Parameter-Efficient Biomedical LLM with Visual Capabilities” [Online; accessed 20-Feb-2024], 2023
Chang Shu et al · 2023
Later among the works it cites.
“Large Language Models Encode Clinical Knowledge”
Karan Singhal et al · 2023
Later among the works it cites.
“Aligning Large Multimodal Models with Factually Augmented RLHF”, 2023
Zhiqing Sun et al · 2023
Later among the works it cites.
“XrayGPT: Chest Radiographs Summarization using Medical Vision-Language Models”, 2023
Omkar Thawkar et al · 2023
Later among the works it cites.
“A Survey on Automatic Generation of Medical Imaging Reports based on Deep Learning”
Pang Ting, Peigao Li and Lijie Zhao · 2023
Later among the works it cites.
“Llama 2: Open Foundation and Fine-Tuned Chat Models”, 2023
Hugo Touvron et al · 2023
Later among the works it cites.
“LLaMA: Open and Efficient Foundation Language Models”, 2023
Hugo Touvron et al · 2023
Later among the works it cites.
“Building Flexible, Scalable, and Machine Learning-ready Multimodal Oncology Datasets”, 2023
Aakash Tripathi et al · 2023
Later among the works it cites.
“A Comprehensive Survey of Continual Learning: Theory, Method and Application”, 2023
Liyuan Wang, Xingxing Zhang, Hang Su and Jun Zhu · 2023
Later among the works it cites.
“Self-Instruct: Aligning Language Models with Self-Generated Instructions”, 2023
Yizhong Wang et al · 2023
Later among the works it cites.
“Multimodal Data Integration for Oncology in the Era of Deep Neural Networks: A Review”, 2023
Asim Waqas et al · 2023
Later among the works it cites.
“Revolutionizing Digital Pathology With the Power of Generative Artificial Intelligence and Foundation Models”
Asim Waqas et al · 2023
Later among the works it cites.
“Evaluating Progress in Automatic Chest X-ray Radiology Report Generation”
Feiyang Yu et al · 2023
Later among the works it cites.
“RAMM: Retrieval-augmented Biomedical Visual Question Answering with Multi-modal Pre-training”, 2023
Zheng Yuan et al · 2023
Later among the works it cites.
“Investigating the Catastrophic Forgetting in Multimodal Large Language Models”, 2023
Yuexiang Zhai et al · 2023
Later among the works it cites.
“Large-Scale Domain-Specific Pretraining for Biomedical Vision-Language Processing”, 2023
Sheng Zhang et al · 2023
Later among the works it cites.
“Adapter Learning in Pretrained Feature Extractor for Continual Learning of Diseases”, 2023
Wentao Zhang et al · 2023
Later among the works it cites.
“Retrieving Multimodal Information for Augmented Generation: A Survey”, 2023
Ruochen Zhao et al · 2023
Later among the works it cites.
“A Survey of Large Language Models in Medicine: Progress, Application, and Challenge”, 2023
Hongjian Zhou et al · 2023
Later among the works it cites.
“Learning without Forgetting for Vision-Language Models”, 2023
Da-Wei Zhou et al · 2023
Later among the works it cites.
“Dynamic Transformer Architecture for Continual Learning of Multimodal Tasks”, 2024
Yuliang Cai and Mohammad Rostami · 2024
Closest in time.
“Brain-Inspired Continual Learning: Robust Feature Distillation and Re-consolidation for Class Incremental Learning”
Hikmat Khan, Nidhal Bouaynaya and Ghulam Rasool · 2024
Closest in time.
“A Survey on Hallucination in Large Vision-Language Models”, 2024
Hanchao Liu et al · 2024
Closest in time.
“Learning or Self-aligning? Rethinking Instruction Fine-tuning”, 2024
Mengjie Ren et al · 2024
Closest in time.