Fetching the paper…
Reading the bibliography…
Medical visual question answering (VQA) is a challenging multimodal task, where Vision-Language Pre-training (VLP) models can effectively improve the generalization performance.
Vaswani, A., Shazeer, N., Parmar, et al.: Attention is all you need. NIPS 30
2017
Earlier work this paper cites.
Anouk Stein, Carol Wu, C.C., et al.: Rsna pneumonia detection challenge (2018), https://kaggle.com/competitions/rsna-pneumonia-detection-challenge
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
Lau, J.J., Gayen, et al.: A dataset of clinically generated visual questions and answers about radiology images. Scientific data 5
2018
Earlier work this paper cites.
Pelka, O., Koitka, S., Rückert, et al.: Radiology objects in context (roco): a multimodal image dataset. In: LABELS 2018, MICCAI 2018. pp. 180–189 (2018)
2018
Earlier work this paper cites.
2019
Earlier work this paper cites.
Nguyen, B.D., Do, T.T., Nguyen, B.X., et al.: Overcoming data limitation in medical visual question answering. In: MICCAI. pp. 522–530. Cham (2019)
2019
Earlier work this paper cites.
2020
Earlier work this paper cites.
He, K., Fan, H., Wu, et al.: Momentum contrast for unsupervised visual representation learning. pp. 9729–9738 (2020)
2020
Earlier work this paper cites.
Sanjay Subramanian, Lucy Lu Wang, S.M., et al.: MedICaT: A Dataset of Medical Images, Captions, and Textual References. In: Findings of EMNLP (2020)
2020
Earlier work this paper cites.
Do, T., Nguyen, B.X., et al.: Multiple meta-model quantifying for medical visual question answering. In: MICCAI. pp. 64–74. Cham (2021)
2021
Cited alongside, same era.
2021
Cited alongside, same era.
Li, J., Selvaraju, R., Gotmare, et al.: Align before fuse: Vision and language representation learning with momentum distillation. NIPS 34
2021
Cited alongside, same era.
Liu, B., Zhan, L.M., Wu, X.M.: Contrastive pre-training and representation distillation for medical visual question answering based on radiology images. In: MICCAI 2021. pp. 210–220. Springer International Publishing, Cham (2021)
2021
Cited alongside, same era.
He, K., Chen, X., Xie, et al.: Masked autoencoders are scalable vision learners. In: CVPR. pp. 16000–16009 (2022)
2022
Later among the works it cites.
Li, J., Li, D., Xiong, et al.: Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In: ICCV. pp. 12888–12900 (2022)
2022
Later among the works it cites.
Pan, H., He, S., Zhang, K., et al.: Amam: An attention-based multimodal alignment model for medical visual question answering. KBS 255
2022
Later among the works it cites.
Chen, C., Zhong, A., Wu, et al.: Contrastive masked image-text modeling for medical visual representation learning. In: MICCAI. pp. 493–503. Springer (2023)
2023
Later among the works it cites.
Lei, Y., Yang, D., Li, M., Wang, S., Chen, J., Zhang, L.: Text-oriented modality reinforcement network for multimodal sentiment analysis from unaligned multimodal sequences. In: CAAI International Conference on Artificial Intelligence. pp. 189–200. Springer (2023)
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2021
Cited alongside, same era.
Radford, A., Kim, J.W., Hallacy, et al.: Learning transferable visual models from natural language supervision. In: International conference on machine learning. pp. 8748–8763. PMLR (2021)
2021
Cited alongside, same era.
Chen, Z., Du, Y., Hu, et al.: Multi-modal masked autoencoders for medical vision-and-language pre-training. In: MICCAI. pp. 679–689. Springer (2022)
2022
Cited alongside, same era.
Cong, F., Xu, S., et al.: Caption-aware medical vqa via semantic focusing and progressive cross-modality comprehension. In: ACM MM. pp. 3569–3577 (2022)
2022
Cited alongside, same era.
Dou, Z.Y., Xu, Y., Gan, Z., Wang, J., Wang, S., Wang, L., Zhu, C., Zhang, P., Yuan, L., Peng, N., et al.: An empirical study of training end-to-end vision-and-language transformers. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 18166–18176 (2022)
2022
Cited alongside, same era.
Gong, H., Chen, G., Mao, et al.: Vqamix: Conditional triplet mixup for medical visual question answering. IEEE Transactions on Medical Imaging 41
2022
Cited alongside, same era.
2023
Later among the works it cites.
2023
Later among the works it cites.
Li, C., Wong, C., Zhang, S., Usuyama, N., Liu, H., Yang, J., Naumann, T., Poon, H., Gao, J.: Llava-med: Training a large language-and-vision assistant for biomedicine in one day. In: Oh, A., Neumann, T., Globerson, A., Saenko, K., Hardt, M., Levine, S. (eds.) Advances in Neural Information Processing Systems. vol. 36, pp. 28541–28564. Curran Associates, Inc. (2023), https://proceedings.neurips.cc/paper_files/paper/2023/file/5abcdf8ecdcacba028c6662789194572-Paper-Datasets_and_Benchmarks.pdf
2023
Later among the works it cites.
Li, P., Liu, G., He, et al.: Masked vision and language pre-training with unimodal and multimodal contrastive losses for medical visual question answering. In: MICCAI. pp. 374–383. Springer (2023)
2023
Later among the works it cites.