Fetching the paper…
Reading the bibliography…
Medical Visual Question Answering~(VQA) is a combination of medical artificial intelligence and popular VQA challenges.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J.D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al., 2020 · 1901
Earlier work this paper cites.
MIMIC-CXR-JPG, a large publicly available database of labeled chest radiographs
Johnson, A.E., Pollard, T.J., Greenbaum, N.R., Lungren, M.P., Deng, C.y., Peng, Y., Lu, Z., Mark, R.G., Berkowitz, S.J., Horng, S., 2019 · 1901
Earlier work this paper cites.
Simpson, A.L., Antonelli, M., Bakas, S., Bilello, M., Farahani, K., van Ginneken, B., Kopp-Schneider, A., Landman, B.A., Litjens, G., Menze, B., Ronneberger, O., Summers, R.M., Bilic, P., Christ, P.F., Do, R.K.G., Gollub, M., Golia-Pernicka, J., Heckers, S.H., Jarnagin, W.R., McHugo, M.K., Napel, S., Vorontsov, E., Maier-Hein, L., Cardoso, M.J., 2019 · 1902
Earlier work this paper cites.
Long short-term memory
Hochreiter, S., Schmidhuber, J., 1997 · 1997
Earlier work this paper cites.
Bidirectional recurrent neural networks
Schuster, M., Paliwal, K.K., 1997 · 1997
Earlier work this paper cites.
BLEU: a method for automatic evaluation of machine translation, in: Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, Association for Computational Linguistics, Philadelphia, Pennsylvania, USA. pp. 311–318
Papineni, K., Roukos, S., Ward, T., Zhu, W.J., 2002 · 2002
Earlier work this paper cites.
PathVQA: 30000+ questions for medical visual question answering
He, X., Zhang, Y., Mou, L., Xing, E., Xie, P., 2020 · 2003
Earlier work this paper cites.
Lin, M., Chen, Q., Yan, S., 2013 · 2013
Earlier work this paper cites.
Learning phrase representations using RNN encoder–decoder for statistical machine translation, in: Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), Association for Computational Linguistics, Doha, Qatar. pp. 1724–1734
Cho, K., van Merriënboer, B., Gulcehre, C., Bahdanau, D., Bougares, F., Schwenk, H., Bengio, Y., 2014 · 2014
Earlier work this paper cites.
Microsoft COCO: Common objects in context, in: European conference on computer vision, Springer. pp. 740–755
Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., Zitnick, C.L., 2014 · 2014
Earlier work this paper cites.
VQA: Visual question answering, in: 2015 IEEE International Conference on Computer Vision (ICCV), IEEE Computer Society, Los Alamitos, CA, USA. pp. 2425–2433
Antol, S., Agrawal, A., Lu, J., Mitchell, M., Batra, D., Zitnick, C., Parikh, D., 2015 · 2015
Earlier work this paper cites.
The effects of changes in utilization and technological advancements of cross-sectional imaging on radiologist workload
McDonald, R.J., Schwartz, K.M., Eckel, L.J., Diehn, F.E., Hunt, C.H., Bartholmai, B.J., Erickson, B.J., Kallmes, D.F., 2015 · 2015
Earlier work this paper cites.
ImageNet large scale visual recognition challenge
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., et al., 2015 · 2015
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition, in: Proceedings of the 3rd International Conference on Learning Representations
Simonyan, K., Zisserman, A., 2015 · 2015
Earlier work this paper cites.
Neural module networks, in: 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), IEEE Computer Society, Los Alamitos, CA, USA. pp. 39–48
Andreas, J., Rohrbach, M., Darrell, T., Klein, D., 2016 · 2016
Earlier work this paper cites.
Human attention in visual question answering: Do humans and deep networks look at the same regions?, in: Conference on Empirical Methods in Natural Language Processing
Das, A., Agrawal, H., Zitnick, C.L., Parikh, D., Batra, D., 2016 · 2016
Earlier work this paper cites.
Multimodal compact bilinear pooling for visual question answering and visual grounding, in: Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, Association for Computational Linguistics, Austin, Texas. pp. 457–468
Fukui, A., Park, D.H., Yang, D., Rohrbach, A., Darrell, T., Rohrbach, M., 2016 · 2016
Earlier work this paper cites.
Deep residual learning for image recognition, in: 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), IEEE Computer Society, Los Alamitos, CA, USA. pp. 770–778
He, K., Zhang, X., Ren, S., Sun, J., 2016 · 2016
Earlier work this paper cites.
Hierarchical question-image co-attention for visual question answering, in: Advances in Neural Information Processing Systems, pp. 289–297
Lu, J., Yang, J., Batra, D., Parikh, D., 2016 · 2016
Earlier work this paper cites.
YFCC100M: The new data in multimedia research
Thomee, B., Shamma, D.A., Friedland, G., Elizalde, B., Ni, K., Poland, D., Borth, D., Li, L.J., 2016 · 2016
Earlier work this paper cites.
Stacked attention networks for image question answering, in: 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), IEEE Computer Society, Los Alamitos, CA, USA. pp. 21–29
Yang, Z., He, X., Gao, J., Deng, L., Smola, A., 2016 · 2016
Earlier work this paper cites.
Making the V in VQA matter: Elevating the role of image understanding in Visual Question Answering, in: Conference on Computer Vision and Pattern Recognition (CVPR)
Goyal, Y., Khot, T., Summers-Stay, D., Batra, D., Parikh, D., 2017 · 2017
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Krishna, R., Zhu, Y., Groth, O., Johnson, J., Hata, K., Kravitz, J., Chen, S., Kalantidis, Y., Li, L.J., Shamma, D.A., et al., 2017 · 2017
Earlier work this paper cites.
Deep EHR: A survey of recent advances in deep learning techniques for electronic health record (EHR) analysis
Shickel, B., Tighe, P.J., Bihorac, A., Rashidi, P., 2018 · 2017
Earlier work this paper cites.
Attention is all you need, in: Advances in Neural Information Processing Systems, pp. 5998–6008
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, Ł., Polosukhin, I., 2017 · 2017
Earlier work this paper cites.
ChestX-Ray8: Hospital-scale chest X-Ray database and benchmarks on weakly-supervised classification and localization of common thorax diseases, in: 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), IEEE Computer Society, Los Alamitos, CA, USA. pp. 3462–3471
Wang, X., Peng, Y., Lu, L., Lu, Z., Bagheri, M., Summers, R.M., 2017b · 2017
Earlier work this paper cites.
Visual question answering: A survey of methods and datasets
Wu, Q., Teney, D., Wang, P., Shen, C., Dick, A., van den Hengel, A., 2017 · 2017
Earlier work this paper cites.
Multi-modal factorized bilinear pooling with co-attention learning for visual question answering, in: 2017 IEEE International Conference on Computer Vision (ICCV), IEEE Computer Society, Los Alamitos, CA, USA. pp. 1839–1848
Yu, Z., Yu, J., Fan, J., Tao, D., 2017 · 2017
Earlier work this paper cites.
NLM at ImageCLEF 2018 visual question answering in the medical domain., in: CLEF (Working Notes)
Abacha, A.B., Gayen, S., Lau, J.J., Rajaraman, S., Demner-Fushman, D., 2018 · 2018
Earlier work this paper cites.
Don’t just assume; look and answer: Overcoming priors for visual question answering, in: 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), IEEE Computer Society, Los Alamitos, CA, USA. pp. 4971–4980
Agrawal, A., Batra, D., Parikh, D., Kembhavi, A., 2018 · 2018
Earlier work this paper cites.
Deep neural networks and decision tree classifier for visual question answering in the medical domain., in: CLEF (Working Notes)
Allaouzi, I., Ahmed, M.B., 2018 · 2018
Earlier work this paper cites.
A sequence-to-sequence model approach for imageclef 2018 medical domain visual question answering, in: 2018 15th IEEE India Council International Conference (INDICON), pp. 1–6
Ambati, R., Reddy Dudyala, C., 2018 · 2018
Earlier work this paper cites.
Bottom-up and top-down attention for image captioning and visual question answering, in: 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), IEEE Computer Society, Los Alamitos, CA, USA. pp. 6077–6086
Anderson, P., He, X., Buehler, C., Teney, D., Johnson, M., Gould, S., Zhang, L., 2018 · 2018
Earlier work this paper cites.
Overview of ImageCLEF 2018 medical domain visual question answering task., in: CLEF (Working Notes)
Hasan, S.A., Ling, Y., Farri, O., Liu, J., Müller, H., Lungren, M.P., 2018 · 2018
Earlier work this paper cites.
Bilinear attention networks, in: Bengio, S., Wallach, H.M., Larochelle, H., Grauman, K., Cesa-Bianchi, N., Garnett, R. (Eds.), Advances in Neural Information Processing Systems, Montréal, Canada. pp. 1571–1581
Kim, J., Jun, J., Zhang, B., 2018 · 2018
Earlier work this paper cites.
A dataset of clinically generated visual questions and answers about radiology images
Lau, J.J., Gayen, S., Abacha, A.B., Demner-Fushman, D., 2018 · 2018
Earlier work this paper cites.
Multimodal explanations: Justifying decisions and pointing to the evidence, in: 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), IEEE Computer Society, Los Alamitos, CA, USA. pp. 8779–8788
Park, D., Hendricks, L., Akata, Z., Rohrbach, A., Schiele, B., Darrell, T., Rohrbach, M., 2018 · 2018
Cited alongside, same era.
Radiology objects in COntext (ROCO): a multimodal image dataset, in: Intravascular Imaging and Computer Assisted Stenting and Large-Scale Annotation of Biomedical Data and Expert Label Synthesis. Springer, pp. 180–189
Pelka, O., Koitka, S., Rückert, J., Nensa, F., Friedrich, C.M., 2018 · 2018
Cited alongside, same era.
UMass at ImageCLEF medical visual question answering (Med-VQA) 2018 task, in: CLEF (Working Notes)
Peng, Y., Liu, F., Rosen, M.P., 2018 · 2018
Cited alongside, same era.
Improving language understanding by generative pre-training
Radford, A., Narasimhan, K., Salimans, T., Sutskever, I., et al., 2018 · 2018
Cited alongside, same era.
Fantastic answers and where to find them: Immersive question-directed visual attention, in: 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), IEEE Computer Society, Los Alamitos, CA, USA. pp. 2977–2986
Jiang, M., Chen, S., Yang, J., Zhao, Q., 2020 · 2020
Later among the works it cites.
bumjun _ \_ jung at VQA-Med 2020: VQA model based on feature extraction and multi-modal feature fusion, in: CLEF 2020 Working Notes
Jung, B., Gu, L., HaradaAl-Sadi, T., 2020 · 2020
Later among the works it cites.
HARENDRAKV at VQA-Med 2020: Sequential VQA with attention for medical visual question answering, in: CLEF 2020 Working Notes
K. Verma, H., Ramachandran S., S., 2020 · 2020
Later among the works it cites.
Towards visual dialog for radiology, in: Proceedings of the 19th SIGBioMed Workshop on Biomedical Language Processing, pp. 60–69
Kovaleva, O., Shivade, C., Kashyap, S., Kanjaria, K., Wu, J., Ballah, D., Coy, A., Karargyris, A., Guo, Y., Beymer, D.B., et al., 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Overcoming language priors in visual question answering with adversarial regularization, in: Advances in Neural Information Processing Systems, pp. 1541–1551
Ramakrishnan, S., Agrawal, A., Lee, S., 2018 · 2018
Cited alongside, same era.
JUST at VQA-Med: A VGG-Seq2Seq model, in: CLEF (Working Notes)
Talafha, B., Al-Ayyoub, M., 2018 · 2018
Cited alongside, same era.
The HAM10000 dataset, a large collection of multi-source dermatoscopic images of common pigmented skin lesions
Tschandl, P., Rosendahl, C., Kittler, H., 2018 · 2018
Cited alongside, same era.
FVQA: Fact-based visual question answering
Wang, P., Wu, Q., Shen, C., Dick, A., van den Hengel, A., 2018 · 2018
Cited alongside, same era.
Beyond bilinear: Generalized multimodal factorized high-order pooling for visual question answering
Yu, Z., Yu, J., Xiang, C., Fan, J., Tao, D., 2018 · 2018
Cited alongside, same era.
Employing Inception-Resnet-v2 and Bi-LSTM for medical domain visual question answering, in: CLEF (Working Notes)
Zhou, Y., Kang, X., Ren, F., 2018 · 2018
Cited alongside, same era.
JUST at ImageCLEF 2019 visual question answering in the medical domain., in: CLEF (Working Notes)
Al-Sadi, A., Talafha, B., Al-Ayyoub, M., Jararweh, Y., Costen, F., 2019 · 2019
Cited alongside, same era.
An encoder-decoder model for visual question answering in the medical domain., in: CLEF (Working Notes)
Allaouzi, I., Ahmed, M.B., Benamrou, B., 2019 · 2019
Cited alongside, same era.
BioBERT: a pre-trained biomedical language representation model for biomedical text mining
Lee, J., Yoon, W., Kim, S., Kim, D., Kim, S., So, C.H., Kang, J., 2020 · 2020
Later among the works it cites.
AIML at VQA-Med 2020: Knowledge inference via a skeleton-based sentence mapping approach for medical domain visual question answering, in: CLEF 2020 Working Notes
Liao, Z., Wu, Q., Shen, C., van den Hengel, A., Verjans, J., 2020 · 2020
Later among the works it cites.
Shengyan at VQA-Med 2020: An encoder-decoder model for medical domain visual question answering task, in: CLEF 2020 Working Notes
Liu, S., Ding, H., Zhou, X., 2020 · 2020
Later among the works it cites.
CGMVQA: A new classification and generative model for medical visual question answering
Ren, F., Zhou, Y., 2020 · 2020
Later among the works it cites.
NLM at VQA-Med 2020: Visual question answering and generation in the medical domain, in: CLEF 2020 Working Notes
Sarrouti, M., 2020 · 2020
Later among the works it cites.
Human-computer collaboration for skin cancer recognition
Tschandl, P., Rinner, C., Apalla, Z., Argenziano, G., Codella, N., Halpern, A., Janda, M., Lallas, A., Longo, C., Malvehy, J., et al., 2020 · 2020
Later among the works it cites.
kdevqa at VQA-Med 2020: focusing on GLU-based classification, in: CLEF 2020 Working Notes
Umada, H., Aono, M., 2020 · 2020
Later among the works it cites.
A question-centric model for visual question answering in medical imaging
Vu, M.H., Löfstedt, T., Nyholm, T., Sznitman, R., 2020 · 2020
Later among the works it cites.
On the general value of evidence, and bilingual scene-text visual question answering, in: 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), IEEE Computer Society, Los Alamitos, CA, USA. pp. 10123–10132
Wang, X., Liu, Y., Shen, C., Ng, C., Luo, C., Jin, L., Chan, C., van den Hengel, A., Wang, L., 2020 · 2020
Later among the works it cites.
Medical visual question answering via conditional reasoning, in: Proceedings of the 28th ACM International Conference on Multimedia (MM ’20), ACM
Zhan, L.M., Liu, B., Fan, L., Chen, J., Wu, X.M., 2020 · 2020
Later among the works it cites.
Learning from the guidance: Knowledge embedded meta-learning for medical visual question answering, in: Yang, H., Pasupa, K., Leung, A.C.S., Kwok, J.T., Chan, J.H., King, I. (Eds.), Neural Information Processing, Springer International Publishing, Cham. pp. 194–202
Zheng, W., Yan, L., Wang, F.Y., Gou, C., 2020 · 2020
Later among the works it cites.
BBN: Bilateral-branch network with cumulative learning for long-tailed visual recognition , 1–8
Zhou, B., Cui, Q., Wei, X.S., Chen, Z.M., 2020 · 2020
Later among the works it cites.
MVQAS: A Medical Visual Question Answering System. Association for Computing Machinery, New York, NY, USA
Bai, H., Shan, X., Huang, Y., Wang, X., 2021 · 2021
Closest in time.
Overview of the VQA-Med task at ImageCLEF 2021: Visual question answering and generation in the medical domain, in: CLEF 2021 Working Notes, CEUR-WS.org, Bucharest, Romania
Ben Abacha, A., Sarrouti, M., Demner-Fushman, D., Hasan, S.A., Müller, H., 2021 · 2021
Closest in time.
Multiple meta-model quantifying for medical visual question answering, in: de Bruijne, M., Cattin, P.C., Cotin, S., Padoy, N., Speidel, S., Zheng, Y., Essert, C. (Eds.), Medical Image Computing and Computer Assisted Intervention – MICCAI 2021, Springer International Publishing, Cham. pp. 64–74
Do, T., Nguyen, B.X., Tjiputra, E., Tran, M., Tran, Q.D., Nguyen, A., 2021 · 2021
Closest in time.
TeamS at VQA-Med 2021: BBN-Orchestra for long-tailed medical visual question answering
Eslami, S., de Melo, G., Meinel, C., 2021 · 2021
Closest in time.
Optimal deep neural network-based model for answering visual medical question
Gasmi, K., Ltaifa, I.B., Lejeune, G., Alshammari, H., Ammar, L.B., Mahmood, M.A., 2021 · 2021
Closest in time.
Cross-modal self-attention with multi-task pre-training for medical visual question answering, in: Proceedings of the 2021 International Conference on Multimedia Retrieval, Association for Computing Machinery, New York, NY, USA. p. 456–460
Gong, H., Chen, G., Liu, S., Yu, Y., Li, G., 2021a · 2021
Closest in time.
SYSU-HCP at VQA-Med 2021: A data-centric model with efficient training methodology for medical visual question answering
Gong, H., Huang, R., Chen, G., Li, G., 2021b · 2021
Closest in time.
Mmbert: Multimodal bert pretraining for improved medical vqa, in: 2021 IEEE 18th International Symposium on Biomedical Imaging (ISBI), pp. 1033–1036
Khare, Y., Bagal, V., Mathew, M., Devi, A., Priyakumar, U.D., Jawahar, C., 2021 · 2021
Closest in time.
Lijie at ImageCLEFmed VQA-Med 2021: Attention model based on efficient interaction between multimodality
Li, J., Liu, S., 2021 · 2021
Closest in time.
TAM at VQA-Med 2021: A hybrid model with feature extraction and fusion for medical visual question answering
Li, Y., Yang, Z., Hao, T., 2021b · 2021
Closest in time.
Contrastive pre-training and representation distillation for medical visual question answering based on radiology images, in: de Bruijne, M., Cattin, P.C., Cotin, S., Padoy, N., Speidel, S., Zheng, Y., Essert, C. (Eds.), Medical Image Computing and Computer Assisted Intervention – MICCAI 2021, Springer International Publishing, Cham. pp. 210–220
Liu, B., Zhan, L.M., Wu, X.M., 2021a · 2021
Closest in time.
PUC chile team at VQA-Med 2021: approaching vqa as a classfication task via fine-tuning a pretrained cnn
Schilling, R., Messina, P., Parra, D., Lobel, H., 2021 · 2021
Closest in time.
Medfusenet: An attention-based multimodal deep learning model for visual question answering in the medical domain
Sharma, D., Purushotham, S., Reddy, C.K., 2021 · 2021
Closest in time.
SSN MLRG at VQA-Med 2021: An approach for VQA to solve abnormality related queries using improved datasets
Sitara, N.M.S., Kavitha, S., 2021 · 2021
Closest in time.
Yunnan university at VQA-Med 2021: Pretrained BioBERT for medical domain visual question answering
Xiao, Q., Zhou, X., Xiao, Y., Zhao, K., 2021 · 2021
Closest in time.
Yang, Y., Panagopoulou, A., Zhou, S., Jin, D., Callison-Burch, C., Yatskar, M., 2022 · 2022
Closest in time.
Sparks of artificial general intelligence: Early experiments with gpt-4
Bubeck, S., Chandrasekaran, V., Eldan, R., Gehrke, J., Horvitz, E., Kamar, E., Lee, P., Lee, Y.T., Li, Y., Lundberg, S., et al., 2023 · 2023
Closest in time.
Capabilities of gpt-4 on medical challenge problems
Nori, H., King, N., McKinney, S.M., Carignan, D., Horvitz, E., 2023 · 2023
Closest in time.
Prompting large language models with answer heuristics for knowledge-based visual question answering
Shao, Z., Yu, Z., Wang, M., Yu, J., 2023 · 2023
Closest in time.
ChatCAD: Interactive computer-aided diagnosis on medical image using large language models
Wang, S., Zhao, Z., Ouyang, X., Wang, Q., Shen, D., 2023 · 2023
Closest in time.