Fetching the paper…
Reading the bibliography…
In recent years, the growing demand for medical imaging diagnosis has placed a significant burden on radiologists.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural computation , vol. 9, no. 8, pp. 1735–1780, 1997
1997
Earlier work this paper cites.
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu, “Bleu: a method for automatic evaluation of machine translation,” in Proceedings of the 40th annual meeting of the Association for Computational Linguistics , 2002, pp. 311–318
2002
Earlier work this paper cites.
C.-Y. Lin, “Rouge: A package for automatic evaluation of summaries,” in Text summarization branches out , 2004, pp. 74–81
2004
Earlier work this paper cites.
S. Banerjee and A. Lavie, “Meteor: An automatic metric for mt evaluation with improved correlation with human judgments,” in Proceedings of the acl workshop on intrinsic and extrinsic evaluation measures for machine translation and/or summarization , 2005, pp. 65–72
2005
Earlier work this paper cites.
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan, “Show and tell: A neural image caption generator,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2015, pp. 3156–3164
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
Z.-Q. Cheng, Y. Liu, X. Wu, and X.-S. Hua, “Video ecommerce: Towards online video advertising,” in Proceedings of the 24th ACM international conference on Multimedia , 2016, pp. 1365–1374
2016
Earlier work this paper cites.
Z. Yang, X. He, J. Gao, L. Deng, and A. Smola, “Stacked attention networks for image question answering,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 21–29
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
D. Demner-Fushman, M. D. Kohli, M. B. Rosenman, S. E. Shooshan, L. Rodriguez, S. Antani, G. R. Thoma, and C. J. McDonald, “Preparing a collection of radiology examinations for distribution and retrieval,” Journal of the American Medical Informatics Association , vol. 23, no. 2, pp. 304–310, 2016
2016
Earlier work this paper cites.
Z.-Q. Cheng, X. Wu, Y. Liu, and X.-S. Hua, “Video2shop: Exact matching clothes in videos to online shopping images,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2017, pp. 4048–4056
2017
Earlier work this paper cites.
——, “Video ecommerce++: Toward large scale online video advertising,” IEEE transactions on multimedia , vol. 19, no. 6, pp. 1170–1183, 2017
2017
Earlier work this paper cites.
P. A. Nguyen, Q. Li, Z.-Q. Cheng, Y.-J. Lu, H. Zhang, X. Wu, and C.-W. Ngo, “Vireo@ trecvid 2017: Video-to-text, ad-hoc video search and video hyperlinking,” 2017
2017
Earlier work this paper cites.
Z.-Q. Cheng, H. Zhang, X. Wu, and C.-W. Ngo, “On the selection of anchors and targets for video hyperlinking,” in Proceedings of the 2017 acm on international conference on multimedia retrieval , 2017, pp. 287–293
2017
Earlier work this paper cites.
S. J. Rennie, E. Marcheret, Y. Mroueh, J. Ross, and V. Goel, “Self-critical sequence training for image captioning,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2017, pp. 7008–7024
2017
Earlier work this paper cites.
J. Lu, C. Xiong, D. Parikh, and R. Socher, “Knowing when to look: Adaptive attention via a visual sentinel for image captioning,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2017, pp. 375–383
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
K. Kafle, M. Yousefhussien, and C. Kanan, “Data augmentation for visual question answering,” in Proceedings of the 10th International Conference on Natural Language Generation , 2017, pp. 198–202
2017
Earlier work this paper cites.
C. Finn, P. Abbeel, and S. Levine, “Model-agnostic meta-learning for fast adaptation of deep networks,” in International conference on machine learning . PMLR, 2017, pp. 1126–1135
2017
Earlier work this paper cites.
R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh, and D. Batra, “Grad-cam: Visual explanations from deep networks via gradient-based localization,” in Proceedings of the IEEE international conference on computer vision , 2017, pp. 618–626
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
O. Pelka, S. Koitka, J. Rückert, F. Nensa, and C. M. Friedrich, “Radiology objects in context (roco): a multimodal image dataset,” in Intravascular Imaging and Computer Assisted Stenting and Large-Scale Annotation of Biomedical Data and Expert Label Synthesis: 7th Joint International Workshop, CVII-STENT 2018 and Third International Workshop, LABELS 2018, Held in Conjunction with MICCAI 2018, Granada, Spain, September 16, 2018, Proceedings 3 . Springer, 2018, pp. 180–189
2018
Earlier work this paper cites.
2018
Cited alongside, same era.
J.-H. Kim, J. Jun, and B.-T. Zhang, “Bilinear attention networks,” Advances in neural information processing systems , vol. 31, 2018
2018
Cited alongside, same era.
Y. Li, X. Liang, Z. Hu, and E. P. Xing, “Hybrid retrieval-generation reinforced agent for medical image report generation,” Advances in neural information processing systems , vol. 31, 2018
2018
Cited alongside, same era.
J. J. Lau, S. Gayen, A. Ben Abacha, and D. Demner-Fushman, “A dataset of clinically generated visual questions and answers about radiology images,” Scientific data , vol. 5, no. 1, pp. 1–10, 2018
2018
Cited alongside, same era.
J. Ma, J. Liu, Q. Lin, B. Wu, Y. Wang, and Y. You, “Multitask learning for visual question answering,” IEEE Transactions on neural networks and learning systems , 2021
2021
Later among the works it cites.
X. He, “Towards visual question answering on pathology images.” in Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing , vol. 2, 2021
2021
Later among the works it cites.
T. Do, B. X. Nguyen, E. Tjiputra, M. Tran, Q. D. Tran, and A. Nguyen, “Multiple meta-model quantifying for medical visual question answering,” in Medical Image Computing and Computer Assisted Intervention–MICCAI 2021: 24th International Conference, Strasbourg, France, September 27–October 1, 2021, Proceedings, Part V 24 . Springer, 2021, pp. 64–74
2021
Later among the works it cites.
H. Gong, G. Chen, S. Liu, Y. Yu, and G. Li, “Cross-modal self-attention with multi-task pre-training for medical visual question answering,” in Proceedings of the 2021 international conference on multimedia retrieval , 2021, pp. 456–460
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2019
Cited alongside, same era.
B. D. Nguyen, T.-T. Do, B. X. Nguyen, T. Do, E. Tjiputra, and Q. D. Tran, “Overcoming data limitation in medical visual question answering,” in Medical Image Computing and Computer Assisted Intervention–MICCAI 2019: 22nd International Conference, Shenzhen, China, October 13–17, 2019, Proceedings, Part IV 22 . Springer, 2019, pp. 522–530
2019
Cited alongside, same era.
J. Irvin, P. Rajpurkar, M. Ko, Y. Yu, S. Ciurea-Ilcus, C. Chute, H. Marklund, B. Haghgoo, R. Ball, K. Shpanskaya et al. , “Chexpert: A large chest radiograph dataset with uncertainty labels and expert comparison,” in Proceedings of the AAAI conference on artificial intelligence , vol. 33, no. 01, 2019, pp. 590–597
2019
Cited alongside, same era.
G. Shih, C. C. Wu, S. S. Halabi, M. D. Kohli, L. M. Prevedello, T. S. Cook, A. Sharma, J. K. Amorosa, V. Arteaga, M. Galperin-Aizenberg et al. , “Augmenting the national institutes of health chest radiograph dataset with expert annotations of possible pneumonia,” Radiology: Artificial Intelligence , vol. 1, no. 1, p. e180041, 2019
2019
Cited alongside, same era.
C. Y. Li, X. Liang, Z. Hu, and E. P. Xing, “Knowledge-driven encode, retrieve, paraphrase for medical image report generation,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 33, no. 01, 2019, pp. 6666–6673
2019
Cited alongside, same era.
Z. Chen, Y. Song, T.-H. Chang, and X. Wan, “Generating radiology reports via memory-driven transformer,” in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) , 2020, pp. 1439–1449
2020
Cited alongside, same era.
L.-M. Zhan, B. Liu, L. Fan, J. Chen, and X.-M. Wu, “Medical visual question answering via conditional reasoning,” in Proceedings of the 28th ACM International Conference on Multimedia , 2020, pp. 2345–2354
2020
Cited alongside, same era.
R. Tang, C. Ma, W. E. Zhang, Q. Wu, and X. Yang, “Semantic equivalent adversarial data augmentation for visual question answering,” in Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XIX 16 . Springer, 2020, pp. 437–453
2020
Cited alongside, same era.
2021
Later among the works it cites.
Y. Khare, V. Bagal, M. Mathew, A. Devi, U. D. Priyakumar, and C. Jawahar, “Mmbert: Multimodal bert pretraining for improved medical vqa,” in 2021 IEEE 18th International Symposium on Biomedical Imaging (ISBI) . IEEE, 2021, pp. 1033–1036
2021
Later among the works it cites.
2022
Later among the works it cites.
Y. Zhang, H. Jiang, Y. Miura, C. D. Manning, and C. P. Langlotz, “Contrastive learning of medical visual representations from paired images and text,” in Machine Learning for Healthcare Conference . PMLR, 2022, pp. 2–25
2022
Later among the works it cites.
K. He, X. Chen, S. Xie, Y. Li, P. Dollár, and R. Girshick, “Masked autoencoders are scalable vision learners,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 16 000–16 009
2022
Later among the works it cites.
Z. Chen, Y. Du, J. Hu, Y. Liu, G. Li, X. Wan, and T.-H. Chang, “Multi-modal masked autoencoders for medical vision-and-language pre-training,” in Medical Image Computing and Computer Assisted Intervention–MICCAI 2022: 25th International Conference, Singapore, September 18–22, 2022, Proceedings, Part V . Springer, 2022, pp. 679–689
2022
Later among the works it cites.
F. Wang, Y. Zhou, S. WANG, V. Vardhanabhuti, and L. Yu, “Multi-granularity cross-modal alignment for generalized medical visual representation learning,” in Advances in Neural Information Processing Systems , vol. 35, 2022, pp. 33 536–33 549
2022
Later among the works it cites.
B. Boecking, N. Usuyama, S. Bannur, D. C. Castro, A. Schwaighofer, S. Hyland, M. Wetscherek, T. Naumann, A. Nori, J. Alvarez-Valle et al. , “Making the most of text semantics to improve biomedical vision–language processing,” in Computer Vision–ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part XXXVI . Springer, 2022, pp. 1–21
2022
Later among the works it cites.
H.-Y. Zhou, X. Chen, Y. Zhang, R. Luo, L. Wang, and Y. Yu, “Generalized radiograph representation learning via cross-supervision between images and free-text radiology reports,” Nature Machine Intelligence , vol. 4, no. 1, pp. 32–40, 2022
2022
Later among the works it cites.
H. Qin and Y. Song, “Reinforced cross-modal alignment for radiology report generation,” in Findings of the Association for Computational Linguistics: ACL 2022 , 2022, pp. 448–458
2022
Later among the works it cites.
Z. Wang, H. Han, L. Wang, X. Li, and L. Zhou, “Automated radiographic report generation purely on transformer: A multicriteria supervised approach,” IEEE Transactions on Medical Imaging , vol. 41, no. 10, pp. 2803–2813, 2022
2022
Later among the works it cites.
B. Yan and M. Pei, “Clinical-bert: Vision-language pre-training for radiograph diagnosis and reports generation,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 36, no. 3, 2022, pp. 2982–2990
2022
Later among the works it cites.
H. Gong, G. Chen, M. Mao, Z. Li, and G. Li, “Vqamix: Conditional triplet mixup for medical visual question answering,” IEEE Transactions on Medical Imaging , vol. 41, no. 11, pp. 3332–3343, 2022
2022
Later among the works it cites.
H.-Y. Zhou, C. Lian, L. Wang, and Y. Yu, “Advancing radiograph representation learning with masked record modeling,” in The Eleventh International Conference on Learning Representations , 2023
2023
Closest in time.
C. Wu, X. Zhang, Y. Zhang, Y. Wang, and W. Xie, “Medklip: Medical knowledge enhanced language-image pre-training,” medRxiv , pp. 2023–01, 2023
2023
Closest in time.
K. Zhang, H. Jiang, J. Zhang, Q. Huang, J. Fan, J. Yu, and W. Han, “Semi-supervised medical report generation via graph-guided hybrid feature consistency,” IEEE Transactions on Multimedia , 2023
2023
Closest in time.
2023
Closest in time.
Y. Tay, M. Dehghani, V. Q. Tran, X. Garcia, J. Wei, X. Wang, H. W. Chung, D. Bahri, T. Schuster, S. Zheng et al. , “Ul2: Unifying language learning paradigms,” in The Eleventh International Conference on Learning Representations , 2023
2023
Closest in time.