Fetching the paper…
Reading the bibliography…
Recently a number of studies demonstrated impressive performance on diverse vision-language multi-modal tasks such as image captioning and visual question answering by extending the BERT architecture with multi-modal pre-training objectives.
A comparison of classification algorithms to automatically identify chest x-ray reports that support pneumonia
W. W. Chapman, M. Fizman, B. E. Chapman, and P. J. Haug · 2001
Earlier work this paper cites.
Toward best practices in radiology reporting
C. E. Kahn Jr, C. P. Langlotz, E. S. Burnside, J. A. Carrino, D. S. Channin, D. M. Hovsepian, and D. L. Rubin · 2009
Earlier work this paper cites.
Multimodal medical image retrieval: image categorization to improve search precision
J. Kalpathy-Cramer and W. Hersh · 2010
Earlier work this paper cites.
Improving communication of diagnostic radiology findings through structured reporting
L. H. Schwartz, D. M. Panicek, A. R. Berk, Y. Li, and H. Hricak · 2011
Earlier work this paper cites.
Predicting semantic descriptions from medical images with convolutional neural networks
T. Schlegl, S. M. Waldstein, W.-D. Vogl, U. Schmidt-Erfurth, and G. Langs · 2015
Earlier work this paper cites.
Preparing a collection of radiology examinations for distribution and retrieval
D. Demner-Fushman, M. D. Kohli, M. B. Rosenman, S. E. Shooshan, L. Rodriguez, S. Antani, G. R. Thoma, and C. J. McDonald · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
S. Ren, K. He, R. Girshick, and J. Sun · 2016
Earlier work this paper cites.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Y. Wu, M. Schuster, Z. Chen, Q. V. Le, M. Norouzi, W. Macherey, M. Krikun, Y. Cao, Q. Gao, K. Macherey, et al · 2016
Earlier work this paper cites.
Densely connected convolutional networks
G. Huang, Z. Liu, L. Van Der Maaten, and K. Q. Weinberger · 2017
Earlier work this paper cites.
A survey on deep learning in medical image analysis
G. Litjens, T. Kooi, B. E. Bejnordi, A. A. A. Setio, F. Ciompi, M. Ghafoorian, J. A. Van Der Laak, B. Van Ginneken, and C. I. Sánchez · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Chestx-ray8: Hospital-scale chest x-ray database and benchmarks on weakly-supervised classification and localization of common thorax diseases
X. Wang, Y. Peng, L. Lu, Z. Lu, M. Bagheri, and R. M. Summers · 2017
Earlier work this paper cites.
Bottom-up and top-down attention for image captioning and visual question answering
P. Anderson, X. He, C. Buehler, D. Teney, M. Johnson, S. Gould, and L. Zhang · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
Earlier work this paper cites.
Unsupervised multimodal representation learning across medical images and reports
T.-M. H. Hsu, W.-H. Weng, W. Boag, M. McDermott, and P. Szolovits · 2018
Earlier work this paper cites.
Tienet: Text-image embedding network for common thorax disease classification and reporting in chest x-rays
X. Wang, Y. Peng, L. Lu, Z. Lu, and R. M. Summers · 2018
Cited alongside, same era.
A database for using machine learning and data mining techniques for coronary artery disease diagnosis
R. Alizadehsani, M. Roshanzamir, M. Abdar, A. Beykikhoshk, A. Khosravi, M. Panahiazar, A. Koohestani, F. Khozeimeh, S. Nahavandi, and N. Sarrafzadegan · 2019
Cited alongside, same era.
Publicly available clinical bert embeddings
E. Alsentzer, J. R. Murphy, W. Boag, W.-H. Weng, D. Jin, T. Naumann, and M. McDermott · 2019
Cited alongside, same era.
Uniter: Learning universal image-text representations
Y.-C. Chen, L. Li, L. Yu, A. E. Kholy, F. Ahmed, Z. Gan, Y. Cheng, and J. Liu · 2019
Cited alongside, same era.
Not-so-supervised: a survey of semi-supervised, multi-instance, and transfer learning in medical image analysis
Emixer: End-to-end multimodal x-ray generation via self-supervision
S. Biswal, P. Zhuang, A. Pyrros, N. Siddiqui, S. Koyejo, and J. Sun · 2020
Later among the works it cites.
Multimodal pretraining unmasked: Unifying the vision and language berts
E. Bugliarello, R. Cotterell, N. Okazaki, and D. Elliott · 2020
Later among the works it cites.
Pixel-bert: Aligning image pixels with text by deep multi-modal transformers
Z. Huang, Z. Zeng, B. Liu, D. Fu, and J. Fu · 2020
Later among the works it cites.
Biobert: a pre-trained biomedical language representation model for biomedical text mining
J. Lee, W. Yoon, S. Kim, D. Kim, S. Kim, C. H. So, and J. Kang · 2020
Later among the works it cites.
A comparison of pre-trained vision-and-language models for multimodal representation learning across medical images and reports
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
V. Cheplygina, M. de Bruijne, and J. P. Pluim · 2019
Cited alongside, same era.
Chexpert: A large chest radiograph dataset with uncertainty labels and expert comparison
J. Irvin, P. Rajpurkar, M. Ko, Y. Yu, S. Ciurea-Ilcus, C. Chute, H. Marklund, B. Haghgoo, R. Ball, K. Shpanskaya, et al · 2019
Cited alongside, same era.
Mimic-cxr, a de-identified publicly available database of chest radiographs with free-text reports
A. E. Johnson, T. J. Pollard, S. J. Berkowitz, N. R. Greenbaum, M. P. Lungren, C.-y. Deng, R. G. Mark, and S. Horng · 2019
Cited alongside, same era.
Visualbert: A simple and performant baseline for vision and language
L. H. Li, M. Yatskar, D. Yin, C.-J. Hsieh, and K.-W. Chang · 2019
Cited alongside, same era.
Clinically accurate chest x-ray report generation
G. Liu, T.-M. H. Hsu, M. McDermott, W. Boag, W.-H. Weng, P. Szolovits, and M. Ghassemi · 2019
Cited alongside, same era.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
J. Lu, D. Batra, D. Parikh, and S. Lee · 2019
Cited alongside, same era.
Overcoming data limitation in medical visual question answering
B. D. Nguyen, T.-T. Do, B. X. Nguyen, T. Do, E. Tjiputra, and Q. D. Tran · 2019
Cited alongside, same era.
Transfusion: Understanding transfer learning for medical imaging
M. Raghu, C. Zhang, J. Kleinberg, and S. Bengio · 2019
Cited alongside, same era.
Y. Li, H. Wang, and Y. Luo · 2020
Later among the works it cites.
Imagebert: Cross-modal pre-training with large-scale weak-supervised image-text data
D. Qi, L. Su, J. Song, E. Cui, T. Bharti, and A. Sacheti · 2020
Later among the works it cites.
Learning medical image denoising with deep dynamic residual attention network
S. Sharif, R. A. Naqvi, and M. Biswas · 2020
Later among the works it cites.
A. Smit, S. Jain, P. Rajpurkar, A. Pareek, A. Y. Ng, and M. P. Lungren · 2020
Later among the works it cites.
Unified vision-language pre-training for image captioning and vqa
L. Zhou, H. Palangi, L. Zhang, H. Hu, J. J. Corso, and J. Gao · 2020
Later among the works it cites.
Exploring and distilling posterior and prior knowledge for radiology report generation
F. Liu, X. Wu, S. Ge, W. Fan, and Y. Zou · 2021
Closest in time.
Use of bert (bidirectional encoder representations from transformers)-based deep learning method for extracting evidences in chinese radiology reports: development of a computer-aided liver cancer diagnosis framework
H. Liu, Z. Zhang, Y. Xu, N. Wang, Y. Huang, Z. Yang, R. Jiang, and H. Chen · 2021
Closest in time.
A self-boosting framework for automated radiographic report generation
Z. Wang, L. Zhou, L. Wang, and X. Li · 2021
Closest in time.
Xgpt: Cross-modal generative pre-training for image captioning
Q. Xia, H. Huang, N. Duan, D. Zhang, L. Ji, Z. Sui, E. Cui, T. Bharti, and M. Zhou · 2021
Closest in time.
Writing by memorizing: Hierarchical retrieval-based medical report generation
X. Yang, M. Ye, Q. You, and F. Ma · 2021
Closest in time.
Deep perceptual enhancement for medical image analysis
S. Sharif, R. A. Naqvi, M. Biswas, and W.-K. Loh · 2022
Closest in time.