Fetching the paper…
Reading the bibliography…
Despite significant advancements in medical vision-language pre-training, existing methods have largely overlooked the inherent linguistic complexity and imbalanced isssue within medical reports, as well as the complex cross-modality contextual relationships between texts and images.
Language models are few-shot learners, pp. 1877–1901
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J.D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al., 2020 · 1901
Earlier work this paper cites.
Mimic-cxr-jpg, a large publicly available database of labeled chest radiographs
Johnson, A.E., Pollard, T.J., Greenbaum, N.R., Lungren, M.P., Deng, C.y., Peng, Y., Lu, Z., Mark, R.G., Berkowitz, S.J., Horng, S., 2019b · 1901
Earlier work this paper cites.
Unifying vision-and-language tasks via text generation, in: International Conference on Machine Learning, PMLR. pp. 1931–1942
Cho, J., Lei, J., Tan, H., Bansal, M., 2021 · 1942
Earlier work this paper cites.
Medical subject headings (mesh)
Lipscomb, C.E., 2000 · 2000
Earlier work this paper cites.
Contrastive learning of medical visual representations from paired images and text
Zhang, Y., Jiang, H., Miura, Y., Manning, C.D., Langlotz, C.P., 2020 · 2010
Earlier work this paper cites.
Unimo: Towards unified-modal understanding and generation via cross-modal contrastive learning
Li, W., Gao, C., Niu, G., Xiao, X., Liu, H., Liu, J., Wu, H., Wang, H., 2020a · 2012
Earlier work this paper cites.
Deep residual learning for image recognition, in: Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 770–778
He, K., Zhang, X., Ren, S., Sun, J., 2016 · 2016
Earlier work this paper cites.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Wu, Y., Schuster, M., Chen, Z., Le, Q.V., Norouzi, M., Macherey, W., Krikun, M., Cao, Y., Gao, Q., Macherey, K., et al., 2016 · 2016
Earlier work this paper cites.
Machine learning for medical imaging
Erickson, B.J., Korfiatis, P., Akkus, Z., Kline, T.L., 2017 · 2017
Earlier work this paper cites.
Dermatologist-level classification of skin cancer with deep neural networks
Esteva, A., Kuprel, B., Novoa, R.A., Ko, J., Swetter, S.M., Blau, H.M., Thrun, S., 2017 · 2017
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., Hinton, G.E., 2017 · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I., Hutter, F., 2017 · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.W., Lee, K., Toutanova, K., 2018 · 2018
Earlier work this paper cites.
Hybrid retrieval-generation reinforced agent for medical image report generation, in: NeurIPS
Li, Y., Liang, X., Hu, Z., Xing, E.P., 2018 · 2018
Earlier work this paper cites.
Automated deep-neural-network surveillance of cranial images for acute neurologic events
Titano, J.J., Badgeley, M., Schefflein, J., Pain, M., Su, A., Cai, M., Swinburne, N., Zech, J., Kim, J., Bederson, J., et al., 2018 · 2018
Earlier work this paper cites.
Chexpert: A large chest radiograph dataset with uncertainty labels and expert comparison, in: Proceedings of the AAAI conference on artificial intelligence, pp. 590–597
Irvin, J., Rajpurkar, P., Ko, M., Yu, Y., Ciurea-Ilcus, S., Chute, C., Marklund, H., Haghgoo, B., Ball, R., Shpanskaya, K., et al., 2019 · 2019
Earlier work this paper cites.
Aptos 2019 blindness detection
Karthik, Maggie, S.D., 2019 · 2019
Earlier work this paper cites.
Knowledge-driven encode, retrieve, paraphrase for medical image report generation, in: AAAI, pp. 6666–6673
Li, C.Y., Liang, X., Hu, Z., Xing, E.P., 2019 · 2019
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., et al., 2019 · 2019
Earlier work this paper cites.
Augmenting the national institutes of health chest radiograph dataset with expert annotations of possible pneumonia
Shih, G., Wu, C.C., Halabi, S.S., Kohli, M.D., Prevedello, L.M., Cook, T.S., Sharma, A., Amorosa, J.K., Arteaga, V., Galperin-Aizenberg, M., et al., 2019 · 2019
Earlier work this paper cites.
Siim-acr pneumothorax segmentation
Steven G. Langer, PhD, C., George Shih, MD, M., 2019 · 2019
Cited alongside, same era.
Padchest: A large chest x-ray image dataset with multi-label annotated reports
Bustos, A., Pertusa, A., Salinas, J.M., de la Iglesia-Vayá, M., 2020 · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale, in: ICLR
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al., 2020 · 2020
Cited alongside, same era.
Covid-net: A tailored deep convolutional neural network design for detection of covid-19 cases from chest x-ray images
Wang, L., Lin, Z.Q., Wong, A., 2020 · 2020
Cited alongside, same era.
Bdkg at mediqa 2021: System report for the radiology report summarization task, in: Proceedings of the 20th Workshop on Biomedical Language Processing, pp. 103–111
Dai, S., Wang, Q., Lyu, Y., Zhu, Y., 2021 · 2021
Cited alongside, same era.
Multitask prompted training enables zero-shot task generalization, in: ICLR
Sanh, V., Webson, A., Raffel, C., Bach, S.H., Sutawika, L., Alyafeai, Z., Chaffin, A., Stiegler, A., Le Scao, T., Raja, A., et al., 2022 · 2022
Later among the works it cites.
Expert-level detection of pathologies from unannotated chest x-ray images via self-supervised learning
Tiu, E., Talius, E., Patel, P., Langlotz, C.P., Ng, A.Y., Rajpurkar, P., 2022 · 2022
Later among the works it cites.
Multi-granularity cross-modal alignment for generalized medical visual representation learning
Wang, F., Zhou, Y., Wang, S., Vardhanabhuti, V., Yu, L., 2022 · 2022
Later among the works it cites.
Generalized radiograph representation learning via cross-supervision between images and free-text radiology reports
Zhou, H.Y., Chen, X., Zhang, Y., Luo, R., Wang, L., Yu, Y., 2022 · 2022
Later among the works it cites.
Learning to exploit temporal structure for biomedical vision-language processing, in: CVPR, pp. 15016–15027
Bannur, S., Hyland, S., Liu, Q., Perez-Garcia, F., Ilse, M., Castro, D.C., Boecking, B., Sharma, H., Bouzid, K., Thieme, A., et al., 2023 · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Masked autoencoders are scalable vision learners, in: CVPR, pp. 15979–15988
He, K., Chen, X., Xie, S., Li, Y., Doll’ar, P., Girshick, R.B., 2021 · 2021
Cited alongside, same era.
Gloria: A multimodal global-local representation learning framework for label-efficient medical image recognition, in: ICCV, pp. 3942–3951
Huang, S.C., Shen, L., Lungren, M.P., Yeung, S., 2021 · 2021
Cited alongside, same era.
Radgraph: Extracting clinical entities and relations from radiology reports, in: Vanschoren, J., Yeung, S. (Eds.), Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks, Curran
Jain, S., Agrawal, A., Saporta, A., Truong, S., Duong, D.N.D.N., Bui, T., Chambon, P., Zhang, Y., Lungren, M., Ng, A., Langlotz, C., Rajpurkar, P., Rajpurkar, P., 2021 · 2021
Cited alongside, same era.
Learning calibrated medical image segmentation via multi-rater agreement modeling, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 12341–12351
Ji, W., Yu, S., Wu, J., Ma, K., Bian, C., Bi, Q., Li, J., Liu, H., Cheng, L., Zheng, Y., 2021 · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision, in: International Conference on Machine Learning, PMLR. pp. 8748–8763
Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al., 2021 · 2021
Cited alongside, same era.
Finetuned language models are zero-shot learners, in: ICLR
Wei, J., Bosma, M., Zhao, V., Guu, K., Yu, A.W., Lester, B., Du, N., Dai, A.M., Le, Q.V., 2021 · 2021
Cited alongside, same era.
Vinvl: Revisiting visual representations in vision-language models, in: CVPR, pp. 5579–5588
Zhang, P., Li, X., Hu, X., Yang, J., Zhang, L., Wang, L., Choi, Y., Gao, J., 2021 · 2021
Cited alongside, same era.
Closest in time.
Contrastive masked image-text modeling for medical visual representation learning, Springer. pp. 493–503
Chen, C., Zhong, A., Wu, D., Luo, J., Li, Q., 2023 · 2023
Closest in time.
Prior: Prototype representation joint learning from medical images and reports, in: ICCV, pp. 21361–21371
Cheng, P., Lin, L., Lyu, J., Huang, Y., Luo, W., Tang, X., 2023 · 2023
Closest in time.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
Chiang, W.L., Li, Z., Lin, Z., Sheng, Y., Wu, Z., Zhang, H., Zheng, L., Zhuang, S., Zhuang, Y., Gonzalez, J.E., Stoica, I., Xing, E.P., 2023 · 2023
Closest in time.
OpenAI, 2023 · 2023
Closest in time.
Xplainer: From x-ray observations to explainable zero-shot diagnosis
Pellegrini, C., Keicher, M., Özsoy, E., Jiraskova, P., Braren, R., Navab, N., 2023 · 2023
Closest in time.
Silva-Rodriguez, J., Chakor, H., Kobbi, R., Dolz, J., Ayed, I.B., 2023 · 2023
Closest in time.
Distilled gpt for source code summarization
Su, C.Y., Mcmillan, C., 2023 · 2023
Closest in time.
Med-unic: Unifying cross-lingual medical vision-language pre-training by diminishing bias, in: NeurIPS
Wan, Z., Liu, C., Zhang, M., Fu, J., Wang, B., Cheng, S., Ma, L., Quilodrán-Casas, C., Arcucci, R., 2023 · 2023
Closest in time.
Medklip: Medical knowledge enhanced language-image pre-training, in: ICCV
Wu, C., Zhang, X., Zhang, Y., Wang, Y., Xie, W., 2023 · 2023
Closest in time.
Medim: Boost medical image representation via radiology report-guided masking, Springer. pp. 13–23
Xie, Y., Gu, L., Harada, T., Zhang, J., Xia, Y., Wu, Q., 2023 · 2023
Closest in time.
Mlip: Enhancing medical visual representation with divergence encoder and knowledge-guided contrastive learning, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 11704–11714
Li, Z., Yang, L.T., Ren, B., Nie, X., Gao, Z., Tan, C., Li, S.Z., 2024 · 2024
Closest in time.
Rethinking masked image modeling for medical image representation
Xie, Y., Gu, L., Harada, T., Zhang, J., Xia, Y., Wu, Q., 2024b · 2024
Closest in time.
Continual self-supervised learning: Towards universal multi-modal medical data representation learning, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 11114–11124
Ye, Y., Xie, Y., Zhang, J., Chen, Z., Wu, Q., Xia, Y., 2024 · 2024
Closest in time.
Chestx-ray8: Hospital-scale chest x-ray database and benchmarks on weakly-supervised classification and localization of common thorax diseases, in: Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 2097–2106
Wang, X., Peng, Y., Lu, L., Lu, Z., Bagheri, M., Summers, R.M., 2017 · 2097
Closest in time.