Fetching the paper…
Reading the bibliography…
Vision-language models (VLMs) have been widely applied to 2D medical image analysis due to their ability to align visual and textual representations.
Neocognitron: A self-organizing neural network model for a mechanism of pattern recognition unaffected by shift in position
Fukushima, K · 1980
Earlier work this paper cites.
Handwritten digit recognition with a back-propagation network
LeCun, Y. et al · 1989
Earlier work this paper cites.
Rectified linear units improve restricted boltzmann machines
Nair, V. & Hinton, G. E · 2010
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P · 2014
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
Ronneberger, O., Fischer, P. & Brox, T · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S. & Sun, J · 2016
Earlier work this paper cites.
Ba, J. L · 2016
Earlier work this paper cites.
Gaussian error linear units (gelus)
Hendrycks, D. & Gimpel, K · 2016
Earlier work this paper cites.
A survey on deep learning in medical image analysis
Litjens, G. et al · 2017
Earlier work this paper cites.
Deep learning in medical image analysis
Shen, D., Wu, G. & Suk, H.-I · 2017
Earlier work this paper cites.
Xception: Deep learning with depthwise separable convolutions
Chollet, F · 2017
Earlier work this paper cites.
Potential, challenges and future directions for deep learning in prognostics and health management applications
Fink, O. et al · 2020
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A · 2020
Earlier work this paper cites.
Deep learning-enabled medical computer vision
Esteva, A. et al · 2021
Earlier work this paper cites.
Transunet: Transformers make strong encoders for medical image segmentation
Chen, J. et al · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A. et al · 2021
Earlier work this paper cites.
Retrieval-based chest x-ray report generation using a pre-trained contrastive language-image model
Endo, M., Krishnan, R., Krishna, V., Ng, A. Y. & Rajpurkar, P · 2021
Cited alongside, same era.
Swin unetr: Swin transformers for semantic segmentation of brain tumors in mri images
Hatamizadeh, A. et al · 2021
Cited alongside, same era.
Swin transformer: Hierarchical vision transformer using shifted windows
Liu, Z. et al · 2021
Cited alongside, same era.
Coatnet: Marrying convolution and attention for all data sizes
Dai, Z., Liu, H., Le, Q. V. & Tan, M · 2021
Cited alongside, same era.
Swin-unet: Unet-like pure transformer for medical image segmentation
Cao, H. et al · 2022
Cited alongside, same era.
Uctransnet: rethinking the skip connections in u-net from a channel-wise perspective with transformer
Scaling up your kernels to 31x31: Revisiting large kernel design in cnns
Ding, X., Zhang, X., Han, J. & Ding, G · 2022
Later among the works it cites.
More convnets in the 2020s: Scaling up kernels beyond 51x51 using sparsity
Liu, S. et al · 2022
Later among the works it cites.
Dual cross-attention for medical image segmentation
Ates, G. C., Mohan, P. & Celik, E · 2023
Later among the works it cites.
Pubmedclip: How much does clip benefit visual question answering in the medical domain?
Eslami, S., Meinel, C. & De Melo, G · 2023
Later among the works it cites.
A visual–language foundation model for pathology image analysis using medical twitter
Huang, Z., Bianchi, F., Yuksekgonul, M., Montine, T. J. & Zou, J · 2023
Later among the works it cites.
Sigmoid loss for language image pre-training
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Wang, H., Cao, P., Wang, J. & Zaiane, O. R · 2022
Cited alongside, same era.
Masked autoencoders are scalable vision learners
He, K. et al · 2022
Cited alongside, same era.
Self-supervised pre-training of swin transformers for 3d medical image analysis
Tang, Y. et al · 2022
Cited alongside, same era.
Scaling vision transformers to gigapixel images via hierarchical self-supervised learning
Chen, R. J. et al · 2022
Cited alongside, same era.
Frozen clip models are efficient video learners
Lin, Z. et al · 2022
Cited alongside, same era.
Clip model is an efficient continual learner
Thengane, V., Khan, S., Hayat, M. & Khan, F · 2022
Cited alongside, same era.
Expert-level detection of pathologies from unannotated chest x-ray images via self-supervised learning
Tiu, E. et al · 2022
Cited alongside, same era.
Zhai, X., Mustafa, B., Kolesnikov, A. & Beyer, L · 2023
Later among the works it cites.
Metaformer baselines for vision
Yu, W. et al · 2023
Later among the works it cites.
Internimage: Exploring large-scale vision foundation models with deformable convolutions
Wang, W. et al · 2023
Later among the works it cites.
Medclip-samv2: Towards universal text-driven medical image segmentation
Koleilat, T., Asgariandehkordi, H., Rivaz, H. & Xiao, Y · 2024
Later among the works it cites.
A multimodal generative ai copilot for human pathology
Lu, M. Y. et al · 2024
Later among the works it cites.
A foundation model utilizing chest ct volumes and radiology reports for supervised-level zero-shot detection of abnormalities
Hamamci, I. E. et al · 2024
Later among the works it cites.
Visual instruction tuning
Liu, H., Li, C., Wu, Q. & Lee, Y. J · 2024
Later among the works it cites.
Lisa: Reasoning segmentation via large language model
Lai, X. et al · 2024
Later among the works it cites.
Inceptionnext: When inception meets convnext
Yu, W., Zhou, P., Yan, S. & Wang, X · 2024
Later among the works it cites.
Generatect: Text-conditional generation of 3d chest ct volumes
Hamamci, I. E. et al · 2025
Closest in time.
Med3dvlm: An efficient vision-language model for 3d medical image analysis
Xin, Y., Ates, G. C., Gong, K. & Shao, W · 2025
Closest in time.