Fetching the paper…

Multi-Modal Masked Autoencoders for Medical Vision-and-Language Pre-Training · Around