Fetching the paper…

DeCap: Decoding CLIP Latents for Zero-Shot Captioning via Text-Only Training · Around