Fetching the paper…

FAME-ViL: Multi-Tasking Vision-Language Model for Heterogeneous Fashion Tasks · Around