Fetching the paper…

E2E-VLP: End-to-End Vision-Language Pre-training Enhanced by Visual Learning · Around