Fetching the paper…

EVE: Towards End-to-End Video Subtitle Extraction with Vision-Language Models · Around