Fetching the paper…

F-VLM: Open-Vocabulary Object Detection upon Frozen Vision and Language Models · Around