Zero-Shot Composed Image Retrieval with Textual Inversion
Alberto Baldrati, Lorenzo Agnolucci, Marco Bertini, and Alberto Del Bimbo · 2023
Later among the works it cites.
Structure and content-guided video synthesis with diffusion models
Patrick Esser, Johnathan Chiu, Parmida Atighehchian, Jonathan Granskog, and Anastasis Germanidis · 2023
Later among the works it cites.
Decap: Decoding CLIP latents for zero-shot captioning via text-only training
Original
Wei Li, Linchao Zhu, Longyin Wen, and Yi Yang · 2023
Later among the works it cites.
Clip-guided vision-language pre-training for question answering in 3d scenes
Maria Parelli, Alexandros Delitzas, Nikolas Hars, Georgios Vlassis, Sotirios Anagnostidis, Gregor Bachmann, and Thomas Hofmann · 2023
Later among the works it cites.
Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation
Nataniel Ruiz, Yuanzhen Li, Varun Jampani, Yael Pritch, Michael Rubinstein, and Kfir Aberman · 2023
Later among the works it cites.
Towards Understanding the Modality Gap in CLIP
Peiyang Shi, Michael C Welle, Mårten Björkman, and Danica Kragic · 2023
Later among the works it cites.
SuS-X: Training-Free Name-Only Transfer of Vision-Language Models
Vishaal Udandarao, Ankush Gupta, and Samuel Albanie · 2023
Later among the works it cites.
Sigmoid loss for language image pre-training
Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, and Lucas Beyer · 2023
Later among the works it cites.
Diagnosing and Rectifying Vision Models using Language
Yuhui Zhang, Jeff Z HaoChen, Shih-Cheng Huang, Kuan-Chieh Wang, James Zou, and Serena Yeung · 2023
Later among the works it cites.
iSEARLE: Improving Textual Inversion for Zero-Shot Composed Image Retrieval
Original
Lorenzo Agnolucci, Alberto Baldrati, Marco Bertini, and Alberto Del Bimbo · 2024
Later among the works it cites.
The llama 3 herd of models
Original
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al · 2024
Later among the works it cites.
Datacomp: In search of the next generation of multimodal datasets
Samir Yitzhak Gadre, Gabriel Ilharco, Alex Fang, Jonathan Hayase, Georgios Smyrnis, Thao Nguyen, Ryan Marten, Mitchell Wortsman, Dhruba Ghosh, Jieyu Zhang, et al · 2024
Later among the works it cites.
Towards flexible perception with visual memory
Original
Robert Geirhos, Priyank Jaini, Austin Stone, Sourabh Medapati, Xi Yi, George Toderici, Abhijit Ogale, and Jonathon Shlens · 2024
Later among the works it cites.
Visual instruction tuning
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee · 2024
Later among the works it cites.
Eclipse: A resource-efficient text-to-image prior for image generations
Maitreya Patel, Changhoon Kim, Sheng Cheng, Chitta Baral, and Yezhou Yang · 2024
Later among the works it cites.
Two Effects, One Trigger: On the Modality Gap, Object Bias, and Information Imbalance in Contrastive Vision-Language Representation Learning
Original
Simon Schrodi, David T Hoffmann, Max Argus, Volker Fischer, and Thomas Brox · 2024
Later among the works it cites.
Leveraging Cross-Modal Neighbor Representation for Improved CLIP Classification
Chao Yi, Lu Ren, De-Chuan Zhan, and Han-Jia Ye · 2024
Later among the works it cites.
Avid: Any-length video inpainting with diffusion model
Zhixing Zhang, Bichen Wu, Xiaoyan Wang, Yaqiao Luo, Luxin Zhang, Yinan Zhao, Peter Vajda, Dimitris Metaxas, and Licheng Yu · 2024
Later among the works it cites.