Fetching the paper…
Reading the bibliography…
Different from Composed Image Retrieval task that requires expensive labels for training task-specific models, Zero-Shot Composed Image Retrieval (ZS-CIR) involves diverse tasks with a broad range of visual content manipulation intent that could be related to domain, scene, object, and attribute.
Language Models are Few-Shot Learners
Brown, T.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J. D.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; Agarwal, S.; Herbert-Voss, A.; Krueger, G.; Henighan, T.; Child, R.; Ramesh, A.; Ziegler, D.; Wu, J.; Winter, C.; Hesse, C.; Chen, M.; Sigler, E.; Litwin, M.; Gray, S.; Chess, B.; Clark, J.; Berner, C.; McCandlish, S.; Radford, A.; Sutskever, I.; and Amodei, D. 2020 · 1901
Earlier work this paper cites.
Multi-Concept Customization of Text-to-Image Diffusion
Kumari, N.; Zhang, B.; Zhang, R.; Shechtman, E.; and Zhu, J.-Y. 2023 · 1941
Earlier work this paper cites.
Long short-term memory
Hochreiter, S.; and Schmidhuber, J. 1997 · 1997
Earlier work this paper cites.
Modality-agnostic attention fusion for visual search with text feedback
Dodds, E.; Culpepper, J.; Herdade, S.; Zhang, Y.; and Boakye, K. 2020 · 2007
Earlier work this paper cites.
Image retrieval: Ideas, influences, and trends of the new age
Datta, R.; Joshi, D.; Li, J.; and Wang, J. Z. 2008 · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J.; Dong, W.; Socher, R.; Li, L.-J.; Li, K.; and Fei-Fei, L. 2009 · 2009
Earlier work this paper cites.
Microsoft COCO: Common Objects in Context
Lin, T.-Y.; Maire, M.; Belongie, S.; Hays, J.; Perona, P.; Ramanan, D.; Dollár, P.; and Zitnick, C. L. 2014 · 2014
Earlier work this paper cites.
Decoupled Weight Decay Regularization
Loshchilov, I.; and Hutter, F. 2018 · 2018
Earlier work this paper cites.
Conceptual Captions: A Cleaned, Hypernymed, Image Alt-text Dataset For Automatic Image Captioning
Sharma, P.; Ding, N.; Goodman, S.; and Soricut, R. 2018 · 2018
Earlier work this paper cites.
End-to-end object detection with transformers
Carion, N.; Massa, F.; Synnaeve, G.; Usunier, N.; Kirillov, A.; and Zagoruyko, S. 2020 · 2020
Earlier work this paper cites.
Image Search With Text Feedback by Visiolinguistic Attention Learning
Chen, Y.; Gong, S.; and Bazzani, L. 2020 · 2020
Earlier work this paper cites.
spaCy: Industrial-strength natural language processing in python
Honnibal, M.; Montani, I.; Van Landeghem, S.; Boyd, A.; et al. 2020 · 2020
Earlier work this paper cites.
Oscar: Object-semantics aligned pre-training for vision-language tasks
Li, X.; Yin, X.; Li, C.; Zhang, P.; Hu, X.; Zhang, L.; Wang, L.; Hu, H.; Dong, L.; Wei, F.; et al. 2020 · 2020
Cited alongside, same era.
Rezero is all you need: Fast convergence at large depth
Bachlechner, T.; Majumder, B. P.; Mao, H.; Cottrell, G.; and McAuley, J. 2021 · 2021
Cited alongside, same era.
The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution Generalization
Hendrycks, D.; Basart, S.; Mu, N.; Kadavath, S.; Wang, F.; Dorundo, E.; Desai, R.; Zhu, T.; Parajuli, S.; Guo, M.; Song, D.; Steinhardt, J.; and Gilmer, J. 2021 · 2021
Cited alongside, same era.
Align before Fuse: Vision and Language Representation Learning with Momentum Distillation
Li, J.; Selvaraju, R.; Gotmare, A.; Joty, S.; Xiong, C.; and Hoi, S. C. H. 2021 · 2021
Cited alongside, same era.
Image Retrieval on Real-Life Images With Pre-Trained Vision-and-Language Models
Liu, Z.; Rodriguez-Opazo, C.; Teney, D.; and Gould, S. 2021 · 2021
Cited alongside, same era.
ARTEMIS: Attention-based Retrieval with Text-Explicit Matching and Implicit Similarity
Delmas, G.; Rezende, R. S.; Csurka, G.; and Larlus, D. 2022 · 2022
Later among the works it cites.
FashionVLP: Vision Language Transformer for Fashion Retrieval With Feedback
Goenka, S.; Zheng, Z.; Jaiswal, A.; Chada, R.; Wu, Y.; Hedau, V.; and Natarajan, P. 2022 · 2022
Later among the works it cites.
FashionViL: Fashion-Focused Vision-and-Language Representation Learning
Han, X.; Yu, L.; Zhu, X.; Zhang, L.; Song, Y.-Z.; and Xiang, T. 2022 · 2022
Later among the works it cites.
BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Li, J.; Li, D.; Xiong, C.; and Hoi, S. 2022 · 2022
Later among the works it cites.
CLIP Models are Few-shot Learners: Empirical Studies on VQA and Visual Entailment
Song, H.; Dong, L.; Zhang, W.-N.; Liu, T.; and Wei, F. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mokady, R.; Hertz, A.; and Bermano, A. H. 2021 · 2021
Cited alongside, same era.
Learning Transferable Visual Models From Natural Language Supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; Krueger, G.; and Sutskever, I. 2021 · 2021
Cited alongside, same era.
Fashion IQ: A New Dataset Towards Retrieving Images by Natural Language Feedback
Wu, H.; Gao, Y.; Guo, X.; Al-Halah, Z.; Rennie, S.; Grauman, K.; and Feris, R. 2021 · 2021
Cited alongside, same era.
Vinvl: Revisiting visual representations in vision-language models
Zhang, P.; Li, X.; Hu, X.; Yang, J.; Zhang, L.; Wang, L.; Choi, Y.; and Gao, J. 2021 · 2021
Cited alongside, same era.
Flamingo: a Visual Language Model for Few-Shot Learning
Alayrac, J.-B.; Donahue, J.; Luc, P.; Miech, A.; Barr, I.; Hasson, Y.; Lenc, K.; Mensch, A.; Millican, K.; Reynolds, M.; Ring, R.; Rutherford, E.; Cabi, S.; Han, T.; Gong, Z.; Samangooei, S.; Monteiro, M.; Menick, J. L.; Borgeaud, S.; Brock, A.; Nematzadeh, A.; Sharifzadeh, S.; Bińkowski, M. a.; Barreira, R.; Vinyals, O.; Zisserman, A.; and Simonyan, K. 2022 · 2022
Cited alongside, same era.
Effective Conditioned and Composed Image Retrieval Combining CLIP-Based Features
Baldrati, A.; Bertini, M.; Uricchio, T.; and Del Bimbo, A. 2022 · 2022
Cited alongside, same era.
“This Is My Unicorn, Fluffy”: Personalizing Frozen Vision-Language Representations
Cohen, N.; Gal, R.; Meirom, E. A.; Chechik, G.; and Atzmon, Y. 2022 · 2022
Cited alongside, same era.
Conditional Prompt Learning for Vision-Language Models
Zhou, K.; Yang, J.; Loy, C. C.; and Liu, Z. 2022 · 2022
Later among the works it cites.
Zero-Shot Composed Image Retrieval with Textual Inversion
Baldrati, A.; Agnolucci, L.; Bertini, M.; and Del Bimbo, A. 2023 · 2023
Closest in time.
Li, J.; Li, D.; Savarese, S.; and Hoi, S. 2023 · 2023
Closest in time.
Pic2Word: Mapping Pictures to Words for Zero-Shot Composed Image Retrieval
Saito, K.; Sohn, K.; Zhang, X.; Li, C.-L.; Lee, C.-Y.; Saenko, K.; and Pfister, T. 2023 · 2023
Closest in time.
Dual Pseudo-Labels Interactive Self-Training for Semi-Supervised Visible-Infrared Person Re-Identification
Shi, J.; Zhang, Y.; Yin, X.; Xie, Y.; Zhang, Z.; Fan, J.; Shi, Z.; and Qu, Y. 2023 · 2023
Closest in time.
Simple Weakly-Supervised Image Captioning via CLIP’s Multimodal Embeddings
Tam, D.; Raffel, C.; and Bansal, M. 2023 · 2023
Closest in time.
Visualize Before You Write: Imagination-Guided Open-Ended Text Generation
Zhu, W.; Yan, A.; Lu, Y.; Xu, W.; Wang, X. E.; Eckstein, M.; and Wang, W. Y. 2023 · 2023
Closest in time.