Fetching the paper…
Reading the bibliography…
Given a query consisting of a reference image and a relative caption, Composed Image Retrieval (CIR) aims to retrieve target images visually similar to the reference one while incorporating the changes specified in the relative caption.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al. , “Language models are few-shot learners,” in Proc. of Advances in Neural Information Processing Systems (NeurIPS) , vol. 33, 2020, pp. 1877–1901
1901
Earlier work this paper cites.
J. MacQueen et al. , “Some methods for classification and analysis of multivariate observations,” in Proceedings of the fifth Berkeley symposium on mathematical statistics and probability , vol. 1, no. 14. Oakland, CA, USA, 1967, pp. 281–297
1967
Earlier work this paper cites.
T. L. Berg, A. C. Berg, and J. Shih, “Automatic attribute discovery and characterization from noisy web data,” in Proc. of the European Conference on Computer Vision (ECCV) . Springer, 2010, pp. 663–676
2010
Earlier work this paper cites.
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick, “Microsoft coco: Common objects in context,” in Proc. of the European Conference on Computer Vision (ECCV) . Springer, 2014, pp. 740–755
2014
Earlier work this paper cites.
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh, “Vqa: Visual question answering,” in Proc. of the IEEE/CVF International Conference on Computer Vision (ICCV) , 2015, pp. 2425–2433
2015
Earlier work this paper cites.
G. Hinton, O. Vinyals, and J. Dean, “Distilling the knowledge in a neural network,” in NIPS Deep Learning and Representation Learning Workshop , 2015
2015
Earlier work this paper cites.
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein et al. , “Imagenet large scale visual recognition challenge,” International Journal of Computer Vision (IJCV) , vol. 115, pp. 211–252, 2015
2015
Earlier work this paper cites.
G. Van Horn, S. Branson, R. Farrell, S. Haber, J. Barry, P. Ipeirotis, P. Perona, and S. Belongie, “Building a bird recognition app and large scale dataset with citizen scientists: The fine print in fine-grained dataset collection,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2015, pp. 595–604
2015
Earlier work this paper cites.
2016
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 770–778
2016
Earlier work this paper cites.
X. Han, Z. Wu, P. X. Huang, X. Zhang, M. Zhu, Y. Li, Y. Zhao, and L. S. Davis, “Automatic spatially-aware fashion concept discovery,” in Proc. of the IEEE/CVF International Conference on Computer Vision (ICCV) , 2017, pp. 1463–1471
2017
Earlier work this paper cites.
G. Chen, W. Choi, X. Yu, T. Han, and M. Chandraker, “Learning efficient object detection models with knowledge distillation,” in Proc. of Advances in Neural Information Processing Systems (NeurIPS) , vol. 30, 2017
2017
Earlier work this paper cites.
X. Guo, H. Wu, Y. Cheng, S. Rennie, G. Tesauro, and R. Feris, “Dialog-based interactive image retrieval,” in Proc. of Advances in Neural Information Processing Systems (NeurIPS) , vol. 31, 2018
2018
Earlier work this paper cites.
P. Sharma, N. Ding, S. Goodman, and R. Soricut, “Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning,” in Proc. of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , 2018, pp. 2556–2565
2018
Earlier work this paper cites.
I. Loshchilov and F. Hutter, “Decoupled weight decay regularization,” in International Conference on Learning Representations , 2018
2018
Earlier work this paper cites.
N. Vo, L. Jiang, C. Sun, K. Murphy, L.-J. Li, L. Fei-Fei, and J. Hays, “Composing text and image for image retrieval-an empirical odyssey,” in Proc. of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2019, pp. 6439–6448
2019
Earlier work this paper cites.
M. Forbes, C. Kaeser-Chen, P. Sharma, and S. Belongie, “Neural naturalist: Generating fine-grained image comparisons,” in Proc. of the Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) , 2019, pp. 708–717
2019
Earlier work this paper cites.
T. Karras, S. Laine, and T. Aila, “A style-based generator architecture for generative adversarial networks,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2019, pp. 4401–4410
2019
Earlier work this paper cites.
M. Cornia, M. Stefanini, L. Baraldi, and R. Cucchiara, “Meshed-memory transformer for image captioning,” in Proc. of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2020, pp. 10 578–10 587
2020
Earlier work this paper cites.
J. Zhu, Y. Shen, D. Zhao, and B. Zhou, “In-domain gan inversion for real image editing,” in Proc. of the European Conference on Computer Vision (ECCV) . Springer, 2020, pp. 592–608
2020
Earlier work this paper cites.
A. Kuznetsova, H. Rom, N. Alldrin, J. Uijlings, I. Krasin, J. Pont-Tuset, S. Kamali, S. Popov, M. Malloci, A. Kolesnikov et al. , “The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale,” International Journal of Computer Vision (IJCV) , vol. 128, no. 7, pp. 1956–1981, 2020
2020
Earlier work this paper cites.
T. Chen, S. Kornblith, M. Norouzi, and G. Hinton, “A simple framework for contrastive learning of visual representations,” in Proc. of International Conference on Machine Learning (ICML) . PMLR, 2020, pp. 1597–1607
2020
Earlier work this paper cites.
J. D. Robinson, C.-Y. Chuang, S. Sra, and S. Jegelka, “Contrastive learning with hard negative samples,” in International Conference on Learning Representations , 2020
2020
Earlier work this paper cites.
Y. Kalantidis, M. B. Sariyildiz, N. Pion, P. Weinzaepfel, and D. Larlus, “Hard negative mixing for contrastive learning,” Advances in Neural Information Processing Systems , vol. 33, pp. 21 798–21 809, 2020
2020
Cited alongside, same era.
Z. Liu, C. Rodriguez-Opazo, D. Teney, and S. Gould, “Image retrieval on real-life images with pre-trained vision-and-language models,” in Proc. of the IEEE/CVF International Conference on Computer Vision (ICCV) , 2021, pp. 2125–2134
2021
Cited alongside, same era.
S. Lee, D. Kim, and B. Han, “Cosmo: Content-style modulation for image retrieval with text feedback,” in Proc. of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2021, pp. 802–812
2021
Cited alongside, same era.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al. , “Learning transferable visual models from natural language supervision,” in Proc. of International Conference on Machine Learning (ICML) . PMLR, 2021, pp. 8748–8763
A. Baldrati, L. Agnolucci, M. Bertini, and A. Del Bimbo, “Zero-shot composed image retrieval with textual inversion,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , October 2023, pp. 15 338–15 347
2023
Later among the works it cites.
R. Gal, Y. Alaluf, Y. Atzmon, O. Patashnik, A. H. Bermano, G. Chechik, and D. Cohen-Or, “An image is worth one word: Personalizing text-to-image generation using textual inversion,” in Proc. of International Conference on Learning Representations (ICLR) , 2023
2023
Later among the works it cites.
K. Saito, K. Sohn, X. Zhang, C.-L. Li, C.-Y. Lee, K. Saenko, and T. Pfister, “Pic2word: Mapping pictures to words for zero-shot composed image retrieval,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 19 305–19 314
2023
Later among the works it cites.
Z. Shao, Z. Yu, M. Wang, and J. Yu, “Prompting large language models with answer heuristics for knowledge-based visual question answering,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 14 974–14 983
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2021
Cited alongside, same era.
H. Wu, Y. Gao, X. Guo, Z. Al-Halah, S. Rennie, K. Grauman, and R. Feris, “Fashion iq: A new dataset towards retrieving images by natural language feedback,” in Proc. of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2021, pp. 11 307–11 317
2021
Cited alongside, same era.
A. Chawla, H. Yin, P. Molchanov, and J. Alvarez, “Data-free knowledge distillation for object detection,” in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision , 2021, pp. 3289–3298
2021
Cited alongside, same era.
O. Tov, Y. Alaluf, Y. Nitzan, O. Patashnik, and D. Cohen-Or, “Designing an encoder for stylegan image manipulation,” ACM Transactions on Graphics (TOG) , vol. 40, no. 4, pp. 1–14, 2021
2021
Cited alongside, same era.
D. Hendrycks, S. Basart, N. Mu, S. Kadavath, F. Wang, E. Dorundo, R. Desai, T. Zhu, S. Parajuli, M. Guo et al. , “The many faces of robustness: A critical analysis of out-of-distribution generalization,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 8340–8349
2021
Cited alongside, same era.
Q. Cao, L. Shen, W. Xie, O. M. Parkhi, and A. Zisserman, “Vggface2: A dataset for recognising faces across pose and age,” in 2018 13th IEEE international conference on automatic face & gesture recognition (FG 2018) . IEEE, 2018, pp. 67–74
2021
Cited alongside, same era.
A. Baldrati, M. Bertini, T. Uricchio, and A. Del Bimbo, “Conditioned and composed image retrieval combining and partially fine-tuning clip-based features,” in Proc. of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2022, pp. 4959–4968
2022
Cited alongside, same era.
——, “Effective conditioned and composed image retrieval combining CLIP-based features,” in Proc. of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2022, pp. 21 466–21 474
2022
Cited alongside, same era.
G. Delmas, R. S. Rezende, G. Csurka, and D. Larlus, “ARTEMIS: Attention-based retrieval with text-explicit matching and implicit similarity,” in Proc. of International Conference on Learning Representations (ICLR) , 2022
2022
Cited alongside, same era.
2023
Later among the works it cites.
M. Barraco, S. Sarto, M. Cornia, L. Baraldi, and R. Cucchiara, “With a little help from your own past: Prototypical memory networks for image captioning,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 3021–3031
2023
Later among the works it cites.
C. Meng, R. Rombach, R. Gao, D. Kingma, S. Ermon, J. Ho, and T. Salimans, “On distillation of guided diffusion models,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 14 297–14 306
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
N. Ruiz, Y. Li, V. Jampani, Y. Pritch, M. Rubinstein, and K. Aberman, “Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 22 500–22 510
2023
Later among the works it cites.
N. Kumari, B. Zhang, R. Zhang, E. Shechtman, and J.-Y. Zhu, “Multi-concept customization of text-to-image diffusion,” in Proc. of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
S. Gu, C. Clark, and A. Kembhavi, “I can’t believe there’s no images! learning visual tasks using only language supervision,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 2672–2683
2023
Later among the works it cites.
R. Gal, M. Arar, Y. Atzmon, A. H. Bermano, G. Chechik, and D. Cohen-Or, “Encoder-based domain tuning for fast personalization of text-to-image models,” ACM Transactions on Graphics (TOG) , vol. 42, no. 4, pp. 1–13, 2023
2023
Later among the works it cites.
W. Li, H. Fan, Y. Wong, M. Kankanhalli, and Y. Yang, “CAT-LLM: Context-aware training enhanced large language models for multi-modal contextual image retrieval,” 2024
2024
Closest in time.
S. Karthik, K. Roth, M. Mancini, and Z. Akata, “Vision-by-language for training-free compositional image retrieval,” in The Twelfth International Conference on Learning Representations , 2024
2024
Closest in time.
M. Mistretta, A. Baldrati, M. Bertini, and A. D. Bagdanov, “Improving zero-shot generalization of learned prompts via unsupervised knowledge distillation,” in European Conference on Computer Vision . Springer, 2024, pp. 459–477
2024
Closest in time.
M. Mistretta, A. Baldrati, L. Agnolucci, M. Bertini, and A. D. Bagdanov, “Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion,” in The Thirteenth International Conference on Learning Representations , 2025
2025
Closest in time.