Fetching the paper…
Reading the bibliography…
In the real world, where information is abundant and diverse across different modalities, understanding and utilizing various data types to improve retrieval systems is a key focus of research.
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al. , “Language models are few-shot learners,” in NeurIPS , 2020, pp. 1877–1901
1901
Earlier work this paper cites.
N. Tumanyan, M. Geyer, S. Bagon, and T. Dekel, “Plug-and-play diffusion features for text-driven image-to-image translation,” in CVPR , 2023, pp. 1921–1930
1930
Earlier work this paper cites.
W. Adams, G. Iyengar, C.-Y. Lin, M. R. Naphade, C. Neti, H. J. Nock, and J. R. Smith, “Semantic indexing of multimedia content using visual, audio, and text cues,” EURASIP Journal on Advances in Signal Processing , vol. 2003, pp. 1–16, 2003
2003
Earlier work this paper cites.
M.-E. Nilsback and A. Zisserman, “Automated flower classification over a large number of classes,” in 2008 Sixth Indian Conference on Computer Vision, Graphics & Image Processing , 2008, pp. 722–729
2008
Earlier work this paper cites.
T. L. Berg, A. C. Berg, and J. Shih, “Automatic attribute discovery and characterization from noisy web data,” in ECCV , 2010, pp. 663–676
2010
Earlier work this paper cites.
V. Bychkovsky, S. Paris, E. Chan, and F. Durand, “Learning photographic global tonal adjustment with a database of input/output image pairs,” in CVPR , 2011, pp. 97–104
2011
Earlier work this paper cites.
C. Wah, S. Branson, P. Welinder, P. Perona, and S. Belongie, “The caltech-ucsd birds-200-2011 dataset,” 2011
2011
Earlier work this paper cites.
A. Graves and A. Graves, “Long short-term memory,” Supervised sequence labelling with recurrent neural networks , pp. 37–45, 2012
2012
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” NIPS , vol. 25, 2012
2012
Earlier work this paper cites.
Z. Si and S.-C. Zhu, “Learning hybrid image templates (hit) by information projection,” IEEE TPAMI , vol. 34, no. 7, pp. 1354–1367, 2012
2012
Earlier work this paper cites.
E. Hassan, S. Chaudhury, and M. Gopal, “Multi-modal information integration for document retrieval,” in DAR , 2013, pp. 1200–1204
2013
Earlier work this paper cites.
T. Mikolov, K. Chen, G. Corrado, and J. Dean, “Efficient estimation of word representations in vector space,” arXiv , 2013
2013
Earlier work this paper cites.
A. Mourão and F. Martins, “Novamedsearch: a multimodal search engine for medical case-based retrieval,” in Proceedings of the 10th Conference on Open Research Areas in Information Retrieval , ser. OAIR ’13. LE CENTRE DE HAUTES ETUDES INTERNATIONALES D’INFORMATIQUE DOCUMENTAIRE, 2013, p. 223–224
2013
Earlier work this paper cites.
A. Babenko, A. Slesarev, A. Chigorin, and V. Lempitsky, “Neural codes for image retrieval,” in ECCV , 2014, pp. 584–599
2014
Earlier work this paper cites.
Y. Cao, S. Steffey, J. He, D. Xiao, C. Tao, P. Chen, and H. Müller, “Medical image retrieval: a multimodal approach,” Cancer informatics , vol. 13, pp. CIN–S14 053, 2014
2014
Earlier work this paper cites.
K. Cho, B. Van Merrienboer, C. Gulcehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y. Bengio, “Learning phrase representations using rnn encoder–decoder for statistical machine translation,” in EMNLP , 2014, pp. 1724–1734
2014
Earlier work this paper cites.
S. Kazemzadeh, V. Ordonez, M. Matten, and T. Berg, “ReferItGame: Referring to objects in photographs of natural scenes,” in EMNLP , A. Moschitti, B. Pang, and W. Daelemans, Eds., Oct. 2014, pp. 787–798
2014
Earlier work this paper cites.
R. Kiros, R. Salakhutdinov, and R. S. Zemel, “Unifying visual-semantic embeddings with multimodal neural language models,” arXiv , 2014
2014
Earlier work this paper cites.
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick, “Microsoft coco: Common objects in context,” in ECCV , 2014, pp. 740–755
2014
Earlier work this paper cites.
M. Mirza and S. Osindero, “Conditional generative adversarial nets,” arXiv , 2014
2014
Earlier work this paper cites.
J. Pennington, R. Socher, and C. D. Manning, “Glove: Global vectors for word representation,” in EMNLP , 2014, pp. 1532–1543
2014
Earlier work this paper cites.
J. Wang, Y. Song, T. Leung, C. Rosenberg, J. Wang, J. Philbin, B. Chen, and Y. Wu, “Learning fine-grained image similarity with deep ranking,” in CVPR , 2014, pp. 1386–1393
2014
Earlier work this paper cites.
P. Isola, J. J. Lim, and E. H. Adelson, “Discovering states and transformations in image collections,” in CVPR , 2015, pp. 1383–1391
2015
Earlier work this paper cites.
Z. Liu, P. Luo, X. Wang, and X. Tang, “Deep learning face attributes in the wild,” in ICCV , 2015, pp. 3730–3738
2015
Earlier work this paper cites.
B. A. Plummer, L. Wang, C. M. Cervantes, J. C. Caicedo, J. Hockenmaier, and S. Lazebnik, “Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models,” in ICCV , 2015, pp. 2641–2649
2015
Earlier work this paper cites.
A. Radford, L. Metz, and S. Chintala, “Unsupervised representation learning with deep convolutional generative adversarial networks,” arXiv , 2015
2015
Earlier work this paper cites.
S. Ren, K. He, R. Girshick, and J. Sun, “Faster r-cnn: Towards real-time object detection with region proposal networks,” NIPS , vol. 28, 2015
2015
Earlier work this paper cites.
K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,” in ICLR , 2015
2015
Earlier work this paper cites.
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich, “Going deeper with convolutions,” in CVPR , 2015, pp. 1–9
2015
Earlier work this paper cites.
F. Yu, A. Seff, Y. Zhang, S. Song, T. Funkhouser, and J. Xiao, “Lsun: Construction of a large-scale image dataset using deep learning with humans in the loop,” arXiv , 2015
2015
Earlier work this paper cites.
M. Cordts, M. Omran, S. Ramos, T. Rehfeld, M. Enzweiler, R. Benenson, U. Franke, S. Roth, and B. Schiele, “The cityscapes dataset for semantic urban scene understanding,” in CVPR , 2016, pp. 3213–3223
2016
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in CVPR , 2016, pp. 770–778
2016
Earlier work this paper cites.
M. Plappert, C. Mandery, and T. Asfour, “The kit motion-language dataset,” Big data , vol. 4, no. 4, pp. 236–252, 2016
2016
Earlier work this paper cites.
P. Sangkloy, N. Burnell, C. Ham, and J. Hays, “The sketchy database: learning to retrieve badly drawn bunnies,” vol. 35, no. 4, jul 2016
2016
Earlier work this paper cites.
H. Dong, S. Yu, C. Wu, and Y. Guo, “Semantic image synthesis via adversarial learning,” in ICCV , 2017, pp. 5706–5714
2017
Earlier work this paper cites.
J. F. Gemmeke, D. P. W. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter, “Audio set: An ontology and human-labeled dataset for audio events,” in ICASSP , 2017, pp. 776–780
2017
Earlier work this paper cites.
T. Han and D. Schlangen, “Draw and tell: Multimodal descriptions outperform verbal-or sketch-only descriptions in an image retrieval task,” in IJNLP , 2017, pp. 361–365
2017
Earlier work this paper cites.
X. Han, Z. Wu, P. X. Huang, X. Zhang, M. Zhu, Y. Li, Y. Zhao, and L. S. Davis, “Automatic spatially-aware fashion concept discovery,” in ICCV , 2017, pp. 1463–1471
2017
Earlier work this paper cites.
S. Hershey, S. Chaudhuri, D. P. Ellis, J. F. Gemmeke, A. Jansen, R. C. Moore, M. Plakal, D. Platt, R. A. Saurous, B. Seybold et al. , “Cnn architectures for large-scale audio classification,” in icassp , 2017, pp. 131–135
2017
Earlier work this paper cites.
R. Hinami, Y. Matsui, and S. Satoh, “Region-based image retrieval revisited,” in ACM MM , 2017, pp. 528–536
2017
Earlier work this paper cites.
A. G. Howard, M. Zhu, B. Chen, D. Kalenichenko, W. Wang, T. Weyand, M. Andreetto, and H. Adam, “Mobilenets: Efficient convolutional neural networks for mobile vision applications,” arXiv , 2017
2017
Earlier work this paper cites.
G. Huang, Z. Liu, L. Van Der Maaten, and K. Q. Weinberger, “Densely connected convolutional networks,” in CVPR , 2017, pp. 4700–4708
2017
Earlier work this paper cites.
J. Johnson, B. Hariharan, L. Van Der Maaten, L. Fei-Fei, C. Lawrence Zitnick, and R. Girshick, “Clevr: A diagnostic dataset for compositional language and elementary visual reasoning,” in CVPR , 2017, pp. 2901–2910
2017
Earlier work this paper cites.
——, “Imagenet classification with deep convolutional neural networks,” Communications of the ACM , vol. 60, no. 6, pp. 84–90, 2017
2017
Earlier work this paper cites.
L. Mai, H. Jin, Z. Lin, C. Fang, J. Brandt, and F. Liu, “Spatial-semantic image search by visual feature synthesis,” in CVPR , 2017, pp. 1121–1130
2017
Earlier work this paper cites.
A. Miech, I. Laptev, and J. Sivic, “Learnable pooling with context gating for video classification,” arXiv , 2017
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” NIPS , vol. 30, 2017
2017
Earlier work this paper cites.
S. Zhu, R. Urtasun, S. Fidler, D. Lin, and C. Change Loy, “Be your own prada: Fashion synthesis with structural coherence,” in ICCV , 2017, pp. 1680–1688
2017
Earlier work this paper cites.
J. Chen, Y. Shen, J. Gao, J. Liu, and X. Liu, “Language-based image editing with recurrent attentive models,” in CVPR , 2018, pp. 8721–8729
2018
Earlier work this paper cites.
L.-C. Chen, Y. Zhu, G. Papandreou, F. Schroff, and H. Adam, “Encoder-decoder with atrous separable convolution for semantic image segmentation,” in ECCV , 2018, pp. 801–818
2018
Earlier work this paper cites.
Y. Chen, Y.-K. Lai, and Y.-J. Liu, “Cartoongan: Generative adversarial networks for photo cartoonization,” in CVPR , 2018, pp. 9465–9474
2018
Earlier work this paper cites.
A. Gonzalez-Garcia, J. v. d. Weijer, and Y. Bengio, “Image-to-image translation for cross-domain disentanglement,” in NIPS , 2018, pp. 1294–1305
2018
Earlier work this paper cites.
X. Guo, H. Wu, Y. Cheng, S. Rennie, G. Tesauro, and R. Feris, “Dialog-based interactive image retrieval,” NIPS , vol. 31, 2018
2018
Earlier work this paper cites.
X. Huang, M.-Y. Liu, S. Belongie, and J. Kautz, “Multimodal unsupervised image-to-image translation,” in ECCV , 2018, pp. 172–189
2018
Earlier work this paper cites.
T. Karras, T. Aila, S. Laine, and J. Lehtinen, “Progressive growing of gans for improved quality, stability, and variation,” in ICLR , 2018
2018
Earlier work this paper cites.
G. Mao, Y. Yuan, and L. Xiaoqiang, “Deep cross-modal retrieval for remote sensing image and audio,” in IAPR workshop , 2018, pp. 1–7
2018
Earlier work this paper cites.
S. Nam, Y. Kim, and S. J. Kim, “Text-adaptive generative adversarial networks: Manipulating images with natural language,” NIPS , vol. 31, 2018
2018
Earlier work this paper cites.
P. Sharma, N. Ding, S. Goodman, and R. Soricut, “Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning,” in ACL , 2018, pp. 2556–2565
2018
Earlier work this paper cites.
——, “Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning,” in ACL , 2018, pp. 2556–2565
2018
Earlier work this paper cites.
S. Shinagawa, K. Yoshino, S. Sakti, Y. Suzuki, and S. Nakamura, “Interactive image manipulation with natural language instruction commands,” arXiv , 2018
2018
Earlier work this paper cites.
H. Wang, J. D. Williams, and S. Kang, “Learning to globally edit images with textual description,” arXiv , 2018
2018
Earlier work this paper cites.
K. Chen, C. B. Choy, M. Savva, A. X. Chang, T. Funkhouser, and S. Savarese, “Text2shape: Generating shapes from natural language by learning joint embeddings,” in ACCV , 2019, pp. 100–116
2019
Earlier work this paper cites.
A. El-Nouby, S. Sharma, H. Schulz, D. Hjelm, L. E. Asri, S. E. Kahou, Y. Bengio, and G. W. Taylor, “Tell, draw, and repeat: Generating and modifying images based on continual linguistic instruction,” in ICCV , 2019, pp. 10 304–10 312
2019
Earlier work this paper cites.
F. Feng, X. He, J. Tang, and T.-S. Chua, “Graph adversarial training: Dynamically regularizing based on graph structure,” IEEE TKDE , vol. 33, no. 6, pp. 2493–2504, 2019
2019
Earlier work this paper cites.
M. Forbes, C. Kaeser-Chen, P. Sharma, and S. Belongie, “Neural naturalist: Generating fine-grained image comparisons,” in EMNLP-IJCNLP , 2019, pp. 708–717
2019
Earlier work this paper cites.
R. Furuta, N. Inoue, and T. Yamasaki, “Efficient and interactive spatial-semantic image retrieval,” Multimedia Tools and Applications , vol. 78, pp. 18 713–18 733, 2019
2019
Earlier work this paper cites.
T. Karras, S. Laine, and T. Aila, “A style-based generator architecture for generative adversarial networks,” in CVPR , 2019, pp. 4401–4410
2019
Earlier work this paper cites.
J. D. M.-W. C. Kenton and L. K. Toutanova, “Bert: Pre-training of deep bidirectional transformers for language understanding,” in Proceedings of NAACL-HLT , 2019, pp. 4171–4186
2019
Earlier work this paper cites.
J.-H. Kim, N. Kitaev, X. Chen, M. Rohrbach, B.-T. Zhang, Y. Tian, D. Batra, and D. Parikh, “Codraw: Collaborative drawing as a testbed for grounded goal-driven communication,” in ACL , 2019, pp. 6495–6513
2019
Earlier work this paper cites.
B. Li, X. Qi, T. Lukasiewicz, and P. Torr, “Controllable text-to-image generation,” NIPS , vol. 32, 2019
2019
Earlier work this paper cites.
X. Mao, Y. Chen, Y. Li, T. Xiong, Y. He, and H. Xue, “Bilinear representation for language-based image editing using conditional generative adversarial networks,” in ICASSP . IEEE, May 2019
2019
Earlier work this paper cites.
Y. Peng and J. Qi, “Cm-gans: Cross-modal generative adversarial networks for common representation learning,” ACM TOMM , vol. 15, no. 1, pp. 1–24, 2019
2019
Earlier work this paper cites.
L. Rossetto, R. Gasser, and H. Schuldt, “Query by semantic sketch,” arXiv , 2019
2019
Earlier work this paper cites.
V. Sanh, “Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter.” in NIPS , 2019
2019
Earlier work this paper cites.
M. Tan and Q. Le, “Efficientnet: Rethinking model scaling for convolutional neural networks,” in ICML , 2019, pp. 6105–6114
2019
Earlier work this paper cites.
I. Tautkute, T. Trzciński, A. P. Skorupa, Ł. Brocki, and K. Marasek, “Deepstyle: Multimodal search engine for fashion and interior design,” IEEE Access , vol. 7, pp. 84 613–84 628, 2019
2019
Earlier work this paper cites.
N. Vo, L. Jiang, C. Sun, K. Murphy, L.-J. Li, L. Fei-Fei, and J. Hays, “Composing text and image for image retrieval-an empirical odyssey,” in CVPR , 2019, pp. 6439–6448
2019
Earlier work this paper cites.
J. Wang, S. Zhu, J. Xu, and D. Cao, “The retrieval of the beautiful: Self-supervised salient object detection for beauty product retrieval,” in ACM MM , 2019, p. 2548–2552
2019
Earlier work this paper cites.
Y. Wei, X. Wang, L. Nie, X. He, R. Hong, and T.-S. Chua, “Mmgcn: Multi-modal graph convolution network for personalized recommendation of micro-video,” in ACM MM , 2019, pp. 1437–1445
2019
Earlier work this paper cites.
L. Zhang, J. Liu, Y. Yang, F. Huang, F. Nie, and D. Zhang, “Optimal projection guided transfer hashing for image retrieval,” IEEE TCSVT , vol. 30, no. 10, pp. 3788–3802, 2019
2019
Earlier work this paper cites.
L. Zhang, G.-J. Qi, L. Wang, and J. Luo, “Aet vs. aed: Unsupervised representation learning by auto-encoding transformations rather than data,” in CVPR , 2019, pp. 2547–2555
2019
Earlier work this paper cites.
L. Zhen, P. Hu, X. Wang, and D. Peng, “Deep supervised cross-modal retrieval,” in CVPR , 2019, pp. 10 394–10 403
2019
Cited alongside, same era.
Y. Chen and L. Bazzani, “Learning joint visual semantic matching embeddings for language-guided retrieval,” in ECCV , 2020, pp. 136–152
2020
Cited alongside, same era.
Y. Chen, S. Gong, and L. Bazzani, “Image search with text feedback by visiolinguistic attention learning,” in CVPR , 2020, pp. 3001–3011
2020
Cited alongside, same era.
Y. Cheng, Z. Gan, Y. Li, J. Liu, and J. Gao, “Sequential attention gan for interactive image editing,” in ACM MM , 2020, pp. 4383–4391
2020
Cited alongside, same era.
Y. Choi, Y. Uh, J. Yoo, and J.-W. Ha, “Stargan v2: Diverse image synthesis for multiple domains,” in CVPR , 2020, pp. 8188–8197
2020
Cited alongside, same era.
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, “High-resolution image synthesis with latent diffusion models,” in CVPR , 2022, pp. 10 684–10 695
2022
Later among the works it cites.
P. Sangkloy, W. Jitkrittum, D. Yang, and J. Hays, “A sketch is worth a thousand words: Image retrieval with text and sketch,” in ECCV , 2022, pp. 251–267
2022
Later among the works it cites.
G. Tevet, B. Gordon, A. Hertz, A. H. Bermano, and D. Cohen-Or, “Motionclip: Exposing human motion generation to clip space,” in ECCV , 2022, pp. 358–374
2022
Later among the works it cites.
Y. Tian, S. Newsam, and K. Boakye, “Image search with text feedback by additive attention compositional learning,” arXiv , 2022
2022
Later among the works it cites.
N. Tumanyan, O. Bar-Tal, S. Bagon, and T. Dekel, “Splicing vit features for semantic appearance transfer,” in CVPR , 2022, pp. 10 748–10 757
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
E. Dodds, J. Culpepper, S. Herdade, Y. Zhang, and K. Boakye, “Modality-agnostic attention fusion for visual search with text feedback,” arXiv , 2020
2020
Cited alongside, same era.
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly et al. , “An image is worth 16x16 words: Transformers for image recognition at scale,” in ICLR , 2020
2020
Cited alongside, same era.
C. Gao, Q. Liu, Q. Xu, L. Wang, J. Liu, and C. Zou, “Sketchycoco: Image generation from freehand scene sketches,” in CVPR , 2020, pp. 5174–5183
2020
Cited alongside, same era.
Y. Gu, K. Vyas, M. Shen, J. Yang, and G.-Z. Yang, “Deep graph-based multimodal feature embedding for endomicroscopy image retrieval,” IEEE TNNLS , vol. 32, no. 2, pp. 481–492, 2020
2020
Cited alongside, same era.
M. Hosseinzadeh and Y. Wang, “Composed query image retrieval using locally bounded features,” in CVPR , 2020, pp. 3596–3605
2020
Cited alongside, same era.
F. Huang, L. Zhang, Y. Yang, and X. Zhou, “Probability weighted compact feature for domain adaptive retrieval,” in CVPR , 2020, pp. 9582–9591
2020
Cited alongside, same era.
S. Jandial, A. Chopra, P. Badjatiya, P. Chawla, M. Sarkar, and B. Krishnamurthy, “Trace: Transform aggregate and compose visiolinguistic representations for image search with text feedback,” arXiv , 2020
2020
Cited alongside, same era.
2022
Later among the works it cites.
T. Wei, D. Chen, W. Zhou, J. Liao, Z. Tan, L. Yuan, W. Zhang, and N. Yu, “Hairclip: Design your hair by text and reference image,” in CVPR , 2022, pp. 18 072–18 081
2022
Later among the works it cites.
Y. Xu, Y. Bin, G. Wang, and Y. Yang, “Hierarchical composition learning for composed query image retrieval,” in MMAsia , 2022
2022
Later among the works it cites.
Z. Xu, T. Lin, H. Tang, F. Li, D. He, N. Sebe, R. Timofte, L. Van Gool, and E. Ding, “Predict, prevent, and evaluate: Disentangled text-driven image manipulation empowered by pre-trained vision-language model,” in CVPR , 2022, pp. 18 229–18 238
2022
Later among the works it cites.
R. Yang, S. Wang, Y. Sun, H. Zhang, Y. Liao, Y. Gu, B. Hou, and L. Jiao, “Multimodal fusion remote sensing image–audio retrieval,” IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing , vol. 15, pp. 6220–6235, 2022
2022
Later among the works it cites.
F. Zhang, M. Xu, and C. Xu, “Geometry sensitive cross-modal reasoning for composed query based image retrieval,” IEEE TIP , vol. 31, pp. 1000–1011, 2022
2022
Later among the works it cites.
——, “Tell, imagine, and search: End-to-end learning for composing text and image to image retrieval,” ACM TOMM , vol. 18, no. 2, pp. 1–23, 2022
2022
Later among the works it cites.
F. Zhang, M. Yan, J. Zhang, and C. Xu, “Comprehensive relationship reasoning for composed query based image retrieval,” in ACM MM , 2022, p. 4655–4664
2022
Later among the works it cites.
G. Zhang, S. Wei, H. Pang, S. Qiu, and Y. Zhao, “Composed image retrieval via explicit erasure and replenishment with semantic alignment,” IEEE TIP , vol. 31, pp. 5976–5988, 2022
2022
Later among the works it cites.
Y. Zhao, Y. Song, and Q. Jin, “Progressive learning for image retrieval with hybrid-modality queries,” in ACM SIGIR , 2022
2022
Later among the works it cites.
Y. Zhu, H. Liu, Y. Song, Z. Yuan, X. Han, C. Yuan, Q. Chen, and J. Wang, “One model to edit them all: Free-form text-driven image manipulation with semantic modulations,” NeurIPS , 2022
2022
Later among the works it cites.
Y. Bai, X. Xu, Y. Liu, S. Khan, F. Khan, W. Zuo, R. S. M. Goh, and C.-M. Feng, “Sentence-level prompts benefit composed image retrieval,” arXiv , 2023
2023
Later among the works it cites.
A. Baldrati, L. Agnolucci, M. Bertini, and A. Del Bimbo, “Zero-shot composed image retrieval with textual inversion,” in ICCV , 2023, pp. 15 338–15 347
2023
Later among the works it cites.
——, “Composed image retrieval using contrastive learning and task-oriented clip-based features,” ACM Transactions on Multimedia Computing, Communications and Applications , vol. 20, no. 3, pp. 1–24, 2023
2023
Later among the works it cites.
T. Brooks, A. Holynski, and A. A. Efros, “Instructpix2pix: Learning to follow image editing instructions,” in CVPR , 2023, pp. 18 392–18 402
2023
Later among the works it cites.
M. Cao, X. Wang, Z. Qi, Y. Shan, X. Qie, and Y. Zheng, “Masactrl: Tuning-free mutual self-attention control for consistent image synthesis and editing,” in ICCV , 2023, pp. 22 560–22 570
2023
Later among the works it cites.
J. Chen and H. Lai, “Pretrain like you inference: Masked tuning improves zero-shot composed image retrieval,” arXiv , 2023
2023
Later among the works it cites.
——, “Ranking-aware uncertainty for text-guided image retrieval,” arXiv , 2023
2023
Later among the works it cites.
J. Choi, Y. Choi, Y. Kim, J. Kim, and S. Yoon, “Custom-edit: Text-guided image editing with customized diffusion models,” arXiv , 2023
2023
Later among the works it cites.
P. N. Chowdhury, A. K. Bhunia, A. Sain, S. Koley, T. Xiang, and Y.-Z. Song, “Scenetrilogy: On human scene-sketch and its complementarity with photo and text,” in CVPR , 2023, pp. 10 972–10 983
2023
Later among the works it cites.
G. Couairon, J. Verbeek, H. Schwenk, and M. Cord, “Diffedit: Diffusion-based semantic image editing with mask guidance,” in ICLR , 2023
2023
Later among the works it cites.
W. Dai, J. Li, D. Li, A. Tiong, J. Zhao, W. Wang, B. Li, P. Fung, and S. Hoi, “InstructBLIP: Towards general-purpose vision-language models with instruction tuning,” in NeurIPS , 2023
2023
Later among the works it cites.
W. Dong, S. Xue, X. Duan, and S. Han, “Prompt tuning inversion for text-driven image editing using diffusion models,” in ICCV , 2023, pp. 7430–7440
2023
Later among the works it cites.
X. Han, X. Zhu, L. Yu, L. Zhang, Y.-Z. Song, and T. Xiang, “Fame-vil: Multi-tasking vision-language model for heterogeneous fashion tasks,” in CVPR , 2023, pp. 2669–2680
2023
Later among the works it cites.
Z. Hu, X. Zhu, S. Tran, R. Vidal, and A. Dhua, “Provla: Compositional image search with progressive vision-language alignment and multimodal fusion,” in ICCV W , October 2023, pp. 2772–2777
2023
Later among the works it cites.
F. Huang, X. Lv, and L. Zhang, “Coarse-to-fine sparse self-attention for vehicle re-identification,” Knowledge-Based Systems , vol. 270, p. 110526, 2023
2023
Later among the works it cites.
F. Huang and L. Zhang, “Language guided local infiltration for interactive image retrieval,” in CVPR , 2023, pp. 6104–6113
2023
Later among the works it cites.
B. Kawar, S. Zada, O. Lang, O. Tov, H. Chang, T. Dekel, I. Mosseri, and M. Irani, “Imagic: Text-based real image editing with diffusion models,” in CVPR , 2023, pp. 6007–6017
2023
Later among the works it cites.
A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y. Lo et al. , “Segment anything,” in ICCV , 2023, pp. 4015–4026
2023
Later among the works it cites.
G. Kwon and J. C. Ye, “Diffusion-based image translation using disentangled style and content representation,” in ICLR , 2023
2023
Later among the works it cites.
J. Li, D. Li, S. Savarese, and S. Hoi, “Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,” in ICML , 2023, pp. 19 730–19 742
2023
Later among the works it cites.
S. Liu, Z. Zeng, T. Ren, F. Li, H. Zhang, J. Yang, C. Li, J. Yang, H. Su, J. Zhu et al. , “Grounding dino: Marrying dino with grounded pre-training for open-set object detection,” arXiv , 2023
2023
Later among the works it cites.
Y. Lyu, K. Zhao, B. Peng, Y. Jiang, Y. Zhang, and J. Dong, “Deltaspace: A semantic-aligned feature space for flexible text-guided image editing,” arXiv , 2023
2023
Later among the works it cites.
R. Mokady, A. Hertz, K. Aberman, Y. Pritch, and D. Cohen-Or, “Null-text inversion for editing real images using guided diffusion models,” in CVPR , 2023, pp. 6038–6047
2023
Later among the works it cites.
A. Pal, S. Wadhwa, A. Jaiswal, X. Zhang, Y. Wu, R. Chada, P. Natarajan, and H. I. Christensen, “Fashionntm: Multi-turn fashion image retrieval via cascaded memory,” in ICCV , 2023, pp. 11 323–11 334
2023
Later among the works it cites.
H. Ravi, S. Kelkar, M. Harikumar, and A. Kale, “Preditor: Text guided image editing with diffusion prior,” arXiv , 2023
2023
Later among the works it cites.
K. Saito, K. Sohn, X. Zhang, C.-L. Li, C.-Y. Lee, K. Saenko, and T. Pfister, “Pic2word: Mapping pictures to words for zero-shot composed image retrieval,” in CVPR , 2023, pp. 19 305–19 314
2023
Later among the works it cites.
M. Tao, B.-K. Bao, H. Tang, F. Wu, L. Wei, and Q. Tian, “De-net: Dynamic text-guided image editing adversarial networks,” in AAAI , vol. 37, no. 8, 2023, pp. 9971–9979
2023
Later among the works it cites.
D. Valevski, M. Kalman, E. Molad, E. Segalis, Y. Matias, and Y. Leviathan, “Unitune: Text-driven image editing by fine tuning a diffusion model on a single image,” ACM TOG , vol. 42, no. 4, pp. 1–10, 2023
2023
Later among the works it cites.
Q. Wang, B. Zhang, M. Birsak, and P. Wonka, “Instructedit: Improving automatic masks for diffusion-based image editing with user instructions,” arXiv , 2023
2023
Later among the works it cites.
——, “Mdp: A generalized framework for text-guided image editing by manipulating the diffusion path,” arXiv , 2023
2023
Later among the works it cites.
Y. Watanabe, R. Togo, K. Maeda, T. Ogawa, and M. Haseyama, “Text-guided image manipulation via generative adversarial network with referring image segmentation-based guidance,” IEEE Access , vol. 11, pp. 42 534–42 545, 2023
2023
Later among the works it cites.
H. Wen, X. Zhang, X. Song, Y. Wei, and L. Nie, “Target-guided composed image retrieval,” in ACM MM , ser. MM ’23. ACM, 2023
2023
Later among the works it cites.
C. Xiao, Q. Yang, X. Xu, J. Zhang, F. Zhou, and C. Zhang, “Where you edit is what you get: Text-guided image editing with region-based attention,” Pattern Recognition , vol. 139, p. 109458, 2023
2023
Later among the works it cites.
Y. Xu, Y. Bin, J. Wei, Y. Yang, G. Wang, and H. T. Shen, “Multi-modal transformer with global-local alignment for composed query image retrieval,” IEEE TMM , vol. 25, pp. 8346–8357, 2023
2023
Later among the works it cites.
Q. Yang, M. Ye, Z. Cai, K. Su, and B. Du, “Composed image retrieval via cross relation network with hierarchical aggregation transformer,” IEEE TIP , 2023
2023
Later among the works it cites.
Z. Zhang, L. Han, A. Ghosh, D. N. Metaxas, and J. Ren, “Sine: Single image editing with text-to-image diffusion models,” in CVPR , 2023, pp. 6027–6037
2023
Later among the works it cites.
Y. Zhou, F. Huang, W. Chen, S. Pu, and L. Zhang, “Stochastic gradient perturbation: An implicit regularizer for person re-identification,” IEEE TCSVT , vol. 33, no. 10, pp. 5894–5907, 2023
2023
Later among the works it cites.
H. Zhu, Y. Wei, Y. Zhao, C. Zhang, and S. Huang, “Amc: Adaptive multi-expert collaborative network for text-guided image retrieval,” vol. 19, no. 6, may 2023
2023
Later among the works it cites.
O. Barbany, M. Huang, X. Zhu, and A. Dhua, “Leveraging large language models for multimodal search,” in CVPR , 2024, pp. 1201–1210
2024
Closest in time.
Z. Feng, R. Zhang, and Z. Nie, “Improving composed image retrieval via contrastive learning with scaling positives and negatives,” arXiv , 2024
2024
Closest in time.
G. Gu, S. Chun, W. Kim, , Y. Kang, and S. Yun, “Language-only training of zero-shot composed image retrieval,” in CVPR , 2024
2024
Closest in time.
G. Gu, S. Chun, W. Kim, H. Jun, Y. Kang, and S. Yun, “Compodiff: Versatile composed image retrieval with latent diffusion,” in CVPR W , 2024
2024
Closest in time.
A. Hu, H. Xu, J. Ye, M. Yan, L. Zhang, B. Zhang, C. Li, J. Zhang, Q. Jin, F. Huang et al. , “mplug-docowl 1.5: Unified structure learning for ocr-free document understanding,” arXiv , 2024
2024
Closest in time.
F. Huang, L. Zhang, X. Fu, and S. Song, “Dynamic weighted combiner for mixed-modal image retrieval,” in AAAI , vol. 38, no. 3, 2024, pp. 2303–2311
2024
Closest in time.
S. Karthik, K. Roth, M. Mancini, and Z. Akata, “Vision-by-language for training-free compositional image retrieval,” in ICLR , 2024
2024
Closest in time.
S. Koley, A. K. Bhunia, A. Sain, P. N. Chowdhury, T. Xiang, and Y.-Z. Song, “You’ll never walk alone: A sketch and text duet for fine-grained image retrieval,” in CVPR , 2024, pp. 16 509–16 519
2024
Closest in time.
M. Levy, R. Ben-Ari, N. Darshan, and D. Lischinski, “Data roaming and quality assessment for composed image retrieval,” in AAAI , vol. 38, no. 4, 2024, pp. 2991–2999
2024
Closest in time.
Z. Liu, W. Sun, Y. Hong, D. Teney, and S. Gould, “Bi-directional training for composed image retrieval via text prompt learning,” in WACV , 2024, pp. 5753–5762
2024
Closest in time.
Z. Liu, W. Sun, D. Teney, and S. Gould, “Candidate set re-ranking for composed image retrieval with dual multi-modal encoder,” IEEE TMLR , 2024
2024
Closest in time.
C. Mou, X. Wang, L. Xie, Y. Wu, J. Zhang, Z. Qi, and Y. Shan, “T2i-adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models,” in AAAI , vol. 38, no. 5, 2024, pp. 4296–4304
2024
Closest in time.
D. H. Park, G. Luo, C. Toste, S. Azadi, X. Liu, M. Karalashvili, A. Rohrbach, and T. Darrell, “Shape-guided diffusion with inside-outside attention,” in WACV , 2024, pp. 4198–4207
2024
Closest in time.
Y. Ruan, H.-H. Lee, Y. Zhang, K. Zhang, and A. X. Chang, “Tricolo: Trimodal contrastive loss for text to shape retrieval,” in WACV , 2024, pp. 5815–5825
2024
Closest in time.
Y. Suo, F. Ma, L. Zhu, and Y. Yang, “Knowledge-enhanced dual-stream zero-shot composed image retrieval,” in CVPR , 2024, pp. 26 951–26 962
2024
Closest in time.
Y. Tang, J. Yu, K. Gai, J. Zhuang, G. Xiong, Y. Hu, and Q. Wu, “Context-i2w: Mapping images to context-dependent words for accurate zero-shot composed image retrieval,” in AAAI , vol. 38, no. 6, 2024, pp. 5180–5188
2024
Closest in time.
H. Wen, X. Song, X. Chen, Y. Wei, L. Nie, and T.-S. Chua, “Simple but effective raw-data level multimodal fusion for composed image retrieval,” in ACM SIGIR , ser. SIGIR 2024. ACM, Jul. 2024
2024
Closest in time.
——, “Align and retrieve: Composition and decomposition learning in image retrieval with text feedback,” IEEE TMM , 2024
2024
Closest in time.
K. Yin, S. Zou, Y. Ge, and Z. Tian, “Tri-modal motion retrieval by learning a joint embedding space,” in CVPR , 2024, pp. 1596–1605
2024
Closest in time.
G. Zhang, S. Li, S. Wei, S. Ge, N. Cai, and Y. Zhao, “Multimodal composition example mining for composed query image retrieval,” IEEE TIP , vol. 33, pp. 1149–1161, 2024
2024
Closest in time.
K. Zhang, Y. Luan, H. Hu, K. Lee, S. Qiao, W. Chen, Y. Su, and M.-W. Chang, “Magiclens: Self-supervised image retrieval with open-ended instructions,” in ICML , 2024
2024
Closest in time.
L. Zhang, X. Fu, F. Huang, Y. Yang, and X. Gao, “An open-world, diverse, cross-spatial-temporal benchmark for dynamic wild person re-identification,” IJCV , pp. 1–24, 2024
2024
Closest in time.
N. Zhang, Y. Liu, Z. Li, J. Xiang, and R. Pan, “Fabric image retrieval based on multi-modal feature fusion,” Signal, Image and Video Processing , vol. 18, no. 3, pp. 2207–2217, 2024
2024
Closest in time.
Z. Zhang, J. Zheng, Z. Fang, and B. A. Plummer, “Text-to-image editing by image information removal,” in WACV , 2024
2024
Closest in time.
L. Zheng, W.-L. Chiang, Y. Sheng, S. Zhuang, Z. Wu, Y. Zhuang, Z. Lin, Z. Li, D. Li, E. Xing et al. , “Judging llm-as-a-judge with mt-bench and chatbot arena,” NeurIPS , vol. 36, 2024
2024
Closest in time.
H. Zhu, J.-H. Huang, S. Rudinac, and E. Kanoulas, “Enhancing interactive image retrieval with query rewriting using large language models and vision language models,” in ICMR , 2024
2024
Closest in time.
O. Patashnik, Z. Wu, E. Shechtman, D. Cohen-Or, and D. Lischinski, “Styleclip: Text-driven manipulation of stylegan imagery,” in ICCV , 2021, pp. 2085–2094
2094
Closest in time.