Fetching the paper…
Reading the bibliography…
Composed Image Retrieval (CoIR) has recently gained popularity as a task that considers both text and image queries together, to search for relevant images in a database.
G. Farnebäck, “Two-frame motion estimation based on polynomial expansion,” in Image Analysis . Springer Berlin Heidelberg, 2003
2003
Earlier work this paper cites.
S. Loria, textblob.readthedocs.io , 2013
2013
Earlier work this paper cites.
T. Lin, M. Maire, S. J. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick, “Microsoft COCO: Common objects in context,” in ECCV , 2014
2014
Earlier work this paper cites.
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh, “VQA: Visual Question Answering,” in ICCV , 2015
2015
Earlier work this paper cites.
P. Sharma, N. Ding, S. Goodman, and R. Soricut, “Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning,” in ACL , 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
X. Guo, H. Wu, Y. Cheng, S. Rennie, G. Tesauro, and R. S. Feris, “Dialog-based interactive image retrieval,” in NeurIPS , 2018
2018
Earlier work this paper cites.
R. Speer, J. Chin, A. Lin, S. Jewett, and L. Nathan, “Luminosoinsight/wordfreq: v2.2,” Oct. 2018. [Online]. Available: https://doi.org/10.5281/zenodo.1443582
2018
Earlier work this paper cites.
A. Fan, M. Lewis, and Y. Dauphin, “Hierarchical neural story generation,” in ACL , 2018
2018
Earlier work this paper cites.
S. N. Thanh, https://pypi.org/project/better-profanity/ , 2018
2018
Earlier work this paper cites.
N. Vo, L. Jiang, C. Sun, K. Murphy, L.-J. Li, L. Fei-Fei, and J. Hays, “Composing text and image for image retrieval - an empirical odyssey,” in CVPR , 2019
2019
Earlier work this paper cites.
W. Su, X. Zhu, Y. Cao, B. Li, L. Lu, F. Wei, and J. Dai, “VL-BERT: Pre-training of generic visual-linguistic representations,” in ICLR , 2019
2019
Earlier work this paper cites.
A. Miech, D. Zhukov, J.-B. Alayrac, M. Tapaswi, I. Laptev, and J. Sivic, “HowTo100M: Learning a text-video embedding by watching hundred million narrated video clips,” in ICCV , 2019
2019
Earlier work this paper cites.
A. Suhr, S. Zhou, A. Zhang, I. Zhang, H. Bai, and Y. Artzi, “A corpus for reasoning about natural language grounded in photographs,” in ACL , 2019
2019
Earlier work this paper cites.
Y. Yu, S. Lee, Y. Choi, and G. Kim, “CurlingNet: Compositional learning between images and text for fashionIQ data,” ICCV Workshop , 2019
2019
Earlier work this paper cites.
J. Johnson, M. Douze, and H. Jégou, “Billion-scale similarity search with GPUs,” IEEE Transactions on Big Data , 2019
2019
Earlier work this paper cites.
I. Loshchilov and F. Hutter, “Decoupled weight decay regularization,” in ICLR , 2019
2019
Earlier work this paper cites.
Y.-C. Chen, L. Li, L. Yu, A. E. Kholy, F. Ahmed, Z. Gan, Y. Cheng, and J. Liu, “UNITER: Universal image-text representation learning,” in ECCV , 2020
2020
Earlier work this paper cites.
X. Li, X. Yin, C. Li, P. Zhang, X. Hu, L. Zhang, L. Wang, H. Hu, L. Dong, F. Wei et al. , “Oscar: Object-semantics aligned pre-training for vision-language tasks,” in ECCV , 2020
2020
Earlier work this paper cites.
L. Li, Y.-C. Chen, Y. Cheng, Z. Gan, L. Yu, and J. Liu, “HERO: Hierarchical encoder for video+language omni-representation pre-training,” in EMNLP , 2020
2020
Earlier work this paper cites.
A. Miech, J.-B. Alayrac, L. Smaira, I. Laptev, J. Sivic, and A. Zisserman, “End-to-end learning of visual representations from uncurated instructional videos,” in CVPR , 2020
2020
Earlier work this paper cites.
A. Nagrani, C. Sun, D. Ross, R. Sukthankar, C. Schmid, and A. Zisserman, “Speech2action: Cross-modal supervision for action recognition,” in CVPR , 2020
2020
Earlier work this paper cites.
Y. Chen and L. Bazzani, “Learning joint visual semantic matching embeddings for language-guided retrieval,” in ECCV , 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
Y. Chen, S. Gong, and L. Bazzani, “Image search with text feedback by visiolinguistic attention learning,” in CVPR , 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
T. Brown et al., “Language models are few-shot learners,” in NeurIPS , 2020
2020
Cited alongside, same era.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever, “Learning transferable visual models from natural language supervision,” in ICML , 2021
2021
Cited alongside, same era.
Z. Liu, C. Rodriguez-Opazo, D. Teney, and S. Gould, “Image retrieval on real-life images with pre-trained vision-and-language models,” in ICCV , 2021
2021
Cited alongside, same era.
H. Wu, Y. Gao, X. Guo, Z. Al-Halah, S. Rennie, K. Grauman, and R. Feris, “Fashion IQ: A new dataset towards retrieving images by natural language feedback,” in CVPR , 2021
2021
Cited alongside, same era.
M. Bain, A. Nagrani, G. Varol, and A. Zisserman, “Frozen in time: A joint video and image encoder for end-to-end retrieval,” in ICCV , 2021
Y. Ge, Y. Ge, X. Liu, D. Li, Y. Shan, X. Qie, and P. Luo, “BridgeFormer: Bridging video-text retrieval with multiple choice questions,” in CVPR , 2022
2022
Later among the works it cites.
Y. Liu, P. Xiong, L. Xu, S. Cao, and Q. Jin, “TS2-Net: Token shift and selection transformer for text-video retrieval,” in ECCV , 2022
2022
Later among the works it cites.
Y. Ma, G. Xu, X. Sun, M. Yan, J. Zhang, and R. Ji, “X-CLIP: End-to-end multi-grained contrastive learning for video-text retrieval,” in ACMMM , 2022
2022
Later among the works it cites.
L. Yao, R. Huang, L. Hou, G. Lu, M. Niu, H. Xu, X. Liang, Z. Li, X. Jiang, and C. Xu, “FILIP: Fine-grained interactive language-image pre-training,” in ICLR , 2022
2022
Later among the works it cites.
Y. Fang, W. Wang, B. Xie, Q. Sun, L. Wu, X. Wang, T. Huang, X. Wang, and Y. Cao, “Eva: Exploring the limits of masked visual representation learning at scale,” 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2021
Cited alongside, same era.
J. Kim, Y. Yu, H. Kim, and G. Kim, “Dual compositional learning in interactive image retrieval,” AAAI , 2021
2021
Cited alongside, same era.
J. Li, R. Selvaraju, A. Gotmare, S. Joty, C. Xiong, and S. C. H. Hoi, “Align before fuse: Vision and language representation learning with momentum distillation,” in NeurIPS , 2021
2021
Cited alongside, same era.
C. Jia, Y. Yang, Y. Xia, Y.-T. Chen, Z. Parekh, H. Pham, Q. Le, Y.-H. Sung, Z. Li, and T. Duerig, “Scaling up visual and vision-language representation learning with noisy text supervision,” in ICML , 2021
2021
Cited alongside, same era.
H. Akbari, L. Yuan, R. Qian, W.-H. Chuang, S.-F. Chang, Y. Cui, and B. Gong, “Vatt: Transformers for multimodal self-supervised learning from raw video, audio and text,” NeurIPS , 2021
2021
Cited alongside, same era.
H. Xu, G. Ghosh, P.-Y. Huang, D. Okhonko, A. Aghajanyan, F. Metze, L. Zettlemoyer, and C. Feichtenhofer, “VideoCLIP: Contrastive pre-training for zero-shot video-text understanding,” in EMNLP , 2021
2021
Cited alongside, same era.
A. Yang, A. Miech, J. Sivic, I. Laptev, and C. Schmid, “Just ask: Learning to answer questions from millions of narrated videos,” in ICCV , 2021
2021
Cited alongside, same era.
2021
Cited alongside, same era.
2022
Later among the works it cites.
S. Jandial, P. Badjatiya, P. Chawla, A. Chopra, M. Sarkar, and B. Krishnamurthy, “SAC: Semantic attention composition for text-conditioned image retrieval,” in WACV , 2022
2022
Later among the works it cites.
S. Goenka, Z. Zheng, A. Jaiswal, R. Chada, Y. Wu, V. Hedau, and P. Natarajan, “FashionVLP: Vision language transformer for fashion retrieval with feedback,” in CVPR , 2022
2022
Later among the works it cites.
A. Baldrati, L. Agnolucci, M. Bertini, and A. Del Bimbo, “Zero-shot composed image retrieval with textual inversion,” ICCV , 2023
2023
Closest in time.
2023
Closest in time.
T. Brooks, A. Holynski, and A. A. Efros, “InstructPix2Pix: Learning to follow image editing instructions,” in CVPR , 2023
2023
Closest in time.
F. Radenovic, A. Dubey, A. Kadian, T. Mihaylov, S. Vandenhende, Y. Patel, Y. Wen, V. Ramanathan, and D. Mahajan, “Filtering, distillation, and hard negatives for vision-language pre-training,” in CVPR , 2023
2023
Closest in time.
J. Li, D. Li, S. Savarese, and S. Hoi, “BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,” in ICML , 2023
2023
Closest in time.
K. Saito, K. Sohn, X. Zhang, C.-L. Li, C.-Y. Lee, K. Saenko, and T. Pfister, “Pic2Word: Mapping pictures to words for zero-shot composed image retrieval,” CVPR , 2023
2023
Closest in time.
2023
Closest in time.
Y. Zhao, I. Misra, P. Krähenbühl, and R. Girdhar, “Learning video representations from large language models,” in CVPR , 2023
2023
Closest in time.
H. Liu, C. Li, Q. Wu, and Y. J. Lee, “Visual instruction tuning,” in NeurIPS , 2023
2023
Closest in time.
H. Rasheed, M. U. Khattak, M. Maaz, S. Khan, and F. S. Khan, “Fine-tuned CLIP models are efficient video learners,” in CVPR , 2023
2023
Closest in time.
H. Xue, Y. Sun, B. Liu, J. Fu, R. Song, H. Li, and J. Luo, “CLIP-vip: Adapting pre-trained image-text model to video-language alignment,” in ICLR , 2023
2023
Closest in time.
H. Touvron et al., “LLaMA 2: Open foundation and fine-tuned chat models,” arXiv:2307.09288 , 2023
2023
Closest in time.
S. A. Lab, https://huggingface.co/datasets/shinonomelab/cleanvid-15m_map , 2023
2023
Closest in time.
G. Gu, S. Chun, W. Kim, H. Jun, Y. Kang, and S. Yun, “Compodiff: Versatile composed image retrieval with latent diffusion,” in TMLR , 2024
2024
Closest in time.
M. Levy, R. Ben-Ari, N. Darshan, and D. Lischinski, “Data roaming and early fusion for composed image retrieval,” in AAAI , 2024
2024
Closest in time.
K. Zhang, Y. Luan, H. Hu, K. Lee, S. Qiao, W. Chen, Y. Su, and M.-W. Chang, “MagicLens: Self-supervised image retrieval with open-ended instructions,” in ICML , ser. Proceedings of Machine Learning Research, vol. 235. PMLR, 21–27 Jul 2024, pp. 59 403–59 420
2024
Closest in time.
G. Gu, S. Chun, W. Kim, , Y. Kang, and S. Yun, “Language-only training of zero-shot composed image retrieval,” in CVPR , 2024
2024
Closest in time.
S. Karthik, K. Roth, M. Mancini, and Z. Akata, “Vision-by-language for training-free compositional image retrieval,” in ICLR , 2024
2024
Closest in time.
L. Ventura, A. Yang, C. Schmid, and G. Varol, “CoVR: Learning composed video retrieval from web video captions,” AAAI , 2024
2024
Closest in time.