Fetching the paper…
Reading the bibliography…
The task of Composed Image Retrieval (CoIR) involves queries that combine image and text modalities, allowing users to express their intent more effectively.
Language Models are Few-Shot Learners
Brown, T.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J. D.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; Agarwal, S.; Herbert-Voss, A.; Krueger, G.; Henighan, T.; Child, R.; Ramesh, A.; Ziegler, D.; Wu, J.; Winter, C.; Hesse, C.; Chen, M.; Sigler, E.; Litwin, M.; Gray, S.; Chess, B.; Clark, J.; Berner, C.; McCandlish, S.; Radford, A.; Sutskever, I.; and Amodei, D. 2020 · 1901
Earlier work this paper cites.
CurlingNet: Compositional Learning between Images and Text for Fashion IQ Data
Yu, Y.; Lee, S.; Choi, Y.; and Kim, G. 2020 · 2003
Earlier work this paper cites.
Modality-Agnostic Attention Fusion for visual search with text feedback
Dodds, E.; Culpepper, J.; Herdade, S.; Zhang, Y.; and Boakye, K. 2020 · 2007
Earlier work this paper cites.
Jandial, S.; Chopra, A.; Badjatiya, P.; Chawla, P.; Sarkar, M.; and Krishnamurthy, B. 2020 · 2009
Earlier work this paper cites.
Collecting Highly Parallel Data for Paraphrase Evaluation
Chen, D. L.; and Dolan, W. B. 2011 · 2011
Earlier work this paper cites.
Microsoft COCO: Common Objects in Context
Lin, T.; Maire, M.; Belongie, S. J.; Hays, J.; Perona, P.; Ramanan, D.; Dollár, P.; and Zitnick, C. L. 2014 · 2014
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Young, P.; Lai, A.; Hodosh, M.; and Hockenmaier, J. 2014 · 2014
Earlier work this paper cites.
VQA: Visual Question Answering
Antol, S.; Agrawal, A.; Lu, J.; Mitchell, M.; Batra, D.; Zitnick, C. L.; and Parikh, D. 2015 · 2015
Earlier work this paper cites.
Discovering states and transformations in image collections
Isola, P.; Lim, J. J.; and Adelson, E. H. 2015 · 2015
Earlier work this paper cites.
Deep Residual Learning for Image Recognition
He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016 · 2016
Earlier work this paper cites.
MSR-VTT: A Large Video Description Dataset for Bridging Video and Language
Xu, J.; Mei, T.; Yao, T.; and Rui, Y. 2016 · 2016
Earlier work this paper cites.
Visual Dialog
Das, A.; Kottur, S.; Gupta, K.; Singh, A.; Yadav, D.; Moura, J. M.; Parikh, D.; and Batra, D. 2017 · 2017
Earlier work this paper cites.
Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering
Goyal, Y.; Khot, T.; Summers-Stay, D.; Batra, D.; and Parikh, D. 2017 · 2017
Earlier work this paper cites.
Automatic Spatially-Aware Fashion Concept Discovery
Han, X.; Wu, Z.; Huang, P. X.; Zhang, X.; Zhu, M.; Li, Y.; Zhao, Y.; and Davis, L. S. 2017 · 2017
Earlier work this paper cites.
Attention Is All You Need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Earlier work this paper cites.
Dialog-based Interactive Image Retrieval
Guo, X.; Wu, H.; Cheng, Y.; Rennie, S.; Tesauro, G.; and Feris, R. S. 2018 · 2018
Earlier work this paper cites.
Nocaps: Novel object captioning at scale
Agrawal, H.; Desai, K.; Wang, Y.; Chen, X.; Jain, R.; Johnson, M.; Batra, D.; Parikh, D.; Lee, S.; and Anderson, P. 2019 · 2019
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Devlin, J.; Chang, M.; Lee, K.; and Toutanova, K. 2019 · 2019
Cited alongside, same era.
ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks
Lu, J.; Batra, D.; Parikh, D.; and Lee, S. 2019 · 2019
Cited alongside, same era.
Howto100m: Learning a text-video embedding by watching hundred million narrated video clips
Miech, A.; Zhukov, D.; Alayrac, J.-B.; Tapaswi, M.; Laptev, I.; and Sivic, J. 2019 · 2019
Cited alongside, same era.
A Corpus for Reasoning about Natural Language Grounded in Photographs
Suhr, A.; Zhou, S.; Zhang, A.; Zhang, I.; Bai, H.; and Artzi, Y. 2019 · 2019
Cited alongside, same era.
Composing Text and Image for Image Retrieval - an Empirical Odyssey
Vo, N.; Jiang, L.; Sun, C.; Murphy, K.; Li, L.-J.; Fei-Fei, L.; and Hays, J. 2019 · 2019
Cited alongside, same era.
Align before Fuse: Vision and Language Representation Learning with Momentum Distillation
Li, J.; Selvaraju, R. R.; Gotmare, A.; Joty, S. R.; Xiong, C.; and Hoi, S. C. 2021 · 2021
Later among the works it cites.
Image Retrieval on Real-life Images with Pre-trained Vision-and-Language Models
Liu, Z.; Rodriguez-Opazo, C.; Teney, D.; and Gould, S. 2021 · 2021
Later among the works it cites.
Learning Transferable Visual Models From Natural Language Supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; Krueger, G.; and Sutskever, I. 2021 · 2021
Later among the works it cites.
RTIC: Residual Learning for Text and Image Composition using Graph Convolutional Network
Shin, M.; Cho, Y.; Ko, B.; and Gu, G. 2021 · 2021
Later among the works it cites.
Explainable, interactive c ontent-based image retrieval
Vasu, B.; Hu, B.; Dong, B.; Collins, R.; and Hoogs, A. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning Joint Visual Semantic Matching Embeddings for Language-Guided Retrieval
Chen, Y.; and Bazzani, L. 2020 · 2020
Cited alongside, same era.
Image Search With Text Feedback by Visiolinguistic Attention Learning
Chen, Y.; Gong, S.; and Bazzani, L. 2020 · 2020
Cited alongside, same era.
Composed Query Image Retrieval Using Locally Bounded Features
Hosseinzadeh, M.; and Wang, Y. 2020 · 2020
Cited alongside, same era.
Oscar: Object-Semantics Aligned Pre-training for Vision-Language Tasks
Li, X.; Yin, X.; Li, C.; Zhang, P.; Hu, X.; Zhang, L.; Wang, L.; Hu, H.; Dong, L.; Wei, F.; Choi, Y.; and Gao, J. 2020 · 2020
Cited alongside, same era.
Few-shot learning for remote sensing image retrieval with maml
Zhong, Q.; Chen, L.; and Qian, Y. 2020 · 2020
Cited alongside, same era.
Content-based image retrieval and the semantic gap in the deep learning era
Barz, B.; and Denzler, J. 2021 · 2021
Cited alongside, same era.
Generic Attention-Model Explainability for Interpreting Bi-Modal and Encoder-Decoder Transformers
Chefer, H.; Gur, S.; and Wolf, L. 2021 · 2021
Cited alongside, same era.
Fashion IQ: A New Dataset Towards Retrieving Images by Natural Language Feedback
Wu, H.; Gao, Y.; Guo, X.; Al-Halah, Z.; Rennie, S.; Grauman, K.; and Feris, R. 2021 · 2021
Later among the works it cites.
Just ask: Learning to answer questions from millions of narrated videos
Yang, A.; Miech, A.; Sivic, J.; Laptev, I.; and Schmid, C. 2021 · 2021
Later among the works it cites.
Effective conditioned and composed image retrieval combining CLIP-based features
Baldrati, A.; Bertini, M.; Uricchio, T.; and Bimbo, A. D. 2022 · 2022
Later among the works it cites.
Embedding Arithmetic of Multimodal Queries for Image Retrieval
Couairon, G.; Douze, M.; Cord, M.; and Schwenk, H. 2022 · 2022
Later among the works it cites.
ARTEMIS: Attention-based Retrieval with Text-Explicit Matching and Implicit Similarity
Delmas, G.; de Rezende, R. S.; Csurka, G.; and Larlus, D. 2022 · 2022
Later among the works it cites.
FashionVLP: Vision Language Transformer for Fashion Retrieval with Feedback
Goenka, S.; Zheng, Z.; Jaiswal, A.; Chada, R.; Wu, Y.; Hedau, V.; and Natarajan, P. 2022 · 2022
Later among the works it cites.
SAC: Semantic Attention Composition for Text-Conditioned Image Retrieval
Jandial, S.; Badjatiya, P.; Chawla, P.; Chopra, A.; Sarkar, M.; and Krishnamurthy, B. 2022 · 2022
Later among the works it cites.
Classification-Regression for Chart comprehension
Levy, M.; Ben-Ari, R.; and Lischinski, D. 2022 · 2022
Later among the works it cites.
BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Li, J.; Li, D.; Xiong, C.; and Hoi, S. C. H. 2022 · 2022
Later among the works it cites.
Learning Audio-Video Modalities from Image Captions
Nagrani, A.; Seo, P. H.; Seybold, B.; Hauth, A.; Manen, S.; Sun, C.; and Schmid, C. 2022 · 2022
Later among the works it cites.
Recall@k surrogate loss with large batches and similarity mixup
Patel, Y.; Tolias, G.; and Matas, J. 2022 · 2022
Later among the works it cites.