Fetching the paper…
Reading the bibliography…
Image retrieval, i.e., finding desired images given a reference image, inherently encompasses rich, multi-faceted search intents that are difficult to capture solely using image-based measures.
Image retrieval: Ideas, influences, and trends of the new age
Datta, R., Joshi, D., Li, J., and Wang, J. Z · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Microsoft COCO: common objects in context
Lin, T., Maire, M., Belongie, S. J., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L · 2014
Earlier work this paper cites.
VQA: visual question answering
Antol, S., Agrawal, A., Lu, J., Mitchell, M., Batra, D., Zitnick, C. L., and Parikh, D · 2015
Earlier work this paper cites.
Microsoft coco captions: Data collection and evaluation server
Chen, X., Fang, H., Lin, T.-Y., Vedantam, R., Gupta, S., Dollar, P., and Zitnick, C. L · 2015
Earlier work this paper cites.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
Plummer, B. A., Wang, L., Cervantes, C. M., Caicedo, J. C., Hockenmaier, J., and Lazebnik, S · 2015
Earlier work this paper cites.
Deep image retrieval: Learning global representations for image search
Gordo, A., Almazán, J., Revaud, J., and Larlus, D · 2016
Earlier work this paper cites.
Sketchnet: Sketch classification with web images
Zhang, H., Liu, S., Zhang, C., Ren, W., Wang, R., and Cao, X · 2016
Earlier work this paper cites.
Vse++: Improving visual-semantic embeddings with hard negatives
Faghri, F., Fleet, D. J., Kiros, J. R., and Fidler, S · 2017
Earlier work this paper cites.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Sharma, P., Ding, N., Goodman, S., and Soricut, R · 2018
Earlier work this paper cites.
Adafactor: Adaptive learning rates with sublinear memory cost
Shazeer, N. and Stern, M · 2018
Earlier work this paper cites.
A zero-shot framework for sketch based image retrieval
Yelamarthi, S. K., Reddy, S. K., Mishra, A., and Mittal, A · 2018
Earlier work this paper cites.
Doodle to search: Practical zero-shot sketch-based image retrieval
Dey, S., Riba, P., Dutta, A., Llados, J., and Song, Y.-Z · 2019
Earlier work this paper cites.
Semantic-aware knowledge preservation for zero-shot sketch-based image retrieval
Liu, Q., Xie, L., Wang, H., and Yuille, A · 2019
Earlier work this paper cites.
A corpus for reasoning about natural language grounded in photographs
Suhr, A., Zhou, S., Zhang, A., Zhang, I., Bai, H., and Artzi, Y · 2019
Earlier work this paper cites.
Composing text and image for image retrieval - an empirical odyssey
Vo, N., Jiang, L., Sun, C., Murphy, K., Li, L.-J., Fei-Fei, L., and Hays, J · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Earlier work this paper cites.
Learning the best pooling strategy for visual semantic embedding
Chen, J., Hu, H., Wu, H., Jiang, Y., and Wang, C · 2021
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N · 2021
Cited alongside, same era.
The many faces of robustness: A critical analysis of out-of-distribution generalization
Hendrycks, D., Basart, S., Mu, N., Kadavath, S., Wang, F., Dorundo, E., Desai, R., Zhu, T., Parajuli, S., Guo, M., Song, D., Steinhardt, J., and Gilmer, J · 2021
Cited alongside, same era.
Scaling up visual and vision-language representation learning with noisy text supervision
Jia, C., Yang, Y., Xia, Y., Chen, Y.-T., Parekh, Z., Pham, H., Le, Q., Sung, Y.-H., Li, Z., and Duerig, T · 2021
Cited alongside, same era.
Vilt: Vision-and-language transformer without convolution or region supervision
Kim, W., Son, B., and Kim, I · 2021
Cited alongside, same era.
Align before fuse: Vision and language representation learning with momentum distillation
Li, J., Selvaraju, R., Gotmare, A., Joty, S., Xiong, C., and Hoi, S. C. H · 2021
Cited alongside, same era.
Palm 2 technical report
Anil, R., Dai, A. M., Firat, O., Johnson, M., Lepikhin, D., Passos, A., Shakeri, S., Taropa, E., Bailey, P., Chen, Z., Chu, E., Clark, J. H., Shafey, L. E., Huang, Y., Meier-Hellstern, K., Mishra, G., Moreira, E., Omernick, M., Robinson, K., Ruder, S., Tay, Y., Xiao, K., Xu, Y., Zhang, Y., Abrego, G. H., Ahn, J., Austin, J., Barham, P., Botha, J., Bradbury, J., Brahma, S., Brooks, K., Catasta, M., Cheng, Y., Cherry, C., Choquette-Choo, C. A., Chowdhery, A., Crepy, C., Dave, S., Dehghani, M., Dev, S., Devlin, J., Díaz, M., Du, N., Dyer, E., Feinberg, V., Feng, F., Fienber, V., Freitag, M., Garcia, X., Gehrmann, S., Gonzalez, L., Gur-Ari, G., Hand, S., Hashemi, H., Hou, L., Howland, J., Hu, A., Hui, J., Hurwitz, J., Isard, M., Ittycheriah, A., Jagielski, M., Jia, W., Kenealy, K., Krikun, M., Kudugunta, S., Lan, C., Lee, K., Lee, B., Li, E., Li, M., Li, W., Li, Y., Li, J., Lim, H., Lin, H., Liu, Z., Liu, F., Maggioni, M., Mahendru, A., Maynez, J., Misra, V., Moussalem, M., Nado, Z., Nham, J., Ni, E., Nystrom, A., Parrish, A., Pellat, M., Polacek, M., Polozov, A., Pope, R., Qiao, S., Reif, E., Richter, B., Riley, P., Ros, A. C., Roy, A., Saeta, B., Samuel, R., Shelby, R., Slone, A., Smilkov, D., So, D. R., Sohn, D., Tokumine, S., Valter, D., Vasudevan, V., Vodrahalli, K., Wang, X., Wang, P., Wang, Z., Wang, T., Wieting, J., Wu, Y., Xu, K., Xu, Y., Xue, L., Yin, P., Yu, J., Zhang, Q., Zheng, S., Zheng, C., Zhou, W., Zhou, D., Petrov, S., and Wu, Y · 2023
Later among the works it cites.
Task-aware retrieval with instructions
Asai, A., Schick, T., Lewis, P., Chen, X., Izacard, G., Riedel, S., Hajishirzi, H., and Yih, W.-t · 2023
Later among the works it cites.
Zero-shot composed image retrieval with textual inversion
Baldrati, A., Agnolucci, L., Bertini, M., and Del Bimbo, A · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Image retrieval on real-life images with pre-trained vision-and-language models
Liu, Z., Rodriguez-Opazo, C., Teney, D., and Gould, S · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I · 2021
Cited alongside, same era.
Fashion iq: A new dataset towards retrieving images by natural language feedback
Wu, H., Gao, Y., Guo, X., Al-Halah, Z., Rennie, S., Grauman, K., and Feris, R · 2021
Cited alongside, same era.
Murag: Multimodal retrieval-augmented generator for open question answering over images and text
Chen, W., Hu, H., Chen, X., Verga, P., and Cohen, W. W · 2022
Cited alongside, same era.
Scaling instruction-finetuned language models
Chung, H. W., Hou, L., Longpre, S., Zoph, B., Tay, Y., Fedus, W., Li, Y., Wang, X., Dehghani, M., Brahma, S., Webson, A., Gu, S. S., Dai, Z., Suzgun, M., Chen, X., Chowdhery, A., Castro-Ros, A., Pellat, M., Robinson, K., Valter, D., Narang, S., Mishra, G., Yu, A., Zhao, V., Huang, Y., Dai, A., Yu, H., Petrov, S., Chi, E. H., Dean, J., Devlin, J., Roberts, A., Zhou, D., Le, Q. V., and Wei, J · 2022
Cited alongside, same era.
“this is my unicorn, fluffy”: Personalizing frozen vision-language representations
Cohen, N., Gal, R., Meirom, E. A., Chechik, G., and Atzmon, Y · 2022
Cited alongside, same era.
BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Li, J., Li, D., Xiong, C., and Hoi, S · 2022
Cited alongside, same era.
Instructpix2pix: Learning to follow image editing instructions
Brooks, T., Holynski, A., and Efros, A. A · 2023
Later among the works it cites.
Pretrain like your inference: Masked tuning improves zero-shot composed image retrieval
Chen, J. and Lai, H · 2023
Later among the works it cites.
Reproducible scaling laws for contrastive language-image learning
Cherti, M., Beaumont, R., Wightman, R., Wortsman, M., Ilharco, G., Gordon, C., Schuhmann, C., Schmidt, L., and Jitsev, J · 2023
Later among the works it cites.
Compodiff: Versatile composed image retrieval with latent diffusion
Gu, G., Chun, S., Kim, W., Jun, H., Kang, Y., and Yun, S · 2023
Later among the works it cites.
BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Li, J., Li, D., Savarese, S., and Hoi, S · 2023
Later among the works it cites.
Zero-shot everything sketch-based image retrieval, and in explainable style
Lin, F., Li, M., Li, D., Hospedales, T., Song, Y.-Z., and Qi, Y · 2023
Later among the works it cites.
Pic2word: Mapping pictures to words for zero-shot composed image retrieval
Saito, K., Sohn, K., Zhang, X., Li, C.-L., Lee, C.-Y., Saenko, K., and Pfister, T · 2023
Later among the works it cites.
One embedder, any task: Instruction-finetuned text embeddings
Su, H., Shi, W., Kasai, J., Wang, Y., Hu, Y., Ostendorf, M., Yih, W.-t., Smith, N. A., Zettlemoyer, L., and Yu, T · 2023
Later among the works it cites.
Genecis: A benchmark for general conditional image similarity
Vaze, S., Carion, N., and Misra, I · 2023
Later among the works it cites.
Image as a foreign language: Beit pretraining for vision and vision-language tasks
Wang, W., Bao, H., Dong, L., Bjorck, J., Peng, Z., Liu, Q., Aggarwal, K., Mohammed, O. K., Singhal, S., Som, S., and Wei, F · 2023
Later among the works it cites.
Uniir: Training and benchmarking universal multimodal information retrievers
Wei, C., Chen, Y., Chen, H., Hu, H., Zhang, G., Fu, J., Ritter, A., and Chen, W · 2023
Later among the works it cites.
Language-only training of zero-shot composed image retrieval
Gu, G., Chun, S., Kim, W., , Kang, Y., and Yun, S · 2024
Closest in time.
Vision-by-language for training-free compositional image retrieval
Karthik, S., Roth, K., Mancini, M., and Akata, Z · 2024
Closest in time.
A comprehensive survey on instruction following
Lou, R., Zhang, K., and Yin, W · 2024
Closest in time.
Context-i2w: Mapping images to context-dependent words for accurate zero-shot composed image retrieval
Tang, Y., Yu, J., Gai, K., Zhuang, J., Xiong, G., Hu, Y., and Wu, Q · 2024
Closest in time.