Fetching the paper…
Reading the bibliography…
Text-image composed retrieval aims to retrieve the target image through the composed query, which is specified in the form of an image plus some text that describes desired modifications to the input image.
Automatic attribute discovery and characterization from noisy web data
Tamara L Berg, Alexander C Berg, and Jonathan Shih · 2010
Earlier work this paper cites.
Relative attributes
Devi Parikh and Kristen Grauman · 2011
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
Discovering states and transformations in image collections
Phillip Isola, Joseph J Lim, and Edward H Adelson · 2015
Earlier work this paper cites.
Automatic spatially-aware fashion concept discovery
Xintong Han, Zuxuan Wu, Phoenix X Huang, Xiao Zhang, Menglong Zhu, Yuan Li, Yang Zhao, and Larry S Davis · 2017
Earlier work this paper cites.
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
Justin Johnson, Bharath Hariharan, Laurens Van Der Maaten, Li Fei-Fei, C Lawrence Zitnick, and Ross Girshick · 2017
Earlier work this paper cites.
Spatial-semantic image search by visual feature synthesis
Long Mai, Hailin Jin, Zhe Lin, Chen Fang, Jonathan Brandt, and Feng Liu · 2017
Earlier work this paper cites.
Robust features in deep-learning-based speech recognition
Vikramjit Mitra, Horacio Franco, Richard M Stern, Julien Van Hout, Luciana Ferrer, Martin Graciarena, Wen Wang, Dimitra Vergyri, Abeer Alwan, and John HL Hansen · 2017
Earlier work this paper cites.
Foil it! find one mismatch between image and language caption
Ravi Shekhar, Sandro Pezzelle, Yauhen Klimovich, Aurélie Herbelot, Moin Nabi, Enver Sangineto, and Raffaella Bernardi · 2017
Earlier work this paper cites.
A corpus for reasoning about natural language grounded in photographs
Alane Suhr, Stephanie Zhou, Ally Zhang, Iris Zhang, Huajun Bai, and Yoav Artzi · 2018
Earlier work this paper cites.
Benchmarking neural network robustness to common corruptions and perturbations
Dan Hendrycks and Thomas Dietterich · 2019
Earlier work this paper cites.
Models in the wild: On corruption robustness of neural nlp systems
Barbara Rychalska, Dominika Basaj, Alicja Gosiewska, and Przemysław Biecek · 2019
Earlier work this paper cites.
Composing text and image for image retrieval-an empirical odyssey
Nam Vo, Lu Jiang, Chen Sun, Kevin Murphy, Li-Jia Li, Li Fei-Fei, and James Hays · 2019
Earlier work this paper cites.
Towards causal vqa: Revealing and reducing spurious correlations by invariant and covariant semantic editing
Vedika Agarwal, Rakshith Shetty, and Mario Fritz · 2020
Earlier work this paper cites.
Image search with text feedback by visiolinguistic attention learning
Yanbei Chen, Shaogang Gong, and Loris Bazzani · 2020
Earlier work this paper cites.
Robustbench: a standardized adversarial robustness benchmark
Francesco Croce, Maksym Andriushchenko, Vikash Sehwag, Edoardo Debenedetti, Nicolas Flammarion, Mung Chiang, Prateek Mittal, and Matthias Hein · 2020
Earlier work this paper cites.
Modality-agnostic attention fusion for visual search with text feedback
Eric Dodds, Jack Culpepper, Simao Herdade, Yang Zhang, and Kofi Boakye · 2020
Cited alongside, same era.
A closer look at the robustness of vision-and-language pre-trained models
Linjie Li, Zhe Gan, and Jingjing Liu · 2020
Cited alongside, same era.
Cosmo: Content-style modulation for image retrieval with text feedback
Seungmin Lee, Dongwan Kim, and Bohyung Han · 2021
Cited alongside, same era.
Image retrieval on real-life images with pre-trained vision-and-language models
Zheyuan Liu, Cristian Rodriguez-Opazo, Damien Teney, and Stephen Gould · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Uigr: unified interactive garment retrieval
Xiao Han, Sen He, Li Zhang, Yi-Zhe Song, and Tao Xiang · 2022
Later among the works it cites.
Fashionvil: Fashion-focused vision-and-language representation learning
Xiao Han, Licheng Yu, Xiatian Zhu, Li Zhang, Yi-Zhe Song, and Tao Xiang · 2022
Later among the works it cites.
Probabilistic compositional embeddings for multimodal image retrieval
Andrei Neculai, Yanbei Chen, and Zeynep Akata · 2022
Later among the works it cites.
Vision transformers are robust learners
Sayak Paul and Pin-Yu Chen · 2022
Later among the works it cites.
Robustlr: A diagnostic benchmark for evaluating logical robustness of deductive reasoners
Soumya Sanyal, Zeyi Liao, and Xiang Ren · 2022
Later among the works it cites.
Winoground: Probing vision and language models for visio-linguistic compositionality
Tristan Thrush, Ryan Jiang, Max Bartolo, Amanpreet Singh, Adina Williams, Douwe Kiela, and Candace Ross · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models, 2021
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer · 2021
Cited alongside, same era.
Adversarial glue: A multi-task benchmark for robustness evaluation of language models
Boxin Wang, Chejian Xu, Shuohang Wang, Zhe Gan, Yu Cheng, Jianfeng Gao, Ahmed Hassan Awadallah, and Bo Li · 2021
Cited alongside, same era.
Textflint: Unified multilingual robustness evaluation toolkit for natural language processing
Xiao Wang, Qin Liu, Tao Gui, Qi Zhang, Yicheng Zou, Xin Zhou, Jiacheng Ye, Yongxin Zhang, Rui Zheng, Zexiong Pang, et al · 2021
Cited alongside, same era.
Measure and improve robustness in nlp models: A survey
Xuezhi Wang, Haohan Wang, and Diyi Yang · 2021
Cited alongside, same era.
The fashion iq dataset: Retrieving images by combining side information and relative natural language feedback
Hui Wu, Yupeng Gao, Xiaoxiao Guo, Ziad Al-Halah, Steven Rennie, Kristen Grauman, and Rogerio Feris · 2021
Cited alongside, same era.
Certified robustness to text adversarial attacks by randomized [mask]
Jiehang Zeng, Xiaoqing Zheng, Jianhan Xu, Linyang Li, Liping Yuan, and Xuanjing Huang · 2021
Cited alongside, same era.
Effective conditioned and composed image retrieval combining clip-based features
Alberto Baldrati, Marco Bertini, Tiberio Uricchio, and Alberto Del Bimbo · 2022
Cited alongside, same era.
Later among the works it cites.
When and why vision-language models behave like bags-of-words, and what to do about it?
Mert Yuksekgonul, Federico Bianchi, Pratyusha Kalluri, Dan Jurafsky, and James Zou · 2022
Later among the works it cites.
A benchmark for compositional visual reasoning
Aimen Zerroug, Mohit Vaishnav, Julien Colin, Sebastian Musslick, and Thomas Serre · 2022
Later among the works it cites.
Compodiff: Versatile composed image retrieval with latent diffusion
Geonmo Gu, Sanghyuk Chun, Wonjae Kim, HeeJae Jun, Yoohoon Kang, and Sangdoo Yun · 2023
Closest in time.
Fame-vil: Multi-tasking vision-language model for heterogeneous fashion tasks
Xiao Han, Xiatian Zhu, Licheng Yu, Li Zhang, Yi-Zhe Song, and Tao Xiang · 2023
Closest in time.
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C. Berg, Wan-Yen Lo, Piotr Dollár, and Ross Girshick · 2023
Closest in time.
Data roaming and early fusion for composed image retrieval
Matan Levy, Rami Ben-Ari, Nir Darshan, and Dani Lischinski · 2023
Closest in time.
Grounding dino: Marrying dino with grounded pre-training for open-set object detection
Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, Hao Zhang, Jie Yang, Chunyuan Li, Jianwei Yang, Hang Su, Jun Zhu, et al · 2023
Closest in time.
Crepe: Can vision-language foundation models reason compositionally?
Zixian Ma, Jerry Hong, Mustafa Omer Gul, Mona Gandhi, Irena Gao, and Ranjay Krishna · 2023
Closest in time.
Pic2word: Mapping pictures to words for zero-shot composed image retrieval
Kuniaki Saito, Kihyuk Sohn, Xiang Zhang, Chun-Liang Li, Chen-Yu Lee, Kate Saenko, and Tomas Pfister · 2023
Closest in time.
Visual chatgpt: Talking, drawing and editing with visual foundation models
Chenfei Wu, Shengming Yin, Weizhen Qi, Xiaodong Wang, Zecheng Tang, and Nan Duan · 2023
Closest in time.