Fetching the paper…
Reading the bibliography…
An intuitive way to search for images is to use queries composed of an example image and a complementary text.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Object retrieval with large vocabularies and fast spatial matching
James Philbin, Ondřej Chum, Michael Isard, Josef Sivic, and Andrew Zisserman · 2007
Earlier work this paper cites.
Automatic attribute discovery and characterization from noisy web data
Tamara L Berg, Alexander C Berg, and Jonathan Shih · 2010
Earlier work this paper cites.
Large-scale image retrieval with compressed Fisher vectors
Florent Perronnin, Yan Liu, Jorge Sánchez, and Hervé Poirier · 2010
Earlier work this paper cites.
Three things everyone should know to improve object retrieval
Relja Arandjelović and Andrew Zisserman · 2012
Earlier work this paper cites.
Devise: A deep visual-semantic embedding model
Andrea Frome, Greg Corrado, Jonathon Shlens, Samy Bengio, Jeffrey Dean, Marc’Aurelio Ranzato, and Tomas Mikolov · 2013
Earlier work this paper cites.
Learning phrase representations using rnn encoder–decoder for statistical machine translation
Kyunghyun Cho, Bart van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D Manning · 2014
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al · 2015
Earlier work this paper cites.
Show, Attend and Tell: Neural Image Caption Generation with Visual Attention
Kelvin Xu, Jimmy Ba, Jamie R. Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhudinov, Rich Zemel, and Yoshua Bengio · 2015
Earlier work this paper cites.
Deep image retrieval: Learning global representations for image search
Albert Gordo, Jon Almazán, Jerome Revaud, and Diane Larlus · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Multimodal residual learning for visual QA
Jin-Hwa Kim, Sang-Woo Lee, Donghyun Kwak, Min-Oh Heo, Jeonghee Kim, Jung-Woo Ha, and Byoung-Tak Zhang · 2016
Earlier work this paper cites.
Hierarchical question-image co-attention for visual question answering
Jiasen Lu, Jianwei Yang, Dhruv Batra, and Devi Parikh · 2016
Earlier work this paper cites.
CNN image retrieval learns from BoW: Unsupervised fine-tuning with hard examples
Filip Radenović, Giorgos Tolias, and Ondřej Chum · 2016
Earlier work this paper cites.
Image search with selective match kernels: aggregation across single and multiple images
Giorgos Tolias, Yannis Avrithis, and Hervé Jégou · 2016
Earlier work this paper cites.
Learning deep structure-preserving image-text embeddings
Liwei Wang, Yin Li, and Svetlana Lazebnik · 2016
Earlier work this paper cites.
Automatic spatially-aware fashion concept discovery
Xintong Han, Zuxuan Wu, Phoenix X Huang, Xiao Zhang, Menglong Zhu, Yuan Li, Yang Zhao, and Larry S Davis · 2017
Earlier work this paper cites.
Person Search with Natural Language Description
Shuang Li, Tong Xiao, Hongsheng Li, Bolei Zhou, Dayu Yue, and Xiaogang Wang · 2017
Cited alongside, same era.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2017
Cited alongside, same era.
Large-scale image retrieval with attentive deep local features
Hyeonwoo Noh, Andre Araujo, Jack Sim, Tobias Weyand, and Bohyung Han · 2017
Cited alongside, same era.
A simple neural network module for relational reasoning
Adam Santoro, David Raposo, David G Barrett, Mateusz Malinowski, Razvan Pascanu, Peter Battaglia, and Timothy Lillicrap · 2017
Cited alongside, same era.
Grad-cam: Visual explanations from deep networks via gradient-based localization
Ramprasaath R Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Batra · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Transferable Attention for Domain Adaptation
Ximei Wang, Liang Li, Weirui Ye, Mingsheng Long, and Jianmin Wang · 2019
Later among the works it cites.
Unifying deep local and global features for image search
Bingyi Cao, Andre Araujo, and Jack Sim · 2020
Later among the works it cites.
Learning joint visual semantic matching embeddings for language-guided retrieval
Yanbei Chen and Loris Bazzani · 2020
Later among the works it cites.
Image search with text feedback by visiolinguistic attention learning
Yanbei Chen, Shaogang Gong, and Loris Bazzani · 2020
Later among the works it cites.
Fashion iq challenge
Yupeng Gao and Xiaoxiao Guo · 2020
Later among the works it cites.
Composed query image retrieval using locally bounded features
Mehrdad Hosseinzadeh and Yang Wang · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Deep attention neural tensor network for visual question answering
Yalong Bai, Jianlong Fu, Tiejun Zhao, and Tao Mei · 2018
Cited alongside, same era.
VSE++: Improving visual-semantic embeddings with hard negatives
Fartash Faghri, David J Fleet, Jamie Ryan Kiros, and Sanja Fidler · 2018
Cited alongside, same era.
Dialog-based interactive image retrieval
Xiaoxiao Guo, Hui Wu, Yu Cheng, Steven Rennie, Gerald Tesauro, and Rogerio Feris · 2018
Cited alongside, same era.
Rethinking floating point for deep learning
Jeff Johnson · 2018
Cited alongside, same era.
Deep Adversarial Attention Alignment for Unsupervised Domain Adaptation: The Benefit of Target Expectation Maximization
Guoliang Kang, Liang Zheng, Yan Yan, and Yi Yang · 2018
Cited alongside, same era.
Stacked cross attention for image-text matching
Kuang-Huei Lee, Xi Chen, Gang Hua, Houdong Hu, and Xiaodong He · 2018
Cited alongside, same era.
Oscar: Object-semantics aligned pre-training for vision-language tasks
Xiujun Li, Xi Yin, Chunyuan Li, Pengchuan Zhang, Xiaowei Hu, Lei Zhang, Lijuan Wang, Houdong Hu, Li Dong, Furu Wei, et al · 2020
Later among the works it cites.
A metric learning reality check
Kevin Musgrave, Serge Belongie, and Ser-Nam Lim · 2020
Later among the works it cites.
Learning and aggregating deep local descriptors for instance-level recognition
Giorgos Tolias, Tomas Jenicek, and Ondřej Chum · 2020
Later among the works it cites.
Compositional learning of image-text query for image retrieval
Muhammad Umer Anwaar, Egor Labintcev, and Martin Kleinsteuber · 2021
Later among the works it cites.
Leveraging style and content features for text conditioned image retrieval
Pranit Chawla, Surgan Jandial, Pinkesh Badjatiya, Ayush Chopra, Mausoom Sarkar, and Balaji Krishnamurthy · 2021
Later among the works it cites.
Probabilistic embeddings for cross-modal retrieval
Sanghyuk Chun, Seong Joon Oh, Rafael S Rezende, Yannis Kalantidis, and Diane Larlus · 2021
Later among the works it cites.
Mostafa Dehghani, Anurag Arnab, Lucas Beyer, Ashish Vaswani, and Yi Tay · 2021
Later among the works it cites.
Cosmo: Content-style modulation for image retrieval with text feedback
Seungmin Lee, Dongwan Kim, and Bohyung Han · 2021
Later among the works it cites.
Image retrieval on real-life images with pre-trained vision-and-language models
Zheyuan Liu, Cristian Rodriguez-Opazo, Damien Teney, and Stephen Gould · 2021
Later among the works it cites.
Thinking fast and slow: Efficient text-to-visual retrieval with transformers
Antoine Miech, Jean-Baptiste Alayrac, Ivan Laptev, Josef Sivic, and Andrew Zisserman · 2021
Later among the works it cites.
Fashion IQ: A new dataset towards retrieving images by natural language feedback
Hui Wu, Yupeng Gao, Xiaoxiao Guo, Ziad Al-Halah, Steven Rennie, Kristen Grauman, and Rogerio Feris · 2021
Later among the works it cites.