Fetching the paper…
Reading the bibliography…
Visual Semantic Embedding (VSE) is a dominant approach for vision-language retrieval, which aims at learning a deep embedding space such that visual data are embedded close to their semantic text labels or descriptions.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Fspool: Learning set representations with featurewise sort pooling
Y. Zhang, Jonathon S. Hare, and A. Prügel-Bennett · 2011
Earlier work this paper cites.
Devise: A deep visual-semantic embedding model
Andrea Frome, Greg S Corrado, Jon Shlens, Samy Bengio, Jeff Dean, Marc’Aurelio Ranzato, and Tomas Mikolov · 2013
Earlier work this paper cites.
Attribute-based classification for zero-shot visual object categorization
Christoph H Lampert, Hannes Nickisch, and Stefan Harmeling · 2013
Earlier work this paper cites.
A convolutional neural network for modelling sentences
Nal Kalchbrenner, Edward Grefenstette, and Phil Blunsom · 2014
Earlier work this paper cites.
Unifying visual-semantic embeddings with multimodal neural language models
Ryan Kiros, Ruslan Salakhutdinov, and Richard S Zemel · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge J. Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick · 2014
Earlier work this paper cites.
Zero-shot learning by convex combination of semantic embeddings
Mohammad Norouzi, Tomas Mikolov, Samy Bengio, Yoram Singer, Jonathon Shlens, Andrea Frome, Greg S Corrado, and Jeffrey Dean · 2014
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael S. Bernstein, Alexander C. Berg, and Fei-Fei Li · 2014
Earlier work this paper cites.
Grounded compositional semantics for finding and describing images with sentences
Richard Socher, Andrej Karpathy, Quoc V. Le, Christopher D. Manning, and Andrew Y. Ng · 2014
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Peter Young, Alice Lai, Micah Hodosh, and Julia Hockenmaier · 2014
Earlier work this paper cites.
Microsoft coco captions: Data collection and evaluation server
Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Dollár, and C Lawrence Zitnick · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Spatial pyramid pooling in deep convolutional networks for visual recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Deep visual-semantic alignments for generating image descriptions
Andrej Karpathy and Li Fei-Fei · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun · 2015
Earlier work this paper cites.
Synthesized classifiers for zero-shot learning
Soravit Changpinyo, Wei-Lun Chao, Boqing Gong, and Fei Sha · 2016
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, J. M. Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A. Shamma, Michael S. Bernstein, and Li Fei-Fei · 2016
Earlier work this paper cites.
Aggregated residual transformations for deep neural networks
Saining Xie, Ross B. Girshick, Piotr Dollár, Zhuowen Tu, and Kaiming He · 2016
Cited alongside, same era.
Msr-vtt: A large video description dataset for bridging video and language
J. Xu, T. Mei, Ting Yao, and Y. Rui · 2016
Cited alongside, same era.
Linking image and text with 2-way nets
Aviv Eisenschtat and Lior Wolf · 2017
Cited alongside, same era.
Vse++: Improved visual-semantic embeddings
Fartash Faghri, David J Fleet, Jamie Ryan Kiros, and Sanja Fidler · 2017
Cited alongside, same era.
Dual attention networks for multimodal reasoning and matching
Hyeonseob Nam, Jung-Woo Ha, and Jeonghee Kim · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Visual semantic reasoning for image-text matching
Kunpeng Li, Yulun Zhang, Kai Li, Yuanyuan Li, and Yun Fu · 2019
Later among the works it cites.
Visualbert: A simple and performant baseline for vision and language
Liunian Harold Li, Mark Yatskar, Da Yin, Cho-Jui Hsieh, and Kai-Wei Chang · 2019
Later among the works it cites.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee · 2019
Later among the works it cites.
Fine-tuning cnn image retrieval with no human annotation
Filip Radenović, Giorgos Tolias, and O. Chum · 2019
Later among the works it cites.
Polysemous visual-semantic embedding for cross-modal retrieval
Yale Song and Mohammad Soleymani · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bottom-up and top-down attention for image captioning and visual question answering
Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang · 2018
Cited alongside, same era.
Look, imagine and match: Improving textual-visual cross-modal retrieval with generative models
Jiuxiang Gu, Jianfei Cai, Shafiq R Joty, Li Niu, and Gang Wang · 2018
Cited alongside, same era.
Learning semantic concepts and order for image and sentence matching
Yan Huang, Qi Wu, Chunfeng Song, and Liang Wang · 2018
Cited alongside, same era.
Stacked cross attention for image-text matching
Kuang-Huei Lee, Xi Chen, Gang Hua, Houdong Hu, and Xiaodong He · 2018
Cited alongside, same era.
Exploring the limits of weakly supervised pretraining
Dhruv Kumar Mahajan, Ross B. Girshick, Vignesh Ramanathan, Kaiming He, Manohar Paluri, Yixuan Li, Ashwin Bharambe, and Laurens van der Maaten · 2018
Cited alongside, same era.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut · 2018
Cited alongside, same era.
Xin Wang, Jiawei Wu, Junkun Chen, Lei Li, Y. Wang, and William Yang Wang · 2019
Later among the works it cites.
Language-agnostic visual-semantic embeddings
Jonatas Wehrmann, Mauricio A. Lopes, Douglas M. Souza, and Rodrigo C. Barros · 2019
Later among the works it cites.
Unified visual-semantic embeddings: Bridging vision and language with structured meaning representations
Hao Wu, Jiayuan Mao, Yufeng Zhang, Yuning Jiang, Lei Li, Weiwei Sun, and Wei-Ying Ma · 2019
Later among the works it cites.
Imram: Iterative matching with recurrent attention memory for cross-modal image-text retrieval
H. Chen, G. Ding, Xudong Liu, Zijia Lin, J. Liu, and J. Han · 2020
Closest in time.
Fine-grained video-text retrieval with hierarchical graph reasoning
Shizhe Chen, Yida Zhao, Qin Jin, and Qi Wu · 2020
Closest in time.
Uniter: Learning universal image-text representations
Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu · 2020
Closest in time.
In defense of grid features for visual question answering
Huaizu Jiang, Ishan Misra, Marcus Rohrbach, Erik G. Learned-Miller, and Xinlei Chen · 2020
Closest in time.
Oscar: Object-semantics aligned pre-training for vision-language tasks
Xiujun Li, Xi Yin, C. Li, X. Hu, Pengchuan Zhang, Lei Zhang, Longguang Wang, H. Hu, Li Dong, Furu Wei, Yejin Choi, and Jianfeng Gao · 2020
Closest in time.
Crisscrossed captions: Extended intramodal and intermodal semantic similarity judgments for ms-coco
Zarana Parekh, Jason Baldridge, Daniel Cer, Austin Waters, and Yinfei Yang · 2020
Closest in time.
Consensus-aware visual-semantic embedding for image-text matching
Haoran Wang, Ying Zhang, Zhong Ji, Yanwei Pang, and Lin Ma · 2020
Closest in time.
Learning to represent image and text with denotation graph
Bowen Zhang, Hexiang Hu, Vihan Jain, Eugene Ie, and Fei Sha · 2020
Closest in time.
Context-aware attention network for image-text retrieval
Qi Zhang, Zhen Lei, Zhaoxiang Zhang, and S. Li · 2020
Closest in time.