Fetching the paper…
Reading the bibliography…
Despite the evolution of deep-learning-based visual-textual processing systems, precise multi-modal matching remains a challenging task.
Rouge: A package for automatic evaluation of summaries. In Text summarization branches out . 74–81
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Efficient Estimation of Word Representations in Vector Space. In 1st International Conference on Learning Representations, ICLR 2013
Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013 · 2013
Earlier work this paper cites.
Microsoft coco: Common objects in context. In European conference on computer vision . Springer, 740–755
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. 2014 · 2014
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Peter Young, Alice Lai, Micah Hodosh, and Julia Hockenmaier. 2014 · 2014
Earlier work this paper cites.
Deep visual-semantic alignments for generating image descriptions. In Proc. of the IEEE conference on computer vision and pattern recognition . 3128–3137
Andrej Karpathy and Li Fei-Fei. 2015 · 2015
Earlier work this paper cites.
Associating neural word embeddings with deep image representations using fisher vectors. In Proc. of the IEEE Conference on Computer Vision and Pattern Recognition . 4437–4446
Benjamin Klein, Guy Lev, Gil Sadeh, and Lior Wolf. 2015 · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks. In Advances in neural information processing systems . 91–99
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. 2015 · 2015
Earlier work this paper cites.
Spice: Semantic propositional image caption evaluation. In European Conference on Computer Vision . Springer, 382–398
Peter Anderson, Basura Fernando, Mark Johnson, and Stephen Gould. 2016 · 2016
Earlier work this paper cites.
Leveraging visual question answering for image-caption ranking. In European Conference on Computer Vision . Springer, 261–277
Xiao Lin and Devi Parikh. 2016 · 2016
Earlier work this paper cites.
Order-Embeddings of Images and Language. In 4th International Conference on Learning Representations, ICLR
Ivan Vendrov, Ryan Kiros, Sanja Fidler, and Raquel Urtasun. 2016 · 2016
Earlier work this paper cites.
Linking image and text with 2-way nets. In Proc. of the IEEE conference on computer vision and pattern recognition . 4601–4611
Aviv Eisenschtat and Lior Wolf. 2017 · 2017
Earlier work this paper cites.
Learning to reason: End-to-end module networks for visual question answering. In Proc. of the IEEE International Conference on Computer Vision . 804–813
Ronghang Hu, Jacob Andreas, Marcus Rohrbach, Trevor Darrell, and Kate Saenko. 2017 · 2017
Earlier work this paper cites.
Instance-aware image and sentence matching with selective multimodal lstm. In Proc. of the IEEE Conference on Computer Vision and Pattern Recognition . 2310–2318
Yan Huang, Wei Wang, and Liang Wang. 2017 · 2017
Earlier work this paper cites.
Inferring and executing programs for visual reasoning. In Proc. of the IEEE International Conference on Computer Vision . 2989–2998
Justin Johnson, Bharath Hariharan, Laurens Van Der Maaten, Judy Hoffman, Li Fei-Fei, C Lawrence Zitnick, and Ross Girshick. 2017 · 2017
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, et al · 2017
Earlier work this paper cites.
Learning a recurrent residual fusion network for multimodal matching. In Proc. of the IEEE International Conference on Computer Vision . 4107–4116
Yu Liu, Yanming Guo, Erwin M Bakker, and Michael S Lew. 2017 · 2017
Earlier work this paper cites.
Self-critical sequence training for image captioning. In Proc. of the IEEE Conference on Computer Vision and Pattern Recognition . 7008–7024
Steven J Rennie, Etienne Marcheret, Youssef Mroueh, Jerret Ross, and Vaibhava Goel. 2017 · 2017
Earlier work this paper cites.
A simple neural network module for relational reasoning
Adam Santoro, David Raposo, David G Barrett, Mateusz Malinowski, Razvan Pascanu, Peter Battaglia, and Timothy Lillicrap. 2017 · 2017
Earlier work this paper cites.
Graph-structured representations for visual question answering. In Proc. of the IEEE conference on computer vision and pattern recognition . 1–9
Damien Teney, Lingqiao Liu, and Anton van Den Hengel. 2017 · 2017
Earlier work this paper cites.
Attention is all you need. In Advances in neural information processing systems . 5998–6008
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Bottom-up and top-down attention for image captioning and visual question answering. In Proc. of the IEEE conference on computer vision and pattern recognition . 6077–6086
Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang. 2018 · 2018
Earlier work this paper cites.
Picture it in your mind: Generating high level visual representations from textual descriptions
Fabio Carrara, Andrea Esuli, Tiziano Fagni, Fabrizio Falchi, and Alejandro Moreo. 2018 · 2018
Earlier work this paper cites.
VSE++: Improving Visual-Semantic Embeddings with Hard Negatives. In BMVC 2018 . BMVA Press, 12
Fartash Faghri, David J. Fleet, Jamie Ryan Kiros, and Sanja Fidler. 2018 · 2018
Cited alongside, same era.
Look, imagine and match: Improving textual-visual cross-modal retrieval with generative models. In Proc. of the IEEE Conf. on Computer Vision and Pattern Recognition . 7181–7189
Jiuxiang Gu, Jianfei Cai, Shafiq R Joty, Li Niu, and Gang Wang. 2018 · 2018
Cited alongside, same era.
Image and sentence matching via semantic concepts and order learning
Yan Huang, Qi Wu, Wei Wang, and Liang Wang. 2018b · 2018
Cited alongside, same era.
Multi-task learning using uncertainty to weigh losses for scene geometry and semantics. In Proc. of the IEEE conference on computer vision and pattern recognition . 7482–7491
Alex Kendall, Yarin Gal, and Roberto Cipolla. 2018 · 2018
Cited alongside, same era.
Stacked cross attention for image-text matching. In Proc. of the European Conference on Computer Vision (ECCV) . 201–216
Learning visual features for relational CBIR
Nicola Messina, Giuseppe Amato, Fabio Carrara, Fabrizio Falchi, and Claudio Gennaro. 2019 · 2019
Later among the works it cites.
Adversarial representation learning for text-to-image matching. In Proc. of the IEEE International Conference on Computer Vision . 5814–5824
Nikolaos Sarafianos, Xiang Xu, and Ioannis A Kakadiaris. 2019 · 2019
Later among the works it cites.
Position focused attention network for image-text matching
Yaxiong Wang, Hao Yang, Xueming Qian, Lin Ma, Jing Lu, Biao Li, and Xin Fan. 2019 · 2019
Later among the works it cites.
Auto-encoding scene graphs for image captioning. In Proc. of the IEEE Conference on Computer Vision and Pattern Recognition . 10685–10694
Xu Yang, Kaihua Tang, Hanwang Zhang, and Jianfei Cai. 2019 · 2019
Later among the works it cites.
IMRAM: Iterative Matching with Recurrent Attention Memory for Cross-Modal Image-Text Retrieval. In Proc. of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 12655–12663
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kuang-Huei Lee, Xi Chen, Gang Hua, Houdong Hu, and Xiaodong He. 2018 · 2018
Cited alongside, same era.
Factorizable net: an efficient subgraph-based framework for scene graph generation. In Proc. of the European Conference on Computer Vision (ECCV) . 335–351
Yikang Li, Wanli Ouyang, Bolei Zhou, Jianping Shi, Chao Zhang, and Xiaogang Wang. 2018 · 2018
Cited alongside, same era.
Learning relationship-aware visual features. In Proc. of the European Conference on Computer Vision (ECCV) . 0–0
Nicola Messina, Giuseppe Amato, Fabio Carrara, Fabrizio Falchi, and Claudio Gennaro. 2018 · 2018
Cited alongside, same era.
Joint global and co-attentive representation learning for image-sentence retrieval. In Proc. of the 26th ACM international conference on Multimedia . 1398–1406
Shuhui Wang, Yangyu Chen, Junbao Zhuo, Qingming Huang, and Qi Tian. 2018 · 2018
Cited alongside, same era.
Graph r-cnn for scene graph generation. In Proc. of the European conference on computer vision (ECCV) . 670–685
Jianwei Yang, Jiasen Lu, Stefan Lee, Dhruv Batra, and Devi Parikh. 2018 · 2018
Cited alongside, same era.
Exploring visual relationship for image captioning. In Proc. of the European conference on computer vision (ECCV) . 684–699
Ting Yao, Yingwei Pan, Yehao Li, and Tao Mei. 2018 · 2018
Cited alongside, same era.
Uniter: Learning universal image-text representations
Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu. 2019 · 2019
Cited alongside, same era.
Show, control and tell: A framework for generating controllable and grounded captions. In Proc. of the IEEE Conf. on Computer Vision and Pattern Recognition . 8307–8316
Marcella Cornia, Lorenzo Baraldi, and Rita Cucchiara. 2019 · 2019
Cited alongside, same era.
Hui Chen, Guiguang Ding, Xudong Liu, Zijia Lin, Ji Liu, and Jungong Han. 2020 · 2020
Closest in time.
Associating Images with Sentences Using Recurrent Canonical Correlation Analysis
Yawen Guo, Hui Yuan, and Kun Zhang. 2020 · 2020
Closest in time.
Bi-directional spatial-semantic attention networks for image-text matching
Feiran Huang, Xiaoming Zhang, Zhonghua Zhao, and Zhoujun Li. 2018c · 2020
Closest in time.
Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers
Zhicheng Huang, Zhaoyang Zeng, Bei Liu, Dongmei Fu, and Jianlong Fu. 2020 · 2020
Closest in time.
Multi-Modal Memory Enhancement Attention Network for Image-Text Matching
Zhong Ji, Zhigang Lin, Haoran Wang, and Yuqing He. 2020a · 2020
Closest in time.
SMAN: Stacked Multimodal Attention Network for Cross-Modal Image-Text Retrieval
Zhong Ji, Haoran Wang, Jungong Han, and Yanwei Pang. 2020b · 2020
Closest in time.
Graph Structured Network for Image-Text Matching. In Proc. of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 10921–10930
Chunxiao Liu, Zhendong Mao, Tianzhu Zhang, Hongtao Xie, Bin Wang, and Yongdong Zhang. 2020 · 2020
Closest in time.
Efficient Document Re-Ranking for Transformers by Precomputing Term Representations
Sean MacAvaney, Franco Maria Nardini, Raffaele Perego, Nicola Tonellotto, Nazli Goharian, and Ophir Frieder. 2020 · 2020
Closest in time.
Transformer Reasoning Network for Image-Text Matching and Retrieval. In International Conference on Pattern Recognition (ICPR) 2020 (Accepted)
Nicola Messina, Fabrizio Falchi, Andrea Esuli, and Giuseppe Amato. 2020 · 2020
Closest in time.
Imagebert: Cross-modal pre-training with large-scale weak-supervised image-text data
Di Qi, Lin Su, Jia Song, Edward Cui, Taroon Bharti, and Arun Sacheti. 2020 · 2020
Closest in time.
Context-Aware Multi-View Summarization Network for Image-Text Matching. In Proc. of the 28th ACM International Conference on Multimedia . 1047–1055
Leigang Qu, Meng Liu, Da Cao, Liqiang Nie, and Qi Tian. 2020 · 2020
Closest in time.
VL-BERT: Pre-training of Generic Visual-Linguistic Representations. In International Conference on Learning Representations
Weijie Su, Xizhou Zhu, Yue Cao, Bin Li, Lewei Lu, Furu Wei, and Jifeng Dai. 2020 · 2020
Closest in time.
Adversarial Attentive Multi-modal Embedding Learning for Image-Text Matching
Kaimin Wei and Zhibo Zhou. 2020 · 2020
Closest in time.
Multi-Modality Cross Attention Network for Image and Sentence Matching. In Proc. of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 10941–10950
Xi Wei, Tianzhu Zhang, Yan Li, Yongdong Zhang, and Feng Wu. 2020 · 2020
Closest in time.
Cross-modal attention with semantic consistence for image-text matching
Xing Xu, Tan Wang, Yang Yang, Lin Zuo, Fumin Shen, and Heng Tao Shen. 2020 · 2020
Closest in time.
Unified Vision-Language Pre-Training for Image Captioning and VQA.. In AAAI . 13041–13049
Luowei Zhou, Hamid Palangi, Lei Zhang, Houdong Hu, Jason J Corso, and Jianfeng Gao. 2020 · 2020
Closest in time.
Learning fragment self-attention embeddings for image-text matching. In Proc. of the 27th ACM International Conference on Multimedia . 2088–2096
Yiling Wu, Shuhui Wang, Guoli Song, and Qingming Huang. 2019 · 2096
Closest in time.