Fetching the paper…
Reading the bibliography…
Image-text matching plays a critical role in bridging the vision and language, and great progress has been made by exploiting the global alignment between image and sentence, or local alignments between regions and words.
Neighbourhood Watch: Referring Expression Comprehension via Language-Guided Graph Attention Networks
Wang, P.; Wu, Q.; Cao, J.; Shen, C.; Gao, L.; and van den Hengel, A. 2019a · 1968
Earlier work this paper cites.
Neighbourhood Watch: Referring Expression Comprehension via Language-Guided Graph Attention Networks
Wang, P.; Wu, Q.; Cao, J.; Shen, C.; Gao, L.; and van den Hengel, A. 2019a · 1968
Earlier work this paper cites.
Polysemous Visual-Semantic Embedding for Cross-Modal Retrieval
Song, Y.; and Soleymani, M. 2019 · 1988
Earlier work this paper cites.
Polysemous Visual-Semantic Embedding for Cross-Modal Retrieval
Song, Y.; and Soleymani, M. 2019 · 1988
Earlier work this paper cites.
Bidirectional recurrent neural networks
Schuster, M.; and Paliwal, K. K. 1997 · 1997
Earlier work this paper cites.
Bidirectional recurrent neural networks
Schuster, M.; and Paliwal, K. K. 1997 · 1997
Earlier work this paper cites.
Fisher Kernels on Visual Vocabularies for Image Categorization
Perronnin, F.; and Dance, C. R. 2007 · 2007
Earlier work this paper cites.
Fisher Kernels on Visual Vocabularies for Image Categorization
Perronnin, F.; and Dance, C. R. 2007 · 2007
Earlier work this paper cites.
DeViSE: A Deep Visual-Semantic Embedding Model
Frome, A.; Corrado, G. S.; Shlens, J.; Bengio, S.; Dean, J.; Ranzato, M.; and Mikolov, T. 2013 · 2013
Earlier work this paper cites.
Efficient Estimation of Word Representations in Vector Space
Mikolov, T.; Chen, K.; Corrado, G.; and Dean, J. 2013 · 2013
Earlier work this paper cites.
DeViSE: A Deep Visual-Semantic Embedding Model
Frome, A.; Corrado, G. S.; Shlens, J.; Bengio, S.; Dean, J.; Ranzato, M.; and Mikolov, T. 2013 · 2013
Earlier work this paper cites.
Efficient Estimation of Word Representations in Vector Space
Mikolov, T.; Chen, K.; Corrado, G.; and Dean, J. 2013 · 2013
Earlier work this paper cites.
Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models
Kiros, R.; Salakhutdinov, R.; and Zemel, R. S. 2014 · 2014
Earlier work this paper cites.
Microsoft COCO: Common Objects in Context
Lin, T.; Maire, M.; Belongie, S. J.; Hays, J.; Perona, P.; Ramanan, D.; Dollár, P.; and Zitnick, C. L. 2014 · 2014
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Young, P.; Lai, A.; Hodosh, M.; and Hockenmaier, J. 2014 · 2014
Earlier work this paper cites.
Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models
Kiros, R.; Salakhutdinov, R.; and Zemel, R. S. 2014 · 2014
Earlier work this paper cites.
Microsoft COCO: Common Objects in Context
Lin, T.; Maire, M.; Belongie, S. J.; Hays, J.; Perona, P.; Ramanan, D.; Dollár, P.; and Zitnick, C. L. 2014 · 2014
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Young, P.; Lai, A.; Hodosh, M.; and Hockenmaier, J. 2014 · 2014
Earlier work this paper cites.
Convolutional Networks on Graphs for Learning Molecular Fingerprints
Duvenaud, D.; Maclaurin, D.; Aguilera-Iparraguirre, J.; Gómez-Bombarelli, R.; Hirzel, T.; Aspuru-Guzik, A.; and Adams, R. P. 2015 · 2015
Earlier work this paper cites.
Deep visual-semantic alignments for generating image descriptions
Karpathy, A.; and Li, F. 2015 · 2015
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization
Kingma, D. P.; and Ba, J. 2015 · 2015
Earlier work this paper cites.
Associating neural word embeddings with deep image representations using Fisher Vectors
Klein, B.; Lev, G.; Sadeh, G.; and Wolf, L. 2015 · 2015
Earlier work this paper cites.
Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks
Ren, S.; He, K.; Girshick, R. B.; and Sun, J. 2015 · 2015
Earlier work this paper cites.
Convolutional Networks on Graphs for Learning Molecular Fingerprints
Duvenaud, D.; Maclaurin, D.; Aguilera-Iparraguirre, J.; Gómez-Bombarelli, R.; Hirzel, T.; Aspuru-Guzik, A.; and Adams, R. P. 2015 · 2015
Earlier work this paper cites.
Deep visual-semantic alignments for generating image descriptions
Karpathy, A.; and Li, F. 2015 · 2015
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization
Kingma, D. P.; and Ba, J. 2015 · 2015
Earlier work this paper cites.
Associating neural word embeddings with deep image representations using Fisher Vectors
Klein, B.; Lev, G.; Sadeh, G.; and Wolf, L. 2015 · 2015
Earlier work this paper cites.
Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks
Ren, S.; He, K.; Girshick, R. B.; and Sun, J. 2015 · 2015
Earlier work this paper cites.
Gated Graph Sequence Neural Networks
Li, Y.; Tarlow, D.; Brockschmidt, M.; and Zemel, R. S. 2016 · 2016
Earlier work this paper cites.
Order-Embeddings of Images and Language
Vendrov, I.; Kiros, R.; Fidler, S.; and Urtasun, R. 2016 · 2016
Earlier work this paper cites.
Learning Deep Structure-Preserving Image-Text Embeddings
Wang, L.; Li, Y.; and Lazebnik, S. 2016 · 2016
Cited alongside, same era.
Gated Graph Sequence Neural Networks
Li, Y.; Tarlow, D.; Brockschmidt, M.; and Zemel, R. S. 2016 · 2016
Cited alongside, same era.
Order-Embeddings of Images and Language
Vendrov, I.; Kiros, R.; Fidler, S.; and Urtasun, R. 2016 · 2016
Cited alongside, same era.
Learning Deep Structure-Preserving Image-Text Embeddings
Wang, L.; Li, Y.; and Lazebnik, S. 2016 · 2016
Cited alongside, same era.
VSE++: Improved Visual-Semantic Embeddings
Faghri, F.; Fleet, D. J.; Kiros, R.; and Fidler, S. 2017 · 2017
Cited alongside, same era.
Semi-Supervised Classification with Graph Convolutional Networks
Kipf, T. N.; and Welling, M. 2017 · 2017
Cited alongside, same era.
Learning Semantic Concepts and Order for Image and Sentence Matching
Huang, Y.; Wu, Q.; Song, C.; and Wang, L. 2018 · 2018
Later among the works it cites.
Stacked Cross Attention for Image-Text Matching
Lee, K.; Chen, X.; Hua, G.; Hu, H.; and He, X. 2018 · 2018
Later among the works it cites.
AttnGAN: Fine-Grained Text to Image Generation With Attentional Generative Adversarial Networks
Xu, T.; Zhang, P.; Huang, Q.; Zhang, H.; Gan, Z.; Huang, X.; and He, X. 2018 · 2018
Later among the works it cites.
Multi-Level Visual-Semantic Alignments with Relation-Wise Dual Attention Network for Image and Text Matching
Hu, Z.; Luo, Y.; Lin, J.; Yan, Y.; and Chen, J. 2019 · 2019
Later among the works it cites.
Saliency-Guided Attention Network for Image-Sentence Matching
Ji, Z.; Wang, H.; Han, J.; and Pang, Y. 2019 · 2019
Later among the works it cites.
Fashion Retrieval via Graph Reasoning Networks on a Similarity Pyramid
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations
Krishna, R.; Zhu, Y.; Groth, O.; Johnson, J.; Hata, K.; Kravitz, J.; Chen, S.; Kalantidis, Y.; Li, L.; Shamma, D. A.; Bernstein, M. S.; and Fei-Fei, L. 2017 · 2017
Cited alongside, same era.
Learning a Recurrent Residual Fusion Network for Multimodal Matching
Liu, Y.; Guo, Y.; Bakker, E. M.; and Lew, M. S. 2017 · 2017
Cited alongside, same era.
Dual Attention Networks for Multimodal Reasoning and Matching
Nam, H.; Ha, J.; and Kim, J. 2017 · 2017
Cited alongside, same era.
Graph-Structured Representations for Visual Question Answering
Teney, D.; Liu, L.; and van den Hengel, A. 2017 · 2017
Cited alongside, same era.
Neural Machine Translation with Latent Semantic of Image and Text
Toyama, J.; Misono, M.; Suzuki, M.; Nakayama, K.; and Matsuo, Y. 2017 · 2017
Cited alongside, same era.
Attention is All you Need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, L.; and Polosukhin, I. 2017 · 2017
Cited alongside, same era.
Kuang, Z.; Gao, Y.; Li, G.; Luo, P.; Chen, Y.; Lin, L.; and Zhang, W. 2019 · 2019
Later among the works it cites.
Visual Semantic Reasoning for Image-Text Matching
Li, K.; Zhang, Y.; Li, K.; Li, Y.; and Fu, Y. 2019 · 2019
Later among the works it cites.
Focus Your Attention: A Bidirectional Focal Attention Network for Image-Text Matching
Liu, C.; Mao, Z.; Liu, A.; Zhang, T.; Wang, B.; and Zhang, Y. 2019 · 2019
Later among the works it cites.
Knowledge Aware Semantic Concept Expansion for Image-Text Matching
Shi, B.; Ji, L.; Lu, P.; Niu, Z.; and Duan, N. 2019 · 2019
Later among the works it cites.
Auto-Encoding Scene Graphs for Image Captioning
Yang, X.; Tang, K.; Zhang, H.; and Cai, J. 2019 · 2019
Later among the works it cites.
Multi-Level Visual-Semantic Alignments with Relation-Wise Dual Attention Network for Image and Text Matching
Hu, Z.; Luo, Y.; Lin, J.; Yan, Y.; and Chen, J. 2019 · 2019
Later among the works it cites.
Saliency-Guided Attention Network for Image-Sentence Matching
Ji, Z.; Wang, H.; Han, J.; and Pang, Y. 2019 · 2019
Later among the works it cites.
Fashion Retrieval via Graph Reasoning Networks on a Similarity Pyramid
Kuang, Z.; Gao, Y.; Li, G.; Luo, P.; Chen, Y.; Lin, L.; and Zhang, W. 2019 · 2019
Later among the works it cites.
Visual Semantic Reasoning for Image-Text Matching
Li, K.; Zhang, Y.; Li, K.; Li, Y.; and Fu, Y. 2019 · 2019
Later among the works it cites.
Focus Your Attention: A Bidirectional Focal Attention Network for Image-Text Matching
Liu, C.; Mao, Z.; Liu, A.; Zhang, T.; Wang, B.; and Zhang, Y. 2019 · 2019
Later among the works it cites.
Knowledge Aware Semantic Concept Expansion for Image-Text Matching
Shi, B.; Ji, L.; Lu, P.; Niu, Z.; and Duan, N. 2019 · 2019
Later among the works it cites.
Auto-Encoding Scene Graphs for Image Captioning
Yang, X.; Tang, K.; Zhang, H.; and Cai, J. 2019 · 2019
Later among the works it cites.
IMRAM: Iterative Matching with Recurrent Attention Memory for Cross-Modal Image-Text Retrieval
Chen, H.; Ding, G.; Liu, X.; Lin, Z.; Liu, J.; and Han, J. 2020 · 2020
Later among the works it cites.
Expressing Objects Just Like Words: Recurrent Visual Embedding for Image-Text Matching
Chen, T.; and Luo, J. 2020 · 2020
Later among the works it cites.
Cross-modal Scene Graph Matching for Relationship-aware Image-Text Retrieval
Wang, S.; Wang, R.; Yao, Z.; Shan, S.; and Chen, X. 2020 · 2020
Later among the works it cites.
Adaptive Cross-Modal Embeddings for Image-Text Alignment
Wehrmann, J.; Kolling, C.; and Barros, R. C. 2020 · 2020
Later among the works it cites.
Multi-Modality Cross Attention Network for Image and Sentence Matching
Wei, X.; Zhang, T.; Li, Y.; Zhang, Y.; and Wu, F. 2020 · 2020
Later among the works it cites.
Context-Aware Attention Network for Image-Text Retrieval
Zhang, Q.; Lei, Z.; Zhang, Z.; and Li, S. Z. 2020 · 2020
Later among the works it cites.
IMRAM: Iterative Matching with Recurrent Attention Memory for Cross-Modal Image-Text Retrieval
Chen, H.; Ding, G.; Liu, X.; Lin, Z.; Liu, J.; and Han, J. 2020 · 2020
Later among the works it cites.
Expressing Objects Just Like Words: Recurrent Visual Embedding for Image-Text Matching
Chen, T.; and Luo, J. 2020 · 2020
Later among the works it cites.
Cross-modal Scene Graph Matching for Relationship-aware Image-Text Retrieval
Wang, S.; Wang, R.; Yao, Z.; Shan, S.; and Chen, X. 2020 · 2020
Later among the works it cites.
Adaptive Cross-Modal Embeddings for Image-Text Alignment
Wehrmann, J.; Kolling, C.; and Barros, R. C. 2020 · 2020
Later among the works it cites.
Multi-Modality Cross Attention Network for Image and Sentence Matching
Wei, X.; Zhang, T.; Li, Y.; Zhang, Y.; and Wu, F. 2020 · 2020
Later among the works it cites.
Context-Aware Attention Network for Image-Text Retrieval
Zhang, Q.; Lei, Z.; Zhang, Z.; and Li, S. Z. 2020 · 2020
Later among the works it cites.