Fetching the paper…
Reading the bibliography…
Image-text matching is an interesting and fascinating task in modern AI research.
C.-Y. Lin, “ROUGE: A package for automatic evaluation of summaries,” in Text Summarization Branches Out . Barcelona, Spain: Association for Computational Linguistics, Jul. 2004, pp. 74–81
2004
Earlier work this paper cites.
T. Mikolov, K. Chen, G. Corrado, and J. Dean, “Efficient estimation of word representations in vector space,” in 1st International Conference on Learning Representations, ICLR 2013 , 2013
2013
Earlier work this paper cites.
T. Lin, M. Maire, S. J. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick, “Microsoft COCO: common objects in context,” in ECCV 2014 , ser. Lecture Notes in Computer Science, vol. 8693. Springer, 2014, pp. 740–755
2014
Earlier work this paper cites.
A. Karpathy and F. Li, “Deep visual-semantic alignments for generating image descriptions,” in CVPR 2015 . IEEE Computer Society, 2015, pp. 3128–3137
2015
Earlier work this paper cites.
B. Klein, G. Lev, G. Sadeh, and L. Wolf, “Associating neural word embeddings with deep image representations using fisher vectors,” in CVPR 2015 . IEEE Computer Society, 2015, pp. 4437–4446
2015
Earlier work this paper cites.
C. Sun, C. Gan, and R. Nevatia, “Automatic concept discovery from parallel text and visual corpora,” in Proceedings of the IEEE international conference on computer vision , 2015, pp. 2596–2604
2015
Earlier work this paper cites.
S. Ren, K. He, R. B. Girshick, and J. Sun, “Faster R-CNN: towards real-time object detection with region proposal networks,” in Advances in Neural Information Processing Systems 28 , 2015, pp. 91–99
2015
Earlier work this paper cites.
I. Vendrov, R. Kiros, S. Fidler, and R. Urtasun, “Order-embeddings of images and language,” in 4th International Conference on Learning Representations, ICLR 2016, San Juan, Puerto Rico, May 2-4, 2016, Conference Track Proceedings , Y. Bengio and Y. LeCun, Eds., 2016
2016
Earlier work this paper cites.
X. Lin and D. Parikh, “Leveraging visual question answering for image-caption ranking,” in ECCV 2016 , ser. Lecture Notes in Computer Science, B. Leibe, J. Matas, N. Sebe, and M. Welling, Eds., vol. 9906. Springer, 2016, pp. 261–277
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
P. Anderson, B. Fernando, M. Johnson, and S. Gould, “SPICE: semantic propositional image caption evaluation,” in ECCV 2016 , ser. Lecture Notes in Computer Science, vol. 9909. Springer, 2016, pp. 382–398
2016
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” in NeurIPS 2017 , 2017, pp. 5998–6008
2017
Earlier work this paper cites.
Y. Huang, W. Wang, and L. Wang, “Instance-aware image and sentence matching with selective multimodal LSTM,” in CVPR 2017 . IEEE Computer Society, 2017, pp. 7254–7262
2017
Earlier work this paper cites.
A. Eisenschtat and L. Wolf, “Linking image and text with 2-way nets,” in CVPR 2017 . IEEE Computer Society, 2017, pp. 1855–1865
2017
Cited alongside, same era.
Y. Liu, Y. Guo, E. M. Bakker, and M. S. Lew, “Learning a recurrent residual fusion network for multimodal matching,” in IEEE International Conference on Computer Vision, ICCV 2017 . IEEE Computer Society, 2017, pp. 4127–4136
2017
Cited alongside, same era.
2017
Cited alongside, same era.
A. Santoro, D. Raposo, D. G. Barrett, M. Malinowski, R. Pascanu, P. Battaglia, and T. Lillicrap, “A simple neural network module for relational reasoning,” in Advances in neural information processing systems , 2017, pp. 4967–4976
2017
Cited alongside, same era.
T. Yao, Y. Pan, Y. Li, and T. Mei, “Exploring visual relationship for image captioning,” in ECCV 2018 , ser. Lecture Notes in Computer Science, vol. 11218. Springer, 2018, pp. 711–727
2018
Later among the works it cites.
J. Yang, J. Lu, S. Lee, D. Batra, and D. Parikh, “Graph R-CNN for scene graph generation,” in ECCV 2018 , ser. Lecture Notes in Computer Science, vol. 11205. Springer, 2018, pp. 690–706
2018
Later among the works it cites.
Y. Li, W. Ouyang, B. Zhou, J. Shi, C. Zhang, and X. Wang, “Factorizable net: An efficient subgraph-based framework for scene graph generation,” in ECCV 2018 , ser. Lecture Notes in Computer Science, vol. 11205. Springer, 2018, pp. 346–363
2018
Later among the works it cites.
P. Anderson, X. He, C. Buehler, D. Teney, M. Johnson, S. Gould, and L. Zhang, “Bottom-up and top-down attention for image captioning and visual question answering,” in CVPR 2018 . IEEE Computer Society, 2018, pp. 6077–6086
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
R. Hu, J. Andreas, M. Rohrbach, T. Darrell, and K. Saenko, “Learning to reason: End-to-end module networks for visual question answering,” in The IEEE International Conference on Computer Vision (ICCV) , 2017
2017
Cited alongside, same era.
J. Johnson, B. Hariharan, L. van der Maaten, J. Hoffman, L. Fei-Fei, C. Lawrence Zitnick, and R. Girshick, “Inferring and executing programs for visual reasoning,” in The IEEE International Conference on Computer Vision (ICCV) , 2017
2017
Cited alongside, same era.
F. Faghri, D. J. Fleet, J. R. Kiros, and S. Fidler, “VSE++: improving visual-semantic embeddings with hard negatives,” in BMVC 2018 . BMVA Press, 2018, p. 12
2018
Cited alongside, same era.
F. Carrara, A. Esuli, T. Fagni, F. Falchi, and A. M. Fernández, “Picture it in your mind: generating high level visual representations from textual descriptions,” Inf. Retr. J. , vol. 21, no. 2-3, pp. 208–229, 2018
2018
Cited alongside, same era.
K. Lee, X. Chen, G. Hua, H. Hu, and X. He, “Stacked cross attention for image-text matching,” in ECCV 2018 , ser. Lecture Notes in Computer Science, vol. 11208. Springer, 2018, pp. 212–228
2018
Cited alongside, same era.
J. Gu, J. Cai, S. R. Joty, L. Niu, and G. Wang, “Look, imagine and match: Improving textual-visual cross-modal retrieval with generative models,” in CVPR 2018 . IEEE Computer Society, 2018, pp. 7181–7189
2018
Cited alongside, same era.
Y. Huang, Q. Wu, C. Song, and L. Wang, “Learning semantic concepts and order for image and sentence matching,” in CVPR 2018 . IEEE Computer Society, 2018, pp. 6163–6171
2018
Cited alongside, same era.
——, “Learning relationship-aware visual features,” in ECCV 2018 Workshops , ser. Lecture Notes in Computer Science, vol. 11132. Springer, 2018, pp. 486–501
2018
Cited alongside, same era.
J. Lu, D. Batra, D. Parikh, and S. Lee, “Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks,” in NeurIPS 2019 , 2019, pp. 13–23
2019
Later among the works it cites.
K. Li, Y. Zhang, K. Li, Y. Li, and Y. Fu, “Visual semantic reasoning for image-text matching,” in ICCV 2019 . IEEE, 2019, pp. 4653–4661
2019
Later among the works it cites.
J. Devlin, M. Chang, K. Lee, and K. Toutanova, “BERT: pre-training of deep bidirectional transformers for language understanding,” in NAACL-HLT 2019 . Association for Computational Linguistics, 2019, pp. 4171–4186
2019
Later among the works it cites.
N. Messina, G. Amato, F. Carrara, F. Falchi, and C. Gennaro, “Learning visual features for relational cbir,” International Journal of Multimedia Information Retrieval , Sep 2019
2019
Later among the works it cites.
X. Yang, K. Tang, H. Zhang, and J. Cai, “Auto-encoding scene graphs for image captioning,” in CVPR 2019 . Computer Vision Foundation / IEEE, 2019, pp. 10 685–10 694
2019
Later among the works it cites.
X. Li and S. Jiang, “Know more say less: Image captioning based on scene graphs,” IEEE Trans. Multimedia , vol. 21, no. 8, pp. 2117–2130, 2019
2019
Later among the works it cites.
2019
Later among the works it cites.