Fetching the paper…
Reading the bibliography…
This paper studies the task of matching image and sentence, where learning appropriate representations across the multi-modal data appears to be the main challenge.
Saliency detection via absorbing markov chain, 2013
Bowen Jiang, Lihe Zhang, Huchuan Lu, Chuan Yang, and Ming-Hsuan Yang · 2013
Earlier work this paper cites.
Salient object detection: A discriminative regional feature integration approach, 2013
Huaizu Jiang, Jingdong Wang, Zejian Yuan, Yang Wu, Nanning Zheng, and Shipeng Li · 2013
Earlier work this paper cites.
Deep fragment embeddings for bidirectional image sentence mapping, 2014
Andrej Karpathy, Armand Joulin, and Fei-Fei Li · 2014
Earlier work this paper cites.
Adam: Amethod for stochastic optimization, 2014
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Unifying visual-semantic embeddings with multimodal neural language models, 2014
Ryan Kiros, Ruslan Salakhutdinov, and Richard S. Zemel · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context, 2014
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition, 2014
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Peter Young, Alice Lai, Micah Hodosh, and Julia Hockenmaier · 2014
Earlier work this paper cites.
Vqa: Visual question answering, 2015
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, Zitnick C Lawrence, and Devi Parikh · 2015
Earlier work this paper cites.
Global contrast based salient region detection
Ming-Ming Cheng, Niloy J Mitra, Xiaolei Huang, Philip HS Torr, and Shi-Min Hu · 2015
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classi-fication, 2015
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Instance-aware image and sentence matching with selective multimodal lstm, 2015
Yan Huang, Wei Wang, and Liang Wang · 2015
Earlier work this paper cites.
Deep visual-semantic alignments for generating image descriptions, 2015
Andrej Karpathy and Li Fei-Fei · 2015
Earlier work this paper cites.
Associating neural word embeddings with deep image representations using fisher vectors, 2015
Benjamin Klein, Guy Lev, Gil Sadeh, and Lior Wolf · 2015
Earlier work this paper cites.
Multimodal convolutional neural networks for matching image and sentence, 2015
Lin Ma, Zhengdong Lu, Lifeng Shang, and Hang Li · 2015
Earlier work this paper cites.
Deep captioning with multimodal recurrent neural networks (m-rnn), 2015
Junhua Mao, Wei Xu, Yi Yang, Jiang Wang, Zhiheng Huang, and Alan Yuille · 2015
Cited alongside, same era.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models, 2015
Bryan A Plummer, Liwei Wang, Chris M Cervantes, Juan C Caicedo, Julia Hockenmaier, and Svetlana Lazebnik · 2015
Cited alongside, same era.
Faster r-cnn: Towards real-time object detection with region proposal networks, 2015
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun · 2015
Cited alongside, same era.
Show, attend and tell: Neural image caption generation with visual attention, 2015
Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhudinov, Rich Zemel, and Yoshua Bengio · 2015
Cited alongside, same era.
Deep residual learning for image recognition, 2016
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Aggregated residual transformations for deep neural networks, 2017
Saining Xie, Ross Girshick, Piotr Dollár, Zhuowen Tu, and Kaiming He · 2017
Later among the works it cites.
Amulet: Aggregating multi-level convolutional features for salient object detection, 2017
Pingping Zhang, Dong Wang, Huchuan Lu, Hongyu Wang, and Xiang Ruan · 2017
Later among the works it cites.
Learning uncertain convolutional features for accurate saliency detection, 2017
Pingping Zhang, Dong Wang, Huchuan Lu, Hongyu Wang, and Baocai Yin · 2017
Later among the works it cites.
Dual-path convolutional image-text embedding, 2017
Zhedong Zheng, Liang Zheng, Michael Garrett, Yi Yang, and Yi-Dong Shen · 2017
Later among the works it cites.
Bottom-up and top-down attention for image captioning and vqa, 2018
Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Order-embeddings of images and language, 2016
Ivan Vendrov, Ryan Kiros, Sanja Fidler, and Raquel Urtasun · 2016
Cited alongside, same era.
Learning deep structure-preserving image-text embeddings, 2016
Liwei Wang, Yin Li, and Svetlana Lazebnik · 2016
Cited alongside, same era.
Hierarchical attention networks for document classification, 2016
Zichao Yang, Diyi Yang, Chris Dyer, Xiaodong He, Alex Smola, and Eduard Hovy · 2016
Cited alongside, same era.
Linking image and text with 2-way nets, 2017
Eisenschtat Aviv and Lior Wolf · 2017
Cited alongside, same era.
Deeply supervised salient object detection with short connections, 2017
Qibin Hou, Ming-Ming Cheng, Xiaowei Hu, Ali Borji, Zhuowen Tu, and Philip HS Torr · 2017
Cited alongside, same era.
A structured self-attentive sentence embedding, 2017
Zhouhan Lin, Minwei Feng, Cicero Nogueira dos Santos, Mo Yu, Bing Xiang, Bowen Zhou, and Yoshua Bengio · 2017
Cited alongside, same era.
Dual attention networks for multimodal reasoning and matching, 2017
Hyeonseob Nam, Jung-Woo Ha, and Jeonghee Kim · 2017
Cited alongside, same era.
Knowledge aided consistency for weakly supervised phrase grounding
Kan Chen, Jiyang Gao, and Ram Nevatia · 2018
Later among the works it cites.
R3net: Recurrent residual refinement network for saliency detection, 2018
Zijun Deng, Xiaowei Hu, Lei Zhu, Xuemiao Xu, Jing Qin, GuoqiangHan, and Pheng-Ann Heng · 2018
Later among the works it cites.
Vse++: improved visual-semantic embeddings, 2018
Fartash Faghri, David J Fleet, Jamie R Kiros, and Sanja Fidler · 2018
Later among the works it cites.
Look, imagine and match: Improving textual-visual cross-modal retrieval with gen-erative models, 2018
Jiuxiang Gu, Jianfei Cai, Shafiq R Joty, Li Niu, and Gang Wang · 2018
Later among the works it cites.
Advanced deep-learning techniques for salient and category-specific object detection: a survey
Junwei Han, Dingwen Zhang, Gong Cheng, Nian Liu, and Dong Xu · 2018
Later among the works it cites.
Self-erasing network for integral object attention, 2018
Qibin Hou, PengTao Jiang, Yunchao Wei, and Ming-Ming Cheng · 2018
Later among the works it cites.
Learning semantic concepts and order for image and sentence matching, 2018
Yan Huang, Qi Wu, Chunfeng Song, and Liang Wang · 2018
Later among the works it cites.
Stacked cross attention for image-text matching, 2018
Kuang-Huei Lee, Xi Chen, Gang Hua, Houdong Hu, and Xiaodong He · 2018
Later among the works it cites.
Deep cross-modal projection learning for image-text matching, 2018
Ying Zhang and Huchuan Lu · 2018
Later among the works it cites.
Polysemous visual-semantic embedding for cross-modal retrieval, 2019
Yale Song and Mohammad Soleymani · 2019
Closest in time.