Fetching the paper…
Reading the bibliography…
We tackle the problem of understanding visual ads where given an ad image, our goal is to rank appropriate human generated statements describing the purpose of the ad.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
D. Bahdanau, K. Cho, and Y. Bengio · 2014
Earlier work this paper cites.
Rich feature hierarchies for accurate object detection and semantic segmentation
R. Girshick, J. Donahue, T. Darrell, and J. Malik · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Earlier work this paper cites.
Unifying visual-semantic embeddings with multimodal neural language models
R. Kiros, R. Salakhutdinov, and R. S. Zemel · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
J. Pennington, R. Socher, and C. Manning · 2014
Earlier work this paper cites.
Learning deep features for scene recognition using places database
B. Zhou, A. Lapedriza, J. Xiao, A. Torralba, and A. Oliva · 2014
Earlier work this paper cites.
Edge boxes: Locating object proposals from edges
C. L. Zitnick and P. Dollár · 2014
Earlier work this paper cites.
Action recognition using visual attention
S. Sharma, R. Kiros, and R. Salakhutdinov · 2015
Cited alongside, same era.
Show, attend and tell: Neural image caption generation with visual attention
K. Xu, J. Ba, R. Kiros, K. Cho, A. Courville, R. Salakhudinov, R. Zemel, and Y. Bengio · 2015
Cited alongside, same era.
The cityscapes dataset for semantic urban scene understanding
M. Cordts, M. Omran, S. Ramos, T. Rehfeld, M. Enzweiler, R. Benenson, U. Franke, S. Roth, and B. Schiele · 2016
Cited alongside, same era.
Weldon: Weakly supervised learning of deep convolutional neural networks
T. Durand, N. Thome, and M. Cord · 2016
Cited alongside, same era.
Hierarchical question-image co-attention for visual question answering
J. Lu, J. Yang, D. Batra, and D. Parikh · 2016
Cited alongside, same era.
Bottom-up and top-down attention for image captioning and vqa
P. Anderson, X. He, C. Buehler, D. Teney, M. Johnson, S. Gould, and L. Zhang · 2017
Later among the works it cites.
Human attention in visual question answering: Do humans and deep networks look at the same regions?
A. Das, H. Agrawal, L. Zitnick, D. Parikh, and D. Batra · 2017
Later among the works it cites.
Attentional pooling for action recognition
R. Girdhar and D. Ramanan · 2017
Later among the works it cites.
Automatic understanding of image and video advertisements
Z. Hussain, M. Zhang, X. Zhang, K. Ye, C. Thomas, Z. Agha, N. Ong, and A. Kovashka · 2017
Later among the works it cites.
Adascan: Adaptive scan pooling in deep convolutional neural networks for human action recognition in videos
A. Kar, N. Rai, K. Sikka, and G. Sharma · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
H. Nam, J.-W. Ha, and J. Kim · 2016
Cited alongside, same era.
Attention networks for weakly supervised object localization
E. W. Teh, M. Rochan, and Y. Wang · 2016
Cited alongside, same era.
Ask, attend and answer: Exploring question-guided spatial attention for visual question answering
H. Xu and K. Saenko · 2016
Cited alongside, same era.
Stacked attention networks for image question answering
Z. Yang, X. He, J. Gao, L. Deng, and A. Smola · 2016
Cited alongside, same era.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma, et al · 2017
Later among the works it cites.
Combining bottom-up, top-down, and smoothness cues for weakly supervised image segmentation
A. Roy and S. Todorovic · 2017
Later among the works it cites.
Advise: Symbolism and external knowledge for decoding advertisements
K. Ye and A. Kovashka · 2017
Later among the works it cites.