Fetching the paper…
Reading the bibliography…
We propose a weakly-supervised approach that takes image-sentence pairs as input and learns to visually ground (i.e., localize) arbitrary linguistic phrases, in the form of spatial attention masks.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Unsupervised Learning of Models for Recognition
M. Weber, M. Welling, and P. Perona · 2000
Earlier work this paper cites.
Object Class Recognition by Unsupervised Scale-Invariant Learning
R. Fergus, P. Perona, and A. Zisserman · 2003
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Automatic Attribute Discovery and Characterization from Noisy Web Data
T. Berg, A. Berg, and J. Shih · 2010
Earlier work this paper cites.
Localizing objects while learning their appearance
T. Deselaers, B. Alexe, and V. Ferrari · 2010
Earlier work this paper cites.
Scene recognition and weakly supervised object localization with deformable part-based models
M. Pandey and S. Lazebnik · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
In Defence of Negative Mining for Annotating Weakly Labelled Data
P. Siva, C. Russell, and T. Xiang · 2012
Earlier work this paper cites.
One billion word benchmark for measuring progress in statistical language modeling
C. Chelba, T. Mikolov, M. Schuster, Q. Ge, T. Brants, P. Koehn, and T. Robinson · 2013
Earlier work this paper cites.
Parsing with compositional vector grammars
R. Socher, J. Bauer, C. D. Manning, and A. Y. Ng · 2013
Earlier work this paper cites.
Weakly supervised learning for attribute localization in outdoor scenes
S. Wang, J. Joo, Y. Wang, and S. C. Zhu · 2013
Earlier work this paper cites.
Multi-fold MIL Training for Weakly Supervised Object Localization
R. Cinbis, J. Verbeek, and C. Schmid · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. Kingma and J. Ba · 2014
Earlier work this paper cites.
Unifying visual-semantic embeddings with multimodal neural language models
R. Kiros, R. Salakhutdinov, and R. S. Zemel · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Earlier work this paper cites.
Weakly-supervised discovery of visual pattern configurations
H. O. Song, Y. J. Lee, S. Jegelka, and T. Darrell · 2014
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
N. Srivastava, G. E. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov · 2014
Cited alongside, same era.
Tell me what you see and i will show you where it is
J. Xu, A. G. Schwing, and R. Urtasun · 2014
Cited alongside, same era.
Vqa: Visual question answering
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. Lawrence Zitnick, and D. Parikh · 2015
Cited alongside, same era.
Mind’s eye: A recurrent visual representation for image caption generation
X. Chen and C. Lawrence Zitnick · 2015
Cited alongside, same era.
Long-term recurrent convolutional networks for visual recognition and description
J. Donahue, L. Anne Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell · 2015
Cited alongside, same era.
Region-based convolutional networks for accurate object detection and segmentation
Neural module networks
J. Andreas, M. Rohrbach, T. Darrell, and D. Klein · 2016
Later among the works it cites.
Weakly supervised object localization with multi-fold multiple instance learning
R. G. Cinbis, J. Verbeek, and C. Schmid · 2016
Later among the works it cites.
Human attention in visual question answering: Do humans and deep networks look at the same regions?
A. Das, H. Agrawal, C. L. Zitnick, D. Parikh, and D. Batra · 2016
Later among the works it cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Later among the works it cites.
Segmentation from natural language expressions
R. Hu, M. Rohrbach, and T. Darrell · 2016
Later among the works it cites.
Natural language object retrieval
R. Hu, H. Xu, M. Rohrbach, J. Feng, K. Saenko, and T. Darrell · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
R. Girshick, J. Donahue, T. Darrell, and J. Malik · 2015
Cited alongside, same era.
Hypercolumns for object segmentation and fine-grained localization
B. Hariharan, P. Arbeláez, R. Girshick, and J. Malik · 2015
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
S. Ioffe and C. Szegedy · 2015
Cited alongside, same era.
Deep visual-semantic alignments for generating image descriptions
A. Karpathy and L. Fei-Fei · 2015
Cited alongside, same era.
Fully convolutional networks for semantic segmentation
J. Long, E. Shelhamer, and T. Darrell · 2015
Cited alongside, same era.
Is object localization for free?-weakly-supervised learning with convolutional neural networks
M. Oquab, L. Bottou, I. Laptev, and J. Sivic · 2015
Cited alongside, same era.
Constrained convolutional neural networks for weakly supervised segmentation
D. Pathak, P. Krähenbühl, and T. Darrell · 2015
Cited alongside, same era.
Densecap: Fully convolutional localization networks for dense captioning
J. Johnson, A. Karpathy, and L. Fei-Fei · 2016
Later among the works it cites.
Contextlocnet: Context-aware deep network models for weakly supervised localization
V. Kantorov, M. Oquab, M. Cho, and I. Laptev · 2016
Later among the works it cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma, M. Bernstein, and L. Fei-Fei · 2016
Later among the works it cites.
Hierarchical question-image co-attention for visual question answering
J. Lu, J. Yang, D. Batra, and D. Parikh · 2016
Later among the works it cites.
Generating images from captions with attention
E. Mansimov, E. Parisotto, J. L. Ba, and R. Salakhutdinov · 2016
Later among the works it cites.
Grounding of textual phrases in images by reconstruction
A. Rohrbach, M. Rohrbach, R. Hu, T. Darrell, and B. Schiele · 2016
Later among the works it cites.
End-to-end localization and ranking for relative attributes
K. K. Singh and Y. J. Lee · 2016
Later among the works it cites.
Track and transfer: Watching videos to simulate strong human supervision for weakly-supervised object detection
K. K. Singh, F. Xiao, and Y. J. Lee · 2016
Later among the works it cites.
Learning deep structure-preserving image-text embeddings
L. Wang, Y. Li, and S. Lazebnik · 2016
Later among the works it cites.
Structured matching for phrase localization
M. Wang, M. Azab, N. Kojima, R. Mihalcea, and J. Deng · 2016
Later among the works it cites.
Top-down neural attention by excitation backprop
J. Zhang, Z. Lin, J. Brandt, X. Shen, and S. Sclaroff · 2016
Later among the works it cites.
Learning deep features for discriminative localization
B. Zhou, A. Khosla, A. Lapedriza, A. Oliva, and A. Torralba · 2016
Later among the works it cites.