Fetching the paper…
Reading the bibliography…
We present a method that learns to answer visual questions by selecting image regions relevant to the text-based query.
Generating typed dependency parses from phrase structure parses
M.-C. De Marneffe, B. MacCartney, C. D. Manning, et al · 2006
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
X. Glorot and Y. Bengio · 2010
Earlier work this paper cites.
Deep sparse rectifier neural networks
X. Glorot, A. Bordes, and Y. Bengio · 2011
Earlier work this paper cites.
Efficient estimation of word representations in vector space
T. Mikolov, K. Chen, G. Corrado, and J. Dean · 2013
Earlier work this paper cites.
Return of the devil in the details: Delving deep into convolutional nets
K. Chatfield, K. Simonyan, A. Vedaldi, and A. Zisserman · 2014
Earlier work this paper cites.
Long-term recurrent convolutional networks for visual recognition and description
J. Donahue, L. A. Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell · 2014
Earlier work this paper cites.
From captions to visual concepts and back
H. Fang, S. Gupta, F. Iandola, R. Srivastava, L. Deng, P. Dollár, J. Gao, X. He, M. Mitchell, J. Platt, et al · 2014
Earlier work this paper cites.
Improving image-sentence embeddings using large weakly annotated photo collections
Y. Gong, L. Wang, M. Hodosh, J. Hockenmaier, and S. Lazebnik · 2014
Earlier work this paper cites.
Explain images with multimodal recurrent neural networks
J. Mao, W. Xu, Y. Yang, J. Wang, and A. L. Yuille · 2014
Cited alongside, same era.
Edge boxes: Locating object proposals from edges
C. L. Zitnick and P. Dollár · 2014
Cited alongside, same era.
Vqa: Visual question answering
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh · 2015
Cited alongside, same era.
Predicting deep zero-shot convolutional neural networks using textual descriptions
J. Ba, K. Swersky, S. Fidler, and R. Salakhutdinov · 2015
Cited alongside, same era.
Mind’s eye: A recurrent visual representation for image caption generation
X. Chen and C. Lawrence Zitnick · 2015
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ask your neurons: A neural-based approach to answering questions about images
M. F. Mateusz Malinowski, Marcus Rohrbach · 2015
Closest in time.
Exploring models and data for image question answering
M. Ren, R. Kiros, and R. Zemel · 2015
Closest in time.
ImageNet Large Scale Visual Recognition Challenge
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. C. Berg, and L. Fei-Fei · 2015
Closest in time.
Weakly supervised memory networks
S. Sukhbaatar, A. Szlam, J. Weston, and R. Fergus · 2015
Closest in time.
Matconvnet – convolutional neural networks for matlab
A. Vedaldi and K. Lenc · 2015
Closest in time.
Show and tell: A neural image caption generator
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Ioffe and C. Szegedy · 2015
Cited alongside, same era.
Deep visual-semantic alignments for generating image descriptions
A. Karpathy and L. Fei-Fei · 2015
Cited alongside, same era.
Deep captioning with multimodal recurrent neural networks (m-rnn)
J. Mao, W. Xu, Y. Yang, J. Wang, Z. Huang, and A. Yuille · 2015
Cited alongside, same era.
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan · 2015
Closest in time.
Show, attend and tell: Neural image caption generation with visual attention
K. Xu, J. Ba, R. Kiros, A. Courville, R. Salakhutdinov, R. Zemel, and Y. Bengio · 2015
Closest in time.
Visual madlibs: Fill in the blank image generation and question answering
L. Yu, E. Park, A. C. Berg, and T. L. Berg · 2015
Closest in time.