Fetching the paper…
Reading the bibliography…
It is commonly assumed that language refers to high-level visual concepts while leaving low-level visual processing unaffected.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Introduction to the special issue on language–vision interactions
F. Ferreira and M. Tanenhaus · 2007
Earlier work this paper cites.
Visualizing data using t-sne
L. Maaten van G. der and Hinton · 2008
Earlier work this paper cites.
Unconscious effects of language-specific terminology on preattentive color perception
G. Thierry, P. Athanasopoulos, A. Wiggett, B. Dering, and JR. Kuipers · 2009
Earlier work this paper cites.
Learning phrase representations using RNN encoder-decoder for statistical machine translation
K. Cho, B. Van Merriënboer, C. Gulcehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y. Bengio · 2014
Earlier work this paper cites.
Prior expectations evoke stimulus templates in the primary visual cortex
P. Kok, M. Failing, and F. de Lange · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
T. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and L. Zitnick · 2014
Earlier work this paper cites.
Glove: Global Vectors for Word Representation
J. Pennington, R. Socher, and C. Manning · 2014
Earlier work this paper cites.
Vqa: Visual question answering
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, Z. Lawrence, and D. Parikh · 2015
Earlier work this paper cites.
Words jump-start vision: A label advantage in object recognition
B. Boutonnet and G. Lupyan · 2015
Earlier work this paper cites.
Ask your neurons: A neural-based approach to answering questions about images
M. Malinowski, M. Rohrbach, and M. Fritz · 2015
Cited alongside, same era.
Exploring models and data for image question answering
M. Ren, R. Kiros, and R. Zemel · 2015
Cited alongside, same era.
Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
I. Sergey and S. Christian · 2015
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2015
Cited alongside, same era.
Ask, attend and answer: Exploring question-guided spatial attention for visual question answering
H. Xu and K. Saenko · 2015
Cited alongside, same era.
Show, attend and tell: Neural image caption generation with visual attention
K. Xu, J. Ba, R. Kiros, K. Cho, A. Courville, R. Salakhutdinov, R. Zemel, and Y. Bengio · 2015
Multimodal Residual Learning for Visual QA
J.-H. Kim, S-W. Lee, D. Kwak, M.-O. Heo, J. Kim, J.-W. Ha, and B.-Y. Zhang · 2016
Later among the works it cites.
Ask your neurons: A deep learning approach to visual question answering
M. Malinowski, M. Rohrbach, and M. Fritz · 2016
Later among the works it cites.
Stacked attention networks for image question answering
Z. Yang, X. He, J. Gao, and L. Deng A. Smola · 2016
Later among the works it cites.
MUTAN: Multimodal Tucker Fusion for Visual Question Answering
H. Ben-Younes, R. Cadène, N. Thome, and M. Cord · 2017
Closest in time.
Visual Dialog
A. Das, S. Kottur, K. Gupta, A. Singh, D. Yadav, J. Moura, D. Parikh, and D. Batra · 2017
Closest in time.
GuessWhat?! Visual object discovery through multi-modal dialogue
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Multimodal Compact Bilinear Pooling for Visual Question Answering and Visual Grounding
A. Fukui, D. Huk Park, D. Yang, A. Rohrbach, T. Darrell, and M. Rohrbach · 2016
Cited alongside, same era.
Hierarchical question-image co-attention for visual question answering
J. Jiasen, J. Yang, D. Batra, and D. Parikh · 2016
Cited alongside, same era.
Deep residual learning for image recognition
K. Kaiming, Z. Xiangyu, S. Ren, and J. Sun · 2016
Cited alongside, same era.
H. de Vries, F. Strub, S. Chandar, O. Pietquin, H. Larochelle, and A. Courville · 2017
Closest in time.
A Learned Representation For Artistic Style
V. Dumoulin, J. Shlens, and M. Kudlur · 2017
Closest in time.
Hadamard Product for Low-rank Bilinear Pooling
J.-H. Kim, K. W. On, W. Lim, J. Kim, J.-W Ha, and B.-T. Zhang · 2017
Closest in time.
Making the V in VQA matter: Elevating the role of image understanding in Visual Question Answering
G. Yashand K. Tejas, S. Douglas, Dhruv B, and P. Devi · 2017
Closest in time.