Fetching the paper…
Reading the bibliography…
Existing methods for visual reasoning attempt to directly map inputs to outputs using black-box architectures without explicitly modeling the underlying reasoning processes.
Understanding Natural Language
T. Winograd · 1972
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
R. J. Williams · 1992
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Accurate unlexicalized parsing
D. Klein and C. D. Manning · 2003
Earlier work this paper cites.
Learning dependency-based compositional semantics
P. Liang, M. I. Jordan, and D. Klein · 2011
Earlier work this paper cites.
ImageNet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
A. Graves, G. Wayne, and I. Danihelka · 2014
Earlier work this paper cites.
A multi-world approach to question answering about real-world scenes based on uncertain input
M. Malinowski and M. Fritz · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
I. Sutskever, O. Vinyals, and Q. V. Le · 2014
Earlier work this paper cites.
W. Zaremba and I. Sutskever · 2014
Earlier work this paper cites.
VQA: Visual question answering
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. Zitnick, and D. Parikh · 2015
Earlier work this paper cites.
Exploring nearest neighbor approaches for image captioning
J. Devlin, S. Gupta, R. Girshick, M. Mitchell, and C. L. Zitnick · 2015
Earlier work this paper cites.
Fast R-CNN
R. Girshick · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
S. Ioffe and C. Szegedy · 2015
Earlier work this paper cites.
Inferring algorithmic patterns with stack-augmented recurrent nets
A. Joulin and T. Mikolov · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. Kingma and J. Ba · 2015
Earlier work this paper cites.
Ask your neurons: A neural-based approach to answering questions about images
M. Malinowski, M. Rohrbach, and M. Fritz · 2015
Earlier work this paper cites.
ImageNet large scale visual recognition challenge
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, et al · 2015
Cited alongside, same era.
End-to-end memory networks
S. Sukhbaatar, A. Szlam, J. Weston, and R. Fergus · 2015
Cited alongside, same era.
Memory networks
J. Weston, S. Chopra, and A. Bordes · 2015
Cited alongside, same era.
Reinforcement learning neural turing machines
W. Zaremba and I. Sutskever · 2015
Cited alongside, same era.
Learning to compose neural networks for question answering
J. Andreas, M. Rohrbach, T. Darrell, and D. Klein · 2016
Cited alongside, same era.
Neural module networks
J. Andreas, M. Rohrbach, T. Darrell, and D. Klein · 2016
Cited alongside, same era.
Learning models for actions and person-object interactions with transfer to question answering
A. Mallya and S. Lazebnik · 2016
Later among the works it cites.
Neural programmer: Inducing latent programs with gradient descent
A. Neelakantan, Q. V. Le, and I. Sutskever · 2016
Later among the works it cites.
Neural programmer-interpreters
S. Reed and N. De Freitas · 2016
Later among the works it cites.
Movieqa: Understanding stories in movies through question-answering
M. Tapaswi, Y. Zhu, R. Stiefelhagen, A. Torralba, R. Urtasun, and S. Fidler · 2016
Later among the works it cites.
Visual question answering: A survey of methods and datasets
Q. Wu, D. Teney, P. Wang, C. Shen, A. Dick, and A. van den Hengel · 2016
Later among the works it cites.
Dynamic memory networks for visual and textual question answering
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Multimodal compact bilinear pooling for visual question answering and visual grounding
A. Fukui, D. H. Park, D. Yang, A. Rohrbach, T. Darrell, and M. Rohrbach · 2016
Cited alongside, same era.
Compact bilinear pooling
Y. Gao, O. Beijbom, N. Zhang, and T. Darrell · 2016
Cited alongside, same era.
Hybrid computing using a neural network with dynamic external memory
A. Graves, G. Wayne, M. Reynolds, T. Harley, I. Danihelka, A. Grabska-Barwinska, S. Colmenarejo, E. Grefenstette, T. Ramalho, J. Agapiou, A. Badia, K. Hermann, Y. Zwols, G. Ostrovski, A. Cain, H. King, C. Summerfield, P. Blunsom, K. . Kavukcuoglu, and D. Hassabis · 2016
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
Revisiting visual question answering baselines
A. Jabri, A. Joulin, and L. van der Maaten · 2016
Cited alongside, same era.
Visual question answering: Datasets, algorithms, and future challenges
K. Kafle and C. Kanan · 2016
Cited alongside, same era.
C. Xiong, S. Merity, and R. Socher · 2016
Later among the works it cites.
Stacked attention networks for image question answering
Z. Yang, X. He, J. Gao, L. Deng, and A. Smola · 2016
Later among the works it cites.
Learning simple algorithms from examples
W. Zaremba, T. Mikolov, A. Joulin, and R. Fergus · 2016
Later among the works it cites.
Yin and yang: Balancing and answering binary visual questions
P. Zhang, Y. Goyal, D. Summers-Stay, D. Batra, and D. Parikh · 2016
Later among the works it cites.
Visual7W: Grounded question answering in images
Y. Zhu, O. Groth, M. Bernstein, and L. Fei-Fei · 2016
Later among the works it cites.
Deepcoder: Learning to write programs
M. Balog, A. Gaunt, M. Brockschmidt, S. Nowozin, and D. Tarlow · 2017
Closest in time.
Making neural programming architectures generalize via recursion
J. Cai, R. Shin, and D. Song · 2017
Closest in time.
Visual dialog
A. Das, S. Kottur, K. Gupta, A. Singh, D. Yadav, J. Moura, D. Parikh, and D. Batra · 2017
Closest in time.
Making the V in VQA matter: Elevating the role of image understanding in visual question answering
Y. Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh · 2017
Closest in time.
Modeling relationships in referential expressions with compositional modular networks
R. Hu, M. Rohrbach, J. Andreas, T. Darrell, and K. Saenko · 2017
Closest in time.
CLEVR: A diagnostic dataset for compositional language and elementary visual reasoning
J. Johnson, B. Hariharan, L. van der Maaten, L. Fei-Fei, C. L. Zitnick, and R. Girshick · 2017
Closest in time.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma, et al · 2017
Closest in time.