Fetching the paper…
Reading the bibliography…
Natural language questions are inherently compositional, and many are most easily answered by reasoning about their decomposition into modular sub-problems.
An analysis of a logical machine using parenthesis-free notation
A. W. Burks, D. W. Warren, and J. B. Wright · 1954
Earlier work this paper cites.
Modeling systems with internal state using evolino
D. Wierstra, F. J. Gomez, and J. Schmidhuber · 2005
Earlier work this paper cites.
Neuroevolution: from architectures to learning
D. Floreano, P. Dürr, and C. Mattiussi · 2008
Earlier work this paper cites.
Making a science of model search: Hyperparameter optimization in hundreds of dimensions for vision architectures
J. Bergstra, D. Yamins, and D. D. Cox · 2013
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
R. Socher, A. Perelygin, J. Y. Wu, J. Chuang, C. D. Manning, A. Y. Ng, C. Potts, et al · 2013
Earlier work this paper cites.
Fast and accurate shift-reduce constituent parsing
M. Zhu, Y. Zhang, W. Chen, M. Zhang, and J. Zhu · 2013
Earlier work this paper cites.
A multi-world approach to question answering about real-world scenes based on uncertain input
M. Malinowski and M. Fritz · 2014
Earlier work this paper cites.
Vqa: Visual question answering
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh · 2015
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
D. Bahdanau, K. Cho, and Y. Bengio · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. Kingma and J. Ba · 2015
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2015
Cited alongside, same era.
TensorFlow: Large-scale machine learning on heterogeneous systems
M. Abadi, A. Agarwal, P. Barham, E. Brevdo, Z. Chen, C. Citro, G. S. Corrado, A. Davis, J. Dean, M. Devin, S. Ghemawat, I. Goodfellow, A. Harp, G. Irving, M. Isard, Y. Jia, R. Jozefowicz, L. Kaiser, M. Kudlur, J. Levenberg, D. Mané, R. Monga, S. Moore, D. Murray, C. Olah, M. Schuster, J. Shlens, B. Steiner, I. Sutskever, K. Talwar, P. Tucker, V. Vanhoucke, V. Vasudevan, F. Viégas, O. Vinyals, P. Warden, M. Wattenberg, M. Wicke, Y. Yu, and X. Zheng · 2016
Cited alongside, same era.
Learning to compose neural networks for question answering
J. Andreas, M. Rohrbach, T. Darrell, and D. Klein · 2016
Cited alongside, same era.
Neural module networks
J. Andreas, M. Rohrbach, T. Darrell, and D. Klein · 2016
Cited alongside, same era.
Multimodal compact bilinear pooling for visual question answering and visual grounding
A. Fukui, D. H. Park, D. Yang, A. Rohrbach, T. Darrell, and M. Rohrbach · 2016
Hierarchical Co-Attention for Visual Question Answering
J. Lu, J. Yang, D. Batra, and D. Parikh · 2016
Later among the works it cites.
Image question answering using convolutional neural network with dynamic parameter prediction
H. Noh, P. H. Seo, and B. Han · 2016
Later among the works it cites.
Dynamic memory networks for visual and textual question answering
C. Xiong, S. Merity, and R. Socher · 2016
Later among the works it cites.
Stacked attention networks for image question answering
Z. Yang, X. He, J. Gao, L. Deng, and A. Smola · 2016
Later among the works it cites.
Simple baseline for visual question answering
B. Zhou, Y. Tian, S. Sukhbaatar, A. Szlam, and R. Fergus · 2016
Later among the works it cites.
Modeling relationships in referential expressions with compositional modular networks
R. Hu, M. Rohrbach, J. Andreas, T. Darrell, and K. Saenko · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Making the V in VQA matter: Elevating the role of image understanding in Visual Question Answering
Y. Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh · 2016
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
Segmentation from natural language expressions
R. Hu, M. Rohrbach, and T. Darrell · 2016
Cited alongside, same era.
Revisiting visual question answering baselines
A. Jabri, A. Joulin, and L. van der Maaten · 2016
Cited alongside, same era.
Closest in time.
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
J. Johnson, B. Hariharan, L. van der Maaten, L. Fei-Fei, C. L. Zitnick, and R. Girshick · 2017
Closest in time.
Inferring and executing programs for visual reasoning
J. Johnson, B. Hariharan, L. van der Maaten, J. Hoffman, L. Fei-Fei, C. L. Zitnick, and R. Girshick · 2017
Closest in time.
Neural architecture search with reinforcement learning
B. Zoph and Q. V. Le · 2017
Closest in time.