Stacked attention networks for image question answering
Z. Yang, X. He, J. Gao, L. Deng, and A. Smola · 2016
Later among the works it cites.
Bottom-up and top-down attention for image captioning and vqa
Original
P. Anderson, X. He, C. Buehler, D. Teney, M. Johnson, S. Gould, and L. Zhang · 2017
Later among the works it cites.
Detecting visual relationships with deep relational networks
B. Dai, Y. Zhang, and D. Lin · 2017
Later among the works it cites.
Neural collaborative filtering
X. He, L. Liao, H. Zhang, L. Nie, X. Hu, and T.-S. Chua · 2017
Later among the works it cites.
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
J. Johnson, B. Hariharan, L. van der Maaten, L. Fei-Fei, C. L. Zitnick, and R. Girshick · 2017
Later among the works it cites.
Inferring and executing programs for visual reasoning
Original
J. Johnson, B. Hariharan, L. van der Maaten, J. Hoffman, L. Fei-Fei, C. L. Zitnick, and R. Girshick · 2017
Later among the works it cites.
One model to learn them all
Original
L. Kaiser, A. N. Gomez, N. Shazeer, A. Vaswani, N. Parmar, L. Jones, and J. Uszkoreit · 2017
Later among the works it cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma, et al · 2017
Later among the works it cites.
Ask your neurons: A deep learning approach to visual question answering
M. Malinowski, M. Rohrbach, and M. Fritz · 2017
Later among the works it cites.
Recurrent relational networks for complex relational reasoning
Original
R. B. Palm, U. Paquet, and O. Winther · 2017
Later among the works it cites.
A simple neural network module for relational reasoning
Original
A. Santoro, D. Raposo, D. G. Barrett, M. Malinowski, R. Pascanu, P. Battaglia, and T. Lillicrap · 2017
Later among the works it cites.
Visual question answering: A tutorial
D. Teney, Q. Wu, and A. van den Hengel · 2017
Later among the works it cites.
FVQA: Fact-based visual question answering
P. Wang, Q. Wu, C. Shen, A. Dick, and A. van den Hengel · 2017
Later among the works it cites.
Visual question answering: A survey of methods and datasets
Q. Wu, D. Teney, P. Wang, C. Shen, A. Dick, and A. van den Hengel · 2017
Later among the works it cites.
Weakly-supervised visual grounding of phrases with linguistic structures
F. Xiao, L. Sigal, and Y. J. Lee · 2017
Later among the works it cites.
Interpretable counting in visual question answering
A. Trott, C. Xiong, and R. Socher · 2018
Closest in time.