Fetching the paper…
Reading the bibliography…
The predominant approach to visual question answering (VQA) relies on encoding the image and question with a "black-box" neural encoder and decoding a single token as the answer like "yes" or "no".
Image retrieval using scene graphs
J. Johnson, R. Krishna, M. Stark, L. Li, D. A. Shamma, M. S. Bernstein, and F. Li · 2015
Earlier work this paper cites.
Analyzing the behavior of visual question answering models
A. Agrawal, D. Batra, and D. Parikh · 2016
Earlier work this paper cites.
Learning to compose neural networks for question answering
J. Andreas, M. Rohrbach, T. Darrell, and D. Klein · 2016
Earlier work this paper cites.
Neural module networks
J. Andreas, M. Rohrbach, T. Darrell, and D. Klein · 2016
Earlier work this paper cites.
Human attention in visual question answering: Do humans and deep networks look at the same regions?
A. Das, H. Agrawal, L. Zitnick, D. Parikh, and D. Batra · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
The color of the cat is gray: 1 million full-sentences visual question answering (FSVQA)
A. Shin, Y. Ushiku, and T. Harada · 2016
Earlier work this paper cites.
spacy 2: Natural language understanding with bloom embeddings, convolutional neural networks and incremental parsing
M. Honnibal and I. Montani · 2017
Earlier work this paper cites.
Learning to reason: End-to-end module networks for visual question answering
R. Hu, J. Andreas, M. Rohrbach, T. Darrell, and K. Saenko · 2017
Earlier work this paper cites.
Inferring and executing programs for visual reasoning
J. Johnson, B. Hariharan, L. van der Maaten, J. Hoffman, L. Fei-Fei, C. L. Zitnick, and R. B. Girshick · 2017
Earlier work this paper cites.
Graph-structured representations for visual question answering
D. Teney, L. Liu, and A. van den Hengel · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Scene graph generation by iterative message passing
D. Xu, Y. Zhu, C. B. Choy, and L. Fei-Fei · 2017
Earlier work this paper cites.
Bottom-up and top-down attention for image captioning and visual question answering
P. Anderson, X. He, C. Buehler, D. Teney, M. Johnson, S. Gould, and L. Zhang · 2018
Earlier work this paper cites.
Relational inductive biases, deep learning, and graph networks
P. W. Battaglia, J. B. Hamrick, V. Bapst, A. Sanchez-Gonzalez, V. F. Zambaldi, M. Malinowski, A. Tacchetti, D. Raposo, A. Santoro, R. Faulkner, Ç. Gülçehre, H. F. Song, A. J. Ballard, J. Gilmer, G. E. Dahl, A. Vaswani, K. R. Allen, C. Nash, V. Langston, C. Dyer, N. Heess, D. Wierstra, P. Kohli, M. Botvinick, O. Vinyals, Y. Li, and R. Pascanu · 2018
Earlier work this paper cites.
Explainable neural computation via stack neural module networks
R. Hu, J. Andreas, T. Darrell, and K. Saenko · 2018
Earlier work this paper cites.
Compositional attention networks for machine reasoning
D. A. Hudson and C. D. Manning · 2018
Cited alongside, same era.
Tell-and-answer: Towards explainable visual question answering using attributes and captions
Q. Li, J. Fu, D. Yu, T. Mei, and J. Luo · 2018
Cited alongside, same era.
VQA-E: explaining, elaborating, and enhancing your answers for visual questions
Q. Li, Q. Tao, S. R. Joty, J. Cai, and J. Luo · 2018
Cited alongside, same era.
Did the model understand the question?
P. K. Mudrakarta, A. Taly, M. Sundararajan, and K. Dhamdhere · 2018
Cited alongside, same era.
Graph R-CNN for scene graph generation
J. Yang, J. Lu, S. Lee, D. Batra, and D. Parikh · 2018
Cited alongside, same era.
Exploring visual relationship for image captioning
T. Yao, Y. Pan, Y. Li, and T. Mei · 2018
Explainable and explicit visual reasoning over scene graphs
J. Shi, H. Zhang, and J. Li · 2019
Later among the works it cites.
LXMERT: learning cross-modality encoder representations from transformers
H. Tan and M. Bansal · 2019
Later among the works it cites.
Probabilistic neural symbolic models for interpretable visual question answering
R. Vedantam, K. Desai, S. Lee, M. Rohrbach, D. Batra, and D. Parikh · 2019
Later among the works it cites.
How powerful are graph neural networks?
K. Xu, W. Hu, J. Leskovec, and S. Jegelka · 2019
Later among the works it cites.
Auto-encoding scene graphs for image captioning
X. Yang, K. Tang, H. Zhang, and J. Cai · 2019
Later among the works it cites.
An empirical study on leveraging scene graphs for visual question answering
C. Zhang, W. Chao, and D. Xuan · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Neural motifs: Scene graph parsing with global context
R. Zellers, M. Yatskar, S. Thomson, and Y. Choi · 2018
Cited alongside, same era.
Knowledge-embedded routing network for scene graph generation
T. Chen, W. Yu, R. Chen, and L. Lin · 2019
Cited alongside, same era.
Meta module network for compositional visual reasoning
W. Chen, Z. Gan, L. Li, Y. Cheng, W. Wang, and J. Liu · 2019
Cited alongside, same era.
Cu-net: Component unmixing network for textile fiber identification
Z. Feng, W. Liang, D. Tao, L. Sun, A. Zeng, and M. Song · 2019
Cited alongside, same era.
Interpretation of neural networks is fragile
A. Ghorbani, A. Abid, and J. Y. Zou · 2019
Cited alongside, same era.
Unpaired image captioning via scene graph alignments
J. Gu, S. R. Joty, J. Cai, H. Zhao, X. Yang, and G. Wang · 2019
Cited alongside, same era.
Neuro-symbolic visual reasoning: Disentangling "visual" from "reasoning"
S. Amizadeh, H. Palangi, O. Polozov, Y. Huang, and K. Koishida · 2020
Closest in time.
End-to-end object detection with transformers
N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, and S. Zagoruyko · 2020
Closest in time.
Graph density-aware losses for novel compositions in scene graph generation
B. Knyazev, H. de Vries, C. Cangea, G. W. Taylor, A. C. Courville, and E. Belilovsky · 2020
Closest in time.
R. Koner, P. Sinhamahapatra, and V. Tresp · 2020
Closest in time.
MOSS: end-to-end dialog system framework with modular supervision
W. Liang, Y. Tian, C. Chen, and Z. Yu · 2020
Closest in time.
ALICE: active learning with contrastive natural language explanations
W. Liang, J. Zou, and Z. Yu · 2020
Closest in time.
Beyond user self-reported likert scale ratings: A comparison model for automatic dialog evaluation
W. Liang, J. Zou, and Z. Yu · 2020
Closest in time.
12-in-1: Multi-task vision and language representation learning
J. Lu, V. Goswami, M. Rohrbach, D. Parikh, and S. Lee · 2020
Closest in time.
Unbiased scene graph generation from biased training
K. Tang, Y. Niu, J. Huang, J. Shi, and H. Zhang · 2020
Closest in time.