Learning to answer questions from image using convolutional neural network
L. Ma, Z. Lu, and H. Li · 2016
Later among the works it cites.
Learning deep representations of fine-grained visual descriptions
S. Reed, Z. Akata, H. Lee, and B. Schiele · 2016
Later among the works it cites.
Where to look: Focus regions for visual question answering
K. J. Shih, S. Singh, and D. Hoiem · 2016
Later among the works it cites.
Ask me anything: Free-form visual question answering based on knowledge from external sources
Q. Wu, P. Wang, C. Shen, A. Dick, and A. van den Hengel · 2016
Later among the works it cites.
Ask, attend and answer: Exploring question-guided spatial attention for visual question answering
H. Xu and K. Saenko · 2016
Later among the works it cites.
Stacked attention networks for image question answering
Z. Yang, X. He, J. Gao, L. Deng, and A. Smola · 2016
Later among the works it cites.
Visual7w: Grounded question answering in images
Y. Zhu, O. Groth, M. Bernstein, and L. Fei-Fei · 2016
Later among the works it cites.
Mutan: Multimodal tucker fusion for visual question answering
H. Ben-Younes, R. Cadene, M. Cord, and N. Thome · 2017
Later among the works it cites.
Variational inference: A review for statisticians
D. M. Blei, A. Kucukelbir, and J. D. McAuliffe · 2017
Later among the works it cites.
Deep bayesian active learning with image data
Y. Gal, R. Islam, and Z. Ghahramani · 2017
Later among the works it cites.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Y. Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh · 2017
Later among the works it cites.
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
J. Johnson, B. Hariharan, L. van der Maaten, L. Fei-Fei, C. Lawrence Zitnick, and R. Girshick · 2017
Later among the works it cites.
What uncertainties do we need in bayesian deep learning for computer vision?
A. Kendall and Y. Gal · 2017
Later among the works it cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma, M. S. Bernstein, and L. Fei-Fei · 2017
Later among the works it cites.
Active learning for visual question answering: An empirical study
Original
X. Lin and D. Parikh · 2017
Later among the works it cites.
Active learning for convolutional neural networks: A core-set approach
Original
O. Sener and S. Savarese · 2017
Later among the works it cites.
Deep active learning for named entity recognition
Original
Y. Shen, H. Yun, Z. C. Lipton, Y. Kronrod, and A. Anandkumar · 2017
Later among the works it cites.
Bottom-up and top-down attention for image captioning and visual question answering
P. Anderson, X. He, C. Buehler, D. Teney, M. Johnson, S. Gould, and L. Zhang · 2018
Later among the works it cites.
Adversarial active learning for sequences labeling and generation
Y. Deng, K. Chen, Y. Shen, and H. Jin · 2018
Later among the works it cites.
Confidence modeling for neural semantic parsing
Original
L. Dong, C. Quirk, and M. Lapata · 2018
Later among the works it cites.
Vizwiz grand challenge: Answering visual questions from blind people
D. Gurari, Q. Li, A. J. Stangl, A. Guo, C. Lin, K. Grauman, J. Luo, and J. P. Bigham · 2018
Later among the works it cites.
Learning by asking questions
I. Misra, R. Girshick, R. Fergus, M. Hebert, A. Gupta, and L. Van Der Maaten · 2018
Later among the works it cites.
Visual question answer diversity
C.-J. Yang, K. Grauman, and D. Gurari · 2018
Later among the works it cites.
Why does a visual question have different answers?
Original
N. Bhattacharya, Q. Li, and D. Gurari · 2019
Closest in time.
Learning to caption images through a lifetime by asking questions
T. Shen, A. Kar, and S. Fidler · 2019
Closest in time.
Bertscore: Evaluating text generation with bert
Original
T. Zhang, V. Kishore, F. Wu, K. Q. Weinberger, and Y. Artzi · 2019
Closest in time.