Fetching the paper…
Reading the bibliography…
We study the problem of visual question answering (VQA) in images by exploiting supervised domain adaptation, where there is a large amount of labeled data in the source domain but only limited labeled data in the target domain with the goal to train a good target model.
Geodesic flow kernel for unsupervised domain adaptation
B. Gong, Y. Shi, F. Sha, and K. Grauman · 2012
Earlier work this paper cites.
Cross language text classification via subspace co-regularized multi-view learning
Y. Guo and M. Xiao · 2012
Earlier work this paper cites.
Learning phrase representations using RNN encoder-decoder for statistical machine translation
K. Cho, B. van Merriënboer, C. Gulcehre, F. Bougares, H. Schwenk, and Y. Bengio · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
J. Pennington, R. Socher, and C. D. Manning · 2014
Earlier work this paper cites.
Vqa: Visual question answering
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh · 2015
Earlier work this paper cites.
Unsupervised domain adaptation by backpropagation
Y. Ganin and V. S. Lempitsky · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2015
Earlier work this paper cites.
Simultaneous deep transfer across domains and tasks
J. Hoffman, E. Tzeng, T. Darrell, and K. Saenko · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
S. Ren, K. He, R. Girshick, and J. Sun · 2015
Earlier work this paper cites.
Semi-supervised domain adaptation with subspace learning for visual recognition
T. Yao, Y. Pan, C.-W. Ngo, H. Li, and T. Mei · 2015
Cited alongside, same era.
Simple baseline for visual question answering
B. Zhou, Y. Tian, S. Sukhbaatar, A. Szlam, and R. Fergus · 2015
Cited alongside, same era.
Revisiting visual question answering baselines
A. Jabri, A. Joulin, and L. van der Maaten · 2016
Cited alongside, same era.
Feature pyramid networks for object detection
T.-Y. Lin, P. Dollár, R. B. Girshick, K. He, B. Hariharan, and S. J. Belongie · 2016
Cited alongside, same era.
Visual7w: Grounded question answering in images
Y. Zhu, O. Groth, M. Bernstein, and F. F. Li · 2016
Cited alongside, same era.
Bottom-up and top-down attention for image captioning and visual question answering
Adversarial discriminative domain adaptation
E. Tzeng, J. Hoffman, K. Saenko, and T. Darrell · 2017
Later among the works it cites.
Aggregated residual transformations for deep neural networks
S. Xie, R. B. Girshick, P. Dollár, Z. Tu, and K. He · 2017
Later among the works it cites.
Detectron
R. Girshick, I. Radosavovic, G. Gkioxari, P. Dollár, and K. He · 2018
Later among the works it cites.
Vizwiz grand challenge: Answering visual questions from blind people
D. Gurari, Q. Li, A. J. Stangl, A. Guo, C. Lin, K. Grauman, J. Luo, and J. P. Bigham · 2018
Later among the works it cites.
Bilinear attention networks
J.-H. Kim, J. Jun, and B.-T. Zhang · 2018
Later among the works it cites.
A unified framework for multimodal domain adaptation
F. Qi, X. Yang, and C. Xu · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
P. Anderson, X. He, C. Buehler, D. Teney, M. Johnson, S. Gould, and L. Zhang · 2017
Cited alongside, same era.
Show, ask, attend, and answer: A strong baseline for visual question answering
V. Kazemi and A. Elqursh · 2017
Cited alongside, same era.
Domain adaptation by mixture of alignments of second- or higher-order scatter tensors
P. Koniusz, Y. Tas, and F. Porikli · 2017
Cited alongside, same era.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma, M. S. Bernstein, and L. Fei-Fei · 2017
Cited alongside, same era.
Wasserstein distance guided representation learning for domain adaptation
J. Shen, Y. Qu, W. Zhang, and Y. Yu · 2017
Cited alongside, same era.
Cross-dataset adaptation for visual question answering
F. Sha, H. Hu, and W.-L. Chao · 2018
Later among the works it cites.
Learning to count objects in natural images for visual question answering
Y. Zhang, J. Hare, and A. Prügel-Bennett · 2018
Later among the works it cites.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Y. Goyal, T. Khot, A. Agrawal, D. Summers-Stay, D. Batra, and D. Parikh · 2019
Closest in time.
Towards VQA models that can read
A. Singh, V. Natarajan, M. Shah, Y. Jiang, X. Chen, D. Batra, D. Parikh, and M. Rohrbach · 2019
Closest in time.