Fetching the paper…
Reading the bibliography…
Visual Question Answering (VQA) has witnessed tremendous progress in recent years.
P. Banerjee, T. Gokhale, Y. Yang, and C. Baral, “Weakly supervised relative spatial reasoning for visual question answering,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 1908–1918
1918
Earlier work this paper cites.
M. A. Stephens, “Edf statistics for goodness of fit and some comparisons,” Journal of the American Statistical Association , vol. 69, no. 347, pp. 730–737, 1974. [Online]. Available: http://www.jstor.org/stable/2286009
1974
Earlier work this paper cites.
J. Johnson, B. Hariharan, L. van der Maaten, L. Fei-Fei, C. Zitnick, and R. Girshick, “Clevr: A diagnostic dataset for compositional language and elementary visual reasoning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 07 2017, pp. 1988–1997
1997
Earlier work this paper cites.
H. A. David and J. L. Gunnink, “The paired t test under artificial pairing,” The American Statistician , vol. 51, no. 1, pp. 9–12, 1997. [Online]. Available: http://www.jstor.org/stable/2684684
1997
Earlier work this paper cites.
M. Malinowski and M. Fritz, “A multi-world approach to question answering about real-world scenes based on uncertain input,” Advances in neural information processing systems , vol. 27, 2014
2014
Earlier work this paper cites.
M. Denkowski and A. Lavie, “Meteor universal: Language specific translation evaluation for any target language,” in Proceedings of the ninth workshop on statistical machine translation , 2014, pp. 376–380
2014
Earlier work this paper cites.
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh, “Vqa: Visual question answering,” in Proceedings of the IEEE international conference on computer vision , 2015, pp. 2425–2433
2015
Earlier work this paper cites.
H. Gao, J. Mao, J. Zhou, Z. Huang, L. Wang, and W. Xu, “Are you talking to a machine? dataset and methods for multilingual image question answering,” NIPS , 2015
2015
Earlier work this paper cites.
M. Ren, R. Kiros, and R. Zemel, “Exploring models and data for image question answering,” Advances in neural information processing systems , vol. 28, pp. 2953–2961, 2015
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
Z. Yang, X. He, J. Gao, L. Deng, and A. Smola, “Stacked attention networks for image question answering,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 21–29
2016
Earlier work this paper cites.
J. Lu, J. Yang, D. Batra, and D. Parikh, “Hierarchical question-image co-attention for visual question answering,” Advances in neural information processing systems , vol. 29, pp. 289–297, 2016
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
M. Tapaswi, Y. Zhu, R. Stiefelhagen, A. Torralba, R. Urtasun, and S. Fidler, “Movieqa: Understanding stories in movies through question-answering,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 4631–4640
2016
Earlier work this paper cites.
Y. Zhu, O. Groth, M. Bernstein, and L. Fei-Fei, “Visual7w: Grounded question answering in images,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 4995–5004
2016
Earlier work this paper cites.
P. Zhang, Y. Goyal, D. Summers-Stay, D. Batra, and D. Parikh, “Yin and yang: Balancing and answering binary visual questions,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2016, pp. 5014–5022
2016
Earlier work this paper cites.
Y. Goyal, T. Khot, D. Summers Stay, D. Batra, and D. Parikh, “Making the v in vqa matter: Elevating the role of image understanding in visual question answering,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 07 2017, pp. 6325–6334
2017
Earlier work this paper cites.
S. Ebrahimi Kahou, V. Michalski, A. Atkinson, A. Kadar, A. Trischler, and Y. Bengio, “Figureqa: An annotated figure dataset for visual reasoning,” arXiv e-prints , pp. arXiv–1710, 2017
2017
Earlier work this paper cites.
P. Wang, Q. Wu, C. Shen, A. Dick, and A. Van Den Hengel, “Fvqa: Fact-based visual question answering,” IEEE transactions on pattern analysis and machine intelligence , vol. 40, no. 10, pp. 2413–2427, 2017
2017
Earlier work this paper cites.
A. Das, S. Kottur, K. Gupta, A. Singh, D. Yadav, J. M. Moura, D. Parikh, and D. Batra, “Visual dialog,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2017, pp. 326–335
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in Advances in neural information processing systems , 2017, pp. 5998–6008
2017
Earlier work this paper cites.
C. R. Qi, L. Yi, H. Su, and L. J. Guibas, “Pointnet++: Deep hierarchical feature learning on point sets in a metric space,” Advances in Neural Information Processing Systems , vol. 30, 2017
2017
Earlier work this paper cites.
A. Dai, A. X. Chang, M. Savva, M. Halber, T. Funkhouser, and M. Nießner, “Scannet: Richly-annotated 3d reconstructions of indoor scenes,” in Proc. Computer Vision and Pattern Recognition (CVPR), IEEE , 2017
2017
Earlier work this paper cites.
Y. Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh, “Making the v in vqa matter: Elevating the role of image understanding in visual question answering,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2017, pp. 6904–6913
2017
Earlier work this paper cites.
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma et al. , “Visual genome: Connecting language and vision using crowdsourced dense image annotations,” International journal of computer vision , vol. 123, no. 1, pp. 32–73, 2017
2017
Earlier work this paper cites.
J. Johnson, B. Hariharan, L. Van Der Maaten, L. Fei-Fei, C. Lawrence Zitnick, and R. Girshick, “Clevr: A diagnostic dataset for compositional language and elementary visual reasoning,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2017, pp. 2901–2910
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
A. Agrawal, D. Batra, D. Parikh, and A. Kembhavi, “Don’t just assume; look and answer: Overcoming priors for visual question answering,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 06 2018, pp. 4971–4980
2018
Earlier work this paper cites.
K. Kafle, B. Price, S. Cohen, and C. Kanan, “Dvqa: Understanding data visualizations via question answering,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2018, pp. 5648–5656
2018
Earlier work this paper cites.
J. Lei, L. Yu, M. Bansal, and T. L. Berg, “Tvqa: Localized, compositional video question answering,” in EMNLP , 2018
2018
Cited alongside, same era.
P. Anderson, Q. Wu, D. Teney, J. Bruce, M. Johnson, N. Sünderhauf, I. Reid, S. Gould, and A. Van Den Hengel, “Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2018, pp. 3674–3683
2018
Cited alongside, same era.
A. Das, S. Datta, G. Gkioxari, S. Lee, D. Parikh, and D. Batra, “Embodied question answering,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2018, pp. 1–10
2018
Cited alongside, same era.
P. Anderson, X. He, C. Buehler, D. Teney, M. Johnson, S. Gould, and L. Zhang, “Bottom-up and top-down attention for image captioning and visual question answering,” in CVPR , 2018
2018
Cited alongside, same era.
J. Lu, V. Goswami, M. Rohrbach, D. Parikh, and S. Lee, “12-in-1: Multi-task vision and language representation learning,” in The IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , June 2020
2020
Later among the works it cites.
W. Su, X. Zhu, Y. Cao, B. Li, L. Lu, F. Wei, and J. Dai, “Vl-bert: Pre-training of generic visual-linguistic representations,” in International Conference on Learning Representations , 2020. [Online]. Available: https://openreview.net/forum?id=SygXPaEYvH
2020
Later among the works it cites.
Z. Lan, M. Chen, S. Goodman, K. Gimpel, P. Sharma, and R. Soricut, “Albert: A lite bert for self-supervised learning of language representations,” in International Conference on Learning Representations , 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Z. Yu, J. Yu, C. Xiang, J. Fan, and D. Tao, “Beyond bilinear: Generalized multimodal factorized high-order pooling for visual question answering,” IEEE Transactions on Neural Networks and Learning Systems , vol. 29, no. 12, pp. 5947–5959, 2018
2018
Cited alongside, same era.
2018
Cited alongside, same era.
A. Singh, V. Natarajan, Y. Jiang, X. Chen, M. Shah, M. Rohrbach, D. Batra, and D. Parikh, “Pythia-a platform for vision & language research,” in SysML Workshop, NeurIPS , vol. 2018, 2018
2018
Cited alongside, same era.
A. Agrawal, D. Batra, D. Parikh, and A. Kembhavi, “Don’t just assume; look and answer: Overcoming priors for visual question answering,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2018, pp. 4971–4980
2018
Cited alongside, same era.
I. Loshchilov and F. Hutter, “Decoupled weight decay regularization,” in International Conference on Learning Representations , 2018
2018
Cited alongside, same era.
S.-M. Hu, J.-X. Cai, and Y.-K. Lai, “Semantic labeling and instance segmentation of 3d point clouds using patch context analysis and multiscale processing,” IEEE transactions on visualization and computer graphics , vol. 26, no. 7, pp. 2485–2498, 2018
2018
Cited alongside, same era.
D. Gurari, Q. Li, A. J. Stangl, A. Guo, C. Lin, K. Grauman, J. Luo, and J. P. Bigham, “Vizwiz grand challenge: Answering visual questions from blind people,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2018, pp. 3608–3617
2018
Cited alongside, same era.
D. Hudson and C. Manning, “Gqa: A new dataset for real-world visual reasoning and compositional question answering,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 06 2019, pp. 6693–6702
2019
Cited alongside, same era.
S. Amizadeh, H. Palangi, A. Polozov, Y. Huang, and K. Koishida, “Neuro-symbolic visual reasoning: Disentangling,” in International Conference on Machine Learning . PMLR, 2020, pp. 279–290
2020
Later among the works it cites.
H. Chen, M. Wei, Y. Sun, X. Xie, and J. Wang, “Multi-patch collaborative point cloud denoising via low-rank recovery with graph constraint,” IEEE Transactions on Visualization and Computer Graphics , vol. 26, no. 11, pp. 3255–3270, 2020
2020
Later among the works it cites.
A. Goyal, K. Yang, D. Yang, and J. Deng, “Rel3d: A minimally contrastive benchmark for grounding spatial relations in 3d,” Advances in Neural Information Processing Systems , vol. 33, 2020
2020
Later among the works it cites.
S.-H. Chou, W.-L. Chao, W.-S. Lai, M. Sun, and M.-H. Yang, “Visual question answering on 360deg images,” in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision , 2020, pp. 1607–1616
2020
Later among the works it cites.
S. Lobry, D. Marcos, J. Murray, and D. Tuia, “Rsvqa: Visual question answering for remote sensing data,” IEEE Transactions on Geoscience and Remote Sensing , vol. 58, no. 12, pp. 8555–8566, 2020
2020
Later among the works it cites.
A. Bansal, Y. Zhang, and R. Chellappa, “Visual question answering on image sets,” in European Conference on Computer Vision . Springer, 2020, pp. 51–67
2020
Later among the works it cites.
S. Chou, W. Chao, W. Lai, M. Sun, and M. Yang, “Visual question answering on 360 images,” in 2020 IEEE Winter Conference on Applications of Computer Vision (WACV) , 2020, pp. 1596–1605. [Online]. Available: https://doi.ieeecomputersociety.org/10.1109/WACV45572.2020.9093452
2020
Later among the works it cites.
C. Kervadec, G. Antipov, M. Baccouche, and C. Wolf, “Roses are red, violets are blue… but should vqa expect them to?” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2021, pp. 1, 2, 3, 4, 7, 8
2021
Closest in time.
M. Mathew, D. Karatzas, and C. Jawahar, “Docvqa: A dataset for vqa on document images,” in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision , 2021, pp. 2200–2209
2021
Closest in time.
S. Ye, D. Chen, S. Han, and J. Liao, “Learning with noisy labels for robust point cloud segmentation,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 6443–6452
2021
Closest in time.
Z. Chen, A. Gholami, M. Niessner, and A. X. Chang, “Scan2cap: Context-aware dense captioning in rgb-d scans,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , June 2021, pp. 3193–3203
2021
Closest in time.
Z. Yuan, X. Yan, Y. Liao, R. Zhang, S. Wang, Z. Li, and S. Cui, “Instancerefer: Cooperative holistic understanding for visual grounding on point clouds through instance multi-level contextual referring,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 1791–1800
2021
Closest in time.
P.-H. Huang, H.-H. Lee, H.-T. Chen, and T.-L. Liu, “Text-guided graph neural networks for referring 3d instance segmentation,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 35, 2021, pp. 1610–1618
2021
Closest in time.
2021
Closest in time.
Z. Yang, S. Zhang, L. Wang, and J. Luo, “Sat: 2d semantics assisted training for 3d visual grounding,” ICCV , 2021
2021
Closest in time.
2021
Closest in time.
W. Chen, Z. Gan, L. Li, Y. Cheng, W. Wang, and J. Liu, “Meta module network for compositional visual reasoning,” in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision , 2021, pp. 655–664
2021
Closest in time.
X. Li, R. Li, G. Chen, C.-W. Fu, D. Cohen-Or, and P.-A. Heng, “A rotation-invariant framework for deep point cloud analysis,” IEEE Transactions on Visualization and Computer Graphics , 2021
2021
Closest in time.
S. Ye, D. Chen, S. Han, Z. Wan, and J. Liao, “Meta-pu: An arbitrary-scale upsampling network for point cloud,” IEEE Transactions on Visualization and Computer Graphics , pp. 1–1, 2021
2021
Closest in time.
2021
Closest in time.
M. Mathew, D. Karatzas, and C. Jawahar, “Docvqa: A dataset for vqa on document images,” in WACV , 2021, pp. 2200–2209
2021
Closest in time.
H. Yun, Y. Yu, W. Yang, K. Lee, and G. Kim, “Pano-avqa: Grounded audio-visual question answering on 360 ∘ videos,” in ICCV , 2021
2021
Closest in time.
F. Han, S. Ye, M. He, M. Chai, and J. Liao, “Exemplar-based 3d portrait stylization,” IEEE Transactions on Visualization and Computer Graphics , 2021
2021
Closest in time.
L. Li, H. Fu, and M. Ovsjanikov, “Wsdesc: Weakly supervised 3d local descriptor learning for point cloud registration,” IEEE Transactions on Visualization and Computer Graphics , pp. 1–1, 2022
2022
Closest in time.
T. Jaunet, C. Kervadec, R. Vuillemot, G. Antipov, M. Baccouche, and C. Wolf, “Visqa: X-raying vision and language reasoning in transformers,” IEEE Transactions on Visualization and Computer Graphics , vol. 28, no. 01, pp. 976–986, jan 2022
2022
Closest in time.