Fetching the paper…
Reading the bibliography…
Visual question answering (VQA) is a challenging task, which has attracted more and more attention in the field of computer vision and natural language processing.
D. Teney and A. van den Hengel, “Actively seeking and learning from live data,” in 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2019, pp. 1940–1949
1949
Earlier work this paper cites.
W. Li, Q. Wu, L. Xu, and C. Shang, “Incremental learning of single-stage detectors with mining memory neurons,” in 2018 IEEE 4th International Conference on Computer and Communications (ICCC) . IEEE, 2018, pp. 1981–1985
1985
Earlier work this paper cites.
P. W. Holland, “Statistics and causal inference,” Journal of the American statistical Association , vol. 81, no. 396, pp. 945–960, 1986
1986
Earlier work this paper cites.
R. Cadene, H. Ben-younes, M. Cord, and N. Thome, “Murel: Multimodal relational reasoning for visual question answering,” in 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . IEEE Computer Society, 2019, pp. 1989–1998
1998
Earlier work this paper cites.
R. Cowie, E. Douglas-Cowie, N. Tsapatsoulis, G. Votsis, S. Kollias, W. Fellenz, and J. G. Taylor, “Emotion recognition in human-computer interaction,” IEEE Signal processing magazine , vol. 18, no. 1, pp. 32–80, 2001
2001
Earlier work this paper cites.
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu, “Bleu: a method for automatic evaluation of machine translation,” in Proceedings of the 40th annual meeting of the Association for Computational Linguistics , 2002, pp. 311–318
2002
Earlier work this paper cites.
J. R. Curran and S. Clark, “Language independent ner using a maximum entropy tagger,” in Proceedings of the seventh conference on Natural language learning at HLT-NAACL 2003 , 2003, pp. 164–167
2003
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in 2009 IEEE conference on computer vision and pattern recognition . Ieee, 2009, pp. 248–255
2009
Earlier work this paper cites.
J. Pearl, “Causal inference in statistics: An overview,” Statistics surveys , vol. 3, pp. 96–146, 2009
2009
Earlier work this paper cites.
M. A. Hernán and J. M. Robins, “Causal inference,” 2010
2010
Earlier work this paper cites.
——, “Causal inference,” Causality: Objectives and Assessment , pp. 39–58, 2010
2010
Earlier work this paper cites.
T. Song, H. Li, F. Meng, Q. Wu, B. Luo, B. Zeng, and M. Gabbouj, “Noise-robust texture description using local contrast patterns via global measures,” IEEE Signal Processing Letters , vol. 21, no. 1, pp. 93–96, 2013
2013
Earlier work this paper cites.
F. Meng, H. Li, K. N. Ngan, L. Zeng, and Q. Wu, “Feature adaptive co-segmentation by complexity awareness,” IEEE Transactions on Image Processing , vol. 22, no. 12, pp. 4809–4824, 2013
2013
Earlier work this paper cites.
2014
Earlier work this paper cites.
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. Lawrence Zitnick, and D. Parikh, “Vqa: Visual question answering,” in ICCV , 2015
2015
Earlier work this paper cites.
D. Tran, L. Bourdev, R. Fergus, L. Torresani, and M. Paluri, “Learning spatio temporal features with 3d convolutional networks,” in Proceedings of the IEEE international conference on computer vision , 2015, pp. 4489–4497
2015
Earlier work this paper cites.
Q. Wu, H. Li, F. Meng, K. N. Ngan, and S. Zhu, “No reference image quality assessment metric via multi-domain structural information and piecewise regression,” Journal of Visual Communication and Image Representation , vol. 32, pp. 205–216, 2015
2015
Earlier work this paper cites.
S. L. Morgan and C. Winship, Counterfactuals and causal inference . Cambridge University Press, 2015
2015
Earlier work this paper cites.
S. Ren, K. He, R. Girshick, and J. Sun, “Faster r-cnn: towards real-time object detection with region proposal networks,” IEEE transactions on pattern analysis and machine intelligence , vol. 39, no. 6, pp. 1137–1149, 2016
2016
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 770–778
2016
Earlier work this paper cites.
K. Ma, Z. Duanmu, Q. Wu, Z. Wang, H. Yong, H. Li, and L. Zhang, “Waterloo exploration database: New challenges for image quality assessment models,” IEEE Transactions on Image Processing , vol. 26, no. 2, pp. 1004–1016, 2016
2016
Earlier work this paper cites.
K. Ma, Q. Wu, Z. Wang, Z. Duanmu, H. Yong, H. Li, and L. Zhang, “Group mad competition-a new methodology to compare objective image quality models,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2016, pp. 1664–1673
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
Z. Yang, X. He, J. Gao, L. Deng, and A. Smola, “Stacked attention networks for image question answering,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 21–29
2016
Earlier work this paper cites.
W. Ouyang, X. Wang, C. Zhang, and X. Yang, “Factors in finetuning deep model for object detection with long-tail distribution,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 864–873
2016
Earlier work this paper cites.
J. Pearl, M. Glymour, and N. P. Jewell, Causal inference in statistics: A primer . John Wiley & Sons, 2016
2016
Earlier work this paper cites.
A. Agrawal, J. Lu, S. Antol, M. Mitchell, C. L. Zitnick, D. Parikh, and D. Batra, “Vqa: Visual question answering,” International Journal of Computer Vision , vol. 123, no. 1, pp. 4–31, 2017
2017
Earlier work this paper cites.
K. He, G. Gkioxari, P. Dollár, and R. Girshick, “Mask r-cnn,” in Proceedings of the IEEE international conference on computer vision , 2017, pp. 2961–2969
2017
Earlier work this paper cites.
T.-Y. Lin, P. Dollár, R. Girshick, K. He, B. Hariharan, and S. Belongie, “Feature pyramid networks for object detection,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2017, pp. 2117–2125
2017
Earlier work this paper cites.
T.-Y. Lin, P. Goyal, R. Girshick, K. He, and P. Dollár, “Focal loss for dense object detection,” in Proceedings of the IEEE international conference on computer vision , 2017, pp. 2980–2988
2017
Earlier work this paper cites.
J. Carreira and A. Zisserman, “Quo vadis, action recognition? a new model and the kinetics dataset,” in Computer Vision & Pattern Recognition , 2017
2017
Earlier work this paper cites.
T. Song, H. Li, F. Meng, Q. Wu, and J. Cai, “Letrist: Locally encoded transform feature histogram for rotation-invariant texture classification,” IEEE Transactions on circuits and systems for video technology , vol. 28, no. 7, pp. 1565–1579, 2017
2017
Earlier work this paper cites.
Q. Wu, H. Li, Z. Wang, F. Meng, B. Luo, W. Li, and K. N. Ngan, “Blind image quality assessment based on rank-order regularized regression,” IEEE Transactions on Multimedia , vol. 19, no. 11, pp. 2490–2504, 2017
2017
Earlier work this paper cites.
Q. Wu, H. Li, K. N. Ngan, and K. Ma, “Blind image quality assessment using local consistency aware retriever and uncertainty aware evaluator,” IEEE Transactions on Circuits and Systems for Video Technology , vol. 28, no. 9, pp. 2078–2089, 2017
2017
Earlier work this paper cites.
D. Tran, H. Wang, L. Torresani, J. Ray, Y. LeCun, and M. Paluri, “A closer look at spatiotemporal convolutions for action recognition,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2018, pp. 6450–6459
2018
Earlier work this paper cites.
K. Luo, F. Meng, Q. Wu, and H. Li, “Weakly supervised semantic segmentation by multiple group cosegmentation,” in 2018 IEEE Visual Communications and Image Processing (VCIP) . IEEE, 2018, pp. 1–4
2018
Earlier work this paper cites.
H. Shi, H. Li, F. Meng, and Q. Wu, “Key-word-aware network for referring expression image segmentation,” in Proceedings of the European Conference on Computer Vision (ECCV) , 2018, pp. 38–54
2018
Earlier work this paper cites.
A. Agrawal, D. Batra, D. Parikh, and A. Kembhavi, “Don’t just assume; look and answer: Overcoming priors for visual question answering,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2018, pp. 4971–4980
2018
Cited alongside, same era.
P. Anderson, X. He, C. Buehler, D. Teney, M. Johnson, S. Gould, and L. Zhang, “Bottom-up and top-down attention for image captioning and visual question answering,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2018, pp. 6077–6086
2018
Cited alongside, same era.
S. Ramakrishnan, A. Agrawal, and S. Lee, “Overcoming language priors in visual question answering with adversarial regularization,” Advances in Neural Information Processing Systems , vol. 31, 2018
2018
Cited alongside, same era.
Y. Zhou, Q. Hu, and Y. Wang, “Deep super-class learning for long-tail distributed image classification,” Pattern Recognition , vol. 80, pp. 118–128, 2018
2018
Cited alongside, same era.
Q. Wu, L. Wang, K. N. Ngan, H. Li, F. Meng, and L. Xu, “Subjective and objective de-raining quality assessment towards authentic rain image,” IEEE Transactions on Circuits and Systems for Video Technology , vol. 30, no. 11, pp. 3883–3897, 2020
2020
Later among the works it cites.
H. Li, Q. Wu, H. Wei, K. N. Ngan, H. Li, F. Meng, and L. Xu, “Haze-robust image understanding via context-aware deep feature refinement,” in 2020 IEEE 22nd International Workshop on Multimedia Signal Processing (MMSP) . IEEE, 2020, pp. 1–6
2020
Later among the works it cites.
Q. Wu, L. Chen, K. N. Ngan, H. Li, F. Meng, and L. Xu, “A unified single image de-raining model via region adaptive coupled network,” in 2020 IEEE International Conference on Visual Communications and Image Processing (VCIP) . IEEE, 2020, pp. 1–4
2020
Later among the works it cites.
L. Peng, Y. Yang, X. Zhang, Y. Ji, H. Lu, and H. T. Shen, “Answer again: Improving vqa with cascaded-answering model,” IEEE Transactions on Knowledge and Data Engineering , 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
W. Li, H. Li, Q. Wu, X. Chen, and K. N. Ngan, “Simultaneously detecting and counting dense vehicles from drone images,” IEEE Transactions on Industrial Electronics , vol. 66, no. 12, pp. 9651–9662, 2019
2019
Cited alongside, same era.
W. Li, H. Li, Q. Wu, F. Meng, L. Xu, and K. N. Ngan, “Headnet: An end-to-end adaptive relational network for head detection,” IEEE Transactions on Circuits and Systems for Video Technology , vol. 30, no. 2, pp. 482–494, 2019
2019
Cited alongside, same era.
H. Qiu, H. Li, Q. Wu, F. Meng, K. N. Ngan, and H. Shi, “A2rmnet: Adaptively aspect ratio multi-scale network for object detection in remote sensing images,” Remote Sensing , vol. 11, no. 13, p. 1594, 2019
2019
Cited alongside, same era.
F. Meng, L. Guo, Q. Wu, and H. Li, “A new deep segmentation quality assessment network for refining bounding box based segmentation,” IEEE Access , vol. 7, pp. 59 514–59 523, 2019
2019
Cited alongside, same era.
Y. Yang, F. Meng, H. Li, K. N. Ngan, and Q. Wu, “A new few-shot segmentation network based on class representation,” in 2019 IEEE Visual Communications and Image Processing (VCIP) . IEEE, 2019, pp. 1–4
2019
Cited alongside, same era.
X. Xu, F. Meng, H. Li, Q. Wu, Y. Yang, and S. Chen, “Bounding box based annotation generation for semantic segmentation by boundary detection,” in 2019 International Symposium on Intelligent Signal Processing and Communication Systems (ISPACS) . IEEE, 2019, pp. 1–2
2019
Cited alongside, same era.
C. Shang, Q. Wu, F. Meng, and L. Xu, “Instance segmentation by learning deep feature in embedding space,” in 2019 IEEE International Conference on Image Processing (ICIP) . IEEE, 2019, pp. 2444–2448
2019
Cited alongside, same era.
Q. Wu, L. Wang, K. N. Ngan, H. Li, and F. Meng, “Beyond synthetic data: A blind deraining quality assessment metric towards authentic rain image,” in 2019 IEEE International Conference on Image Processing (ICIP) . IEEE, 2019, pp. 2364–2368
2019
Cited alongside, same era.
2020
Later among the works it cites.
R. Shrestha, K. Kafle, and C. Kanan, “A negative case analysis of visual grounding methods for vqa,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , 2020, pp. 8172–8181
2020
Later among the works it cites.
G. KV and A. Mittal, “Reducing language biases in visual question answering with visually-grounded question encoder,” in Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XIII 16 . Springer, 2020, pp. 18–34
2020
Later among the works it cites.
L. Zhang, S. Liu, D. Liu, P. Zeng, X. Li, J. Song, and L. Gao, “Rich visual knowledge-based augmentation network for visual question answering,” IEEE Transactions on Neural Networks and Learning Systems , 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
I. Gat, I. Schwartz, A. G. Schwing, and T. Hazan, “Removing bias in multi-modal classifiers: Regularization by maximizing functional entropies,” in Advances in Neural Information Processing Systems , vol. 33, 2020, pp. 3197–3208
2020
Later among the works it cites.
L. Chen, X. Yan, J. Xiao, H. Zhang, S. Pu, and Y. Zhuang, “Counterfactual samples synthesizing for robust visual question answering,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2020, pp. 10 800–10 809
2020
Later among the works it cites.
Z. Liang, W. Jiang, H. Hu, and J. Zhu, “Learning to contrast the counterfactual samples for robust visual question answering,” in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) , 2020, pp. 3285–3292
2020
Later among the works it cites.
D. Teney, E. Abbasnedjad, and A. van den Hengel, “Learning what makes a difference from counterfactual examples and gradient supervision,” in Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part X 16 . Springer, 2020, pp. 580–599
2020
Later among the works it cites.
2020
Later among the works it cites.
T. Gokhale, P. Banerjee, C. Baral, and Y. Yang, “Mutant: A training paradigm for out-of-distribution generalization in visual question answering,” in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) , 2020, pp. 878–892
2020
Later among the works it cites.
D. Teney, E. Abbasnejad, K. Kafle, R. Shrestha, C. Kanan, and A. van den Hengel, “On the value of out-of-distribution testing: An example of goodhart's law,” in Advances in Neural Information Processing Systems , H. Larochelle, M. Ranzato, R. Hadsell, M. F. Balcan, and H. Lin, Eds., vol. 33. Curran Associates, Inc., 2020, pp. 407–417. [Online]. Available: https://proceedings.neurips.cc/paper/2020/file/045117b0e0a11a242b9765e79cbf113f-Paper.pdf
2020
Later among the works it cites.
X. Zhu, Z. Mao, C. Liu, P. Zhang, B. Wang, and Y. Zhang, “Overcoming language priors with self-supervised learning for visual question answering,” in Proceedings of the Twenty-Ninth International Joint Conference on Artificial Intelligence, IJCAI-20 , C. Bessiere, Ed. International Joint Conferences on Artificial Intelligence Organization, 2020, pp. 1083–1089
2020
Later among the works it cites.
M. Ziaeefard and F. Lécué, “Towards knowledge-augmented visual question answering,” in Proceedings of the 28th International Conference on Computational Linguistics , 2020, pp. 1863–1873
2020
Later among the works it cites.
A. K. Menon, S. Jayasumana, A. S. Rawat, H. Jain, A. Veit, and S. Kumar, “Long-tail learning via logit adjustment,” in International Conference on Learning Representations , 2020
2020
Later among the works it cites.
W. Wang, M. Wang, S. Wang, G. Long, L. Yao, G. Qi, and Y. Chen, “One-shot learning for long-tail visual relation detection,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 34, no. 07, 2020, pp. 12 225–12 232
2020
Later among the works it cites.
X. Chen, H. Li, Q. Wu, F. Meng, and H. Qiu, “Bal-r2cnn: High quality recurrent object detection with balance optimization,” IEEE Transactions on Multimedia , 2021
2021
Closest in time.
H. Luo, H. Luo, Q. Wu, K. N. Ngan, H. Li, F. Meng, and L. Xu, “Single image deraining via multi-scale gated feature enhancement network,” in Digital TV and Wireless Multimedia Communication: 17th International Forum, IFTC 2020, Shanghai, China, December 2, 2020: Revised Selected Papers , vol. 1390. Springer Nature, 2021, p. 73
2021
Closest in time.
2021
Closest in time.
H. Li, Q. Wu, K. N. Ngan, H. Li, and F. Meng, “Single image dehazing via region adaptive two-shot network,” IEEE MultiMedia , 2021
2021
Closest in time.
2021
Closest in time.
Q. Si, Z. Lin, M. y. Zheng, P. Fu, and W. Wang, “Check it again:progressive visual question answering via visual entailment,” in Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers) . Online: Association for Computational Linguistics, Aug. 2021, pp. 4101–4110. [Online]. Available: https://aclanthology.org/2021.acl-long.317
2021
Closest in time.
Y. Niu, K. Tang, H. Zhang, Z. Lu, X.-S. Hua, and J.-R. Wen, “Counterfactual vqa: A cause-effect look at language bias,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2021, pp. 12 700–12 710
2021
Closest in time.
X. Han, S. Wang, C. Su, Q. Huang, and Q. Tian, “Greedy gradient ensemble for robust visual question answering,” in ICCV , 2021
2021
Closest in time.
Z. Liang, H. Hu, and J. Zhu, “Lpf: A language-prior feedback objective function for de-biased visual question answering,” in Proceedings of the 44th International Conference on Research and Development in Information Retrieval (SIGIR) , 2021
2021
Closest in time.
D. Teney, E. Abbasnejad, and A. van den Hengel, “Unshuffling data for improved generalization in visual question answering,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 1417–1427
2021
Closest in time.
M. Lao, Y. Guo, Y. Liu, and M. S. Lew, “A language prior based focal loss for visual question answering,” in 2021 IEEE International Conference on Multimedia and Expo (ICME) . IEEE, 2021, pp. 1–6
2021
Closest in time.
Y. Guo, L. Nie, Z. Cheng, F. Ji, J. Zhang, and A. Del Bimbo, “Adavqa: Overcoming language priors with adapted margin cosine loss,” in Proceedings of the Thirtieth International Joint Conference on Artificial Intelligence, IJCAI-21 , Z.-H. Zhou, Ed. International Joint Conferences on Artificial Intelligence Organization, 2021, pp. 708–714
2021
Closest in time.
C. Yang, S. Feng, D. Li, H. Shen, G. Wang, and B. Jiang, “Learning content and context with language bias for visual question answering,” in 2021 IEEE International Conference on Multimedia and Expo (ICME) . IEEE, 2021, pp. 1–6
2021
Closest in time.
N. Ouyang, Q. Huang, P. Li, C. Yi, B. Liu, H.-f. Leung, and Q. Li, “Suppressing biased samples for robust vqa,” IEEE Transactions on Multimedia , 2021
2021
Closest in time.
P. Banerjee, T. Gokhale, Y. Yang, and C. Baral, “Weaqa: Weak supervision via captions for visual question answering,” in Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021 , 2021, pp. 3420–3435
2021
Closest in time.
J. Jiang, Z. Liu, Y. Liu, Z. Nan, and N. Zheng, “X-ggm: Graph generative modeling for out-of-distribution generalization in visual question answering,” in Proceedings of the 29th ACM International Conference on Multimedia , 2021, pp. 199–208
2021
Closest in time.
Q. Cao, B. Li, X. Liang, K. Wang, and L. Lin, “Knowledge-routed visual question reasoning: Challenges for deep representation embedding,” IEEE Transactions on Neural Networks and Learning Systems , 2021
2021
Closest in time.
D. Yuan, X. Liu, Q. Wu, H. Li, F. Meng, K. N. Ngan, and L. Xu, “Empower counterfactual thinking via contrastive learning for robust visual question answering,” in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process. (ICASSP) . IEEE, 2022, p. under review
2022
Closest in time.