Fetching the paper…
Reading the bibliography…
Data Augmentation (DA) -- generating extra training samples beyond original training set -- has been widely-used in today's unbiased VQA models to mitigate the language biases.
Antol, S., Agrawal, A., Lu, J., Mitchell, M., Batra, D., Zitnick, C.L., Parikh, D.: Vqa: Visual question answering. In: ICCV. pp. 2425–2433 (2015)
2015
Earlier work this paper cites.
Geman, D., Geman, S., Hallonquist, N., Younes, L.: Visual turing test for computer vision systems. PNAS pp. 3618–3623 (2015)
2015
Earlier work this paper cites.
Kingma, D.P., Ba, J.: Adam: A method for stochastic optimization. In: ICLR (2015)
2015
Earlier work this paper cites.
Ren, S., He, K., Girshick, R., Sun, J.: Faster r-cnn: Towards real-time object detection with region proposal networks. In: NeurIPS. pp. 91–99 (2015)
2015
Earlier work this paper cites.
Zhu, Y., Kiros, R., Zemel, R., Salakhutdinov, R., Urtasun, R., Torralba, A., Fidler, S.: Aligning books and movies: Towards story-like visual explanations by watching movies and reading books. In: ICCV. pp. 19–27 (2015)
2015
Earlier work this paper cites.
Agrawal, A., Batra, D., Parikh, D.: Analyzing the behavior of visual question answering models. In: EMNLP (2016)
2016
Earlier work this paper cites.
Zhang, P., Goyal, Y., Summers-Stay, D., Batra, D., Parikh, D.: Yin and yang: Balancing and answering binary visual questions. In: CVPR (2016)
2016
Earlier work this paper cites.
Chen, G., Choi, W., Yu, X., Han, T., Chandraker, M.: Learning efficient object detection models with knowledge distillation. NeurIPS (2017)
2017
Earlier work this paper cites.
Chen, L., Zhang, H., Xiao, J., Nie, L., Shao, J., Liu, W., Chua, T.S.: Sca-cnn: Spatial and channel-wise attention in convolutional networks for image captioning. In: CVPR. pp. 5659–5667 (2017)
2017
Earlier work this paper cites.
Goyal, Y., Khot, T., Summers-Stay, D., Batra, D., Parikh, D.: Making the v in vqa matter: Elevating the role of image understanding in visual question answering. In: CVPR. pp. 6904–6913 (2017)
2017
Earlier work this paper cites.
Honnibal, M., Montani, I.: spacy 2: Natural language understanding with bloom embeddings, convolutional neural networks and incremental parsing. To appear (2017)
2017
Earlier work this paper cites.
Johnson, J., Hariharan, B., van der Maaten, L., Fei-Fei, L., Lawrence Zitnick, C., Girshick, R.: Clevr: A diagnostic dataset for compositional language and elementary visual reasoning. In: CVPR (2017)
2017
Earlier work this paper cites.
Kafle, K., Yousefhussien, M., Kanan, C.: Data augmentation for visual question answering. In: INLG. pp. 198–202 (2017)
2017
Earlier work this paper cites.
Krishna, R., Zhu, Y., Groth, O., Johnson, J., Hata, K., Kravitz, J., Chen, S., Kalantidis, Y., Li, L.J., Shamma, D.A., et al.: Visual genome: Connecting language and vision using crowdsourced dense image annotations. In: IJCV. pp. 32–73 (2017)
2017
Earlier work this paper cites.
Agrawal, A., Batra, D., Parikh, D., Kembhavi, A.: Don’t just assume; look and answer: Overcoming priors for visual question answering. In: CVPR (2018)
2018
Earlier work this paper cites.
Anderson, P., He, X., Buehler, C., Teney, D., Johnson, M., Gould, S., Zhang, L.: Bottom-up and top-down attention for image captioning and visual question answering. In: CVPR (2018)
2018
Earlier work this paper cites.
Radosavovic, I., Dollár, P., Girshick, R., Gkioxari, G., He, K.: Data distillation: Towards omni-supervised learning. In: CVPR. pp. 4119–4128 (2018)
2018
Earlier work this paper cites.
Ramakrishnan, S., Agrawal, A., Lee, S.: Overcoming language priors in visual question answering with adversarial regularization. In: NeurIPS (2018)
2018
Earlier work this paper cites.
Cadene, R., Ben-Younes, H., Cord, M., Thome, N.: Murel: Multimodal relational reasoning for visual question answering. In: CVPR (2019)
2019
Earlier work this paper cites.
Cadene, R., Dancette, C., Ben-younes, H., Cord, M., Parikh, D.: Rubi: Reducing unimodal biases in visual question answering. In: NeurIPS (2019)
2019
Earlier work this paper cites.
Clark, C., Yatskar, M., Zettlemoyer, L.: Don’t take the easy way out: Ensemble based methods for avoiding known dataset biases. In: EMNLP (2019)
2019
Cited alongside, same era.
Devlin, J., Chang, M., Lee, K., Toutanova, K.: BERT: pre-training of deep bidirectional transformers for language understanding. In: NAACL. pp. 4171–4186 (2019)
2019
Cited alongside, same era.
Grand, G., Belinkov, Y.: Adversarial regularization for visual question answering: Strengths, shortcomings, and side effects. In: ACLW (2019)
2019
Cited alongside, same era.
Lu, C., Chen, L., Tan, C., Li, X., Xiao, J.: Debug: A dense bottom-up grounding approach for natural language video localization. In: EMNLP. pp. 5144–5153 (2019)
2019
Cited alongside, same era.
Wang, T., Yuan, L., Zhang, X., Feng, J.: Distilling object detectors with fine-grained feature imitation. In: CVPR. pp. 4933–4942 (2019)
Chen, L., Ma, W., Xiao, J., Zhang, H., Chang, S.F.: Ref-nms: Breaking proposal bottlenecks in two-stage referring expression grounding. In: AAAI. pp. 1036–1044 (2021)
2021
Later among the works it cites.
Chen, L., Zheng, Y., Niu, Y., Zhang, H., Xiao, J.: Counterfactual samples synthesizing and training for robust visual question answering. arXiv (2021)
2021
Later among the works it cites.
Han, X., Wang, S., Su, C., Huang, Q., Tian, Q.: Greedy gradient ensemble for robust visual question answering. In: ICCV (2021)
2021
Later among the works it cites.
Kant, Y., Moudgil, A., Batra, D., Parikh, D., Agrawal, H.: Contrast and classify: Training robust vqa models. In: ICCV. pp. 1604–1613 (2021)
2021
Later among the works it cites.
Kil, J., Zhang, C., Xuan, D., Chao, W.L.: Discovering the unknown knowns: Turning implicit knowledge in the dataset into explicit training examples for visual question answering. In: EMNLP (2021)
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2019
Cited alongside, same era.
Abbasnejad, E., Teney, D., Parvaneh, A., Shi, J., Hengel, A.v.d.: Counterfactual vision and language learning. In: CVPR (2020)
2020
Cited alongside, same era.
Agarwal, V., Shetty, R., Fritz, M.: Towards causal vqa: Reveling and reducing spurious correlations by invariant and covariant semantic editing. In: CVPR (2020)
2020
Cited alongside, same era.
Chen, L., Lu, C., Tang, S., Xiao, J., Zhang, D., Tan, C., Li, X.: Rethinking the bottom-up framework for query-based video localization. In: AAAI. pp. 10551–10558 (2020)
2020
Cited alongside, same era.
Chen, L., Yan, X., Xiao, J., Zhang, H., Pu, S., Zhuang, Y.: Counterfactual samples synthesizing for robust visual question answering. In: CVPR. pp. 10800–10809 (2020)
2020
Cited alongside, same era.
Gokhale, T., Banerjee, P., Baral, C., Yang, Y.: Mutant: A training paradigm for out-of-distribution generalization in visual question answering. In: EMNLP (2020)
2020
Cited alongside, same era.
Gokhale, T., Banerjee, P., Baral, C., Yang, Y.: Vqa-lol: Visual question answering under the lens of logic. In: ECCV. pp. 379–396 (2020)
2020
Cited alongside, same era.
Liang, Z., Jiang, W., Hu, H., Zhu, J.: Learning to contrast the counterfactual samples for robust visual question answering. In: EMNLP (2020)
2020
Cited alongside, same era.
2021
Later among the works it cites.
Liang, Z., Hu, H., Zhu, J.: Lpf: A language-prior feedback objective function for de-biased visual question answering. In: ACM SIGIR. pp. 1955–1959 (2021)
2021
Later among the works it cites.
Niu, Y., Tang, K., Zhang, H., Lu, Z., Hua, X.S., Wen, J.R.: Counterfactual vqa: A cause-effect look at language bias. In: CVPR (2021)
2021
Later among the works it cites.
Niu, Y., Zhang, H.: Introspective distillation for robust question answering. In: NeurIPS (2021)
2021
Later among the works it cites.
Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., Sutskever, I.: Learning transferable visual models from natural language supervision. In: ICML. pp. 8748–8763 (2021)
2021
Later among the works it cites.
Teney, D., Abbasnejad, E., Hengel, A.v.d.: Unshuffling data for improved generalization. In: ICCV (2021)
2021
Later among the works it cites.
Wang, L., Yoon, K.J.: Knowledge distillation and student-teacher learning for visual intelligence: A review and new outlooks. IEEE TPAMI (2021)
2021
Later among the works it cites.
Wang, Z., Miao, Y., Specia, L.: Cross-modal generative augmentation for visual question answering. In: BMVC (2021)
2021
Later among the works it cites.
Wen, Z., Xu, G., Tan, M., Wu, Q., Wu, Q.: Debiased visual question answering from feature and sample perspectives. In: NeurIPS (2021)
2021
Later among the works it cites.
Xiao, S., Chen, L., Zhang, S., Ji, W., Shao, J., Ye, L., Xiao, J.: Boundary proposal network for two-stage natural language video localization. In: AAAI. pp. 2986–2994 (2021)
2021
Later among the works it cites.
Askarian, N., Abbasnejad, E., Zukerman, I., Buntine, W., Haffari, G.: Inductive biases for low data vqa: A data augmentation approach. In: WACV. pp. 231–240 (2022)
2022
Closest in time.
Boukhers, Z., Hartmann, T., Jürjens, J.: Coin: Counterfactual image generation for vqa interpretation. In: arXiv (2022)
2022
Closest in time.
Kolling, C., More, M., Gavenski, N., Pooch, E., Parraga, O., Barros, R.C.: Efficient counterfactual debiasing for visual question answering. In: WACV. pp. 3001–3010 (2022)
2022
Closest in time.
Li, X., Chen, L., Ma, W., Yang, Y., Xiao, J.: Integrating object-aware and interaction-aware knowledge for weakly supervised scene graph generation. In: ACM MM (2022)
2022
Closest in time.
Mao, Y., Chen, L., Jiang, Z., Zhang, D., Zhang, Z., Shao, J., Xiao, J.: Rethinking the reference-based distinctive image captioning. In: ACM MM (2022)
2022
Closest in time.