Fetching the paper…
Reading the bibliography…
Existing visual question reasoning methods usually fail to explicitly discover the inherent causal mechanism and ignore jointly modeling cross-modal event temporality and causality.
Bareinboim, E., Pearl, J.: Controlling selection bias in causal inference. In: Artificial Intelligence and Statistics. pp. 100–108. PMLR (2012)
2012
Earlier work this paper cites.
Pennington, J., Socher, R., Manning, C.D.: Glove: Global vectors for word representation. In: Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP). pp. 1532–1543 (2014)
2014
Earlier work this paper cites.
Antol, S., Agrawal, A., Lu, J., Mitchell, M., Batra, D., Zitnick, C.L., Parikh, D.: Vqa: Visual question answering. In: Proceedings of the IEEE international conference on computer vision. pp. 2425–2433 (2015)
2015
Earlier work this paper cites.
Sukhbaatar, S., Szlam, A., Weston, J., Fergus, R.: End-to-end memory networks. Advances in Neural Information Processing Systems 2015
2015
Earlier work this paper cites.
Xu, K., Ba, J., Kiros, R., Cho, K., Courville, A., Salakhudinov, R., Zemel, R., Bengio, Y.: Show, attend and tell: Neural image caption generation with visual attention. In: International conference on machine learning. pp. 2048–2057 (2015)
2015
Earlier work this paper cites.
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
Liu, Y., Li, J., Lu, Z., Yang, T., Liu, Z.: Combining multiple features for cross-domain face sketch recognition. In: Biometric Recognition: 11th Chinese Conference, CCBR 2016, Chengdu, China, October 14-16, 2016, Proceedings 11. pp. 139–146. Springer (2016)
2016
Earlier work this paper cites.
Pearl, J., Glymour, M., Jewell, N.P.: Causal inference in statistics: A primer. John Wiley & Sons (2016)
2016
Earlier work this paper cites.
Yang, Z., He, X., Gao, J., Deng, L., Smola, A.: Stacked attention networks for image question answering. In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 21–29 (2016)
2016
Earlier work this paper cites.
Das, A., Kottur, S., Gupta, K., Singh, A., Yadav, D., Moura, J.M., Parikh, D., Batra, D.: Visual dialog. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 326–335 (2017)
2017
Earlier work this paper cites.
Jang, Y., Song, Y., Yu, Y., Kim, Y., Kim, G.: Tgif-qa: Toward spatio-temporal reasoning in visual question answering. In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 2758–2766 (2017)
2017
Earlier work this paper cites.
Krishna, R., Hata, K., Ren, F., Fei-Fei, L., Carlos Niebles, J.: Dense-captioning events in videos. In: Proceedings of the IEEE international conference on computer vision. pp. 706–715 (2017)
2017
Earlier work this paper cites.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, Ł., Polosukhin, I.: Attention is all you need. In: Advances in neural information processing systems. pp. 5998–6008 (2017)
2017
Earlier work this paper cites.
Xu, D., Zhao, Z., Xiao, J., Wu, F., Zhang, H., He, X., Zhuang, Y.: Video question answering via gradually refined attention over appearance and motion. In: Proceedings of the 25th ACM international conference on Multimedia. pp. 1645–1653 (2017)
2017
Earlier work this paper cites.
Anderson, P., He, X., Buehler, C., Teney, D., Johnson, M., Gould, S., Zhang, L.: Bottom-up and top-down attention for image captioning and visual question answering. In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 6077–6086 (2018)
2018
Earlier work this paper cites.
Anderson, P., Wu, Q., Teney, D., Bruce, J., Johnson, M., Sünderhauf, N., Reid, I., Gould, S., Van Den Hengel, A.: Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 3674–3683 (2018)
2018
Earlier work this paper cites.
Chowdhury, M.I.H., Nguyen, K., Sridharan, S., Fookes, C.: Hierarchical relational attention for video question answering. In: 2018 25th IEEE International Conference on Image Processing (ICIP). pp. 599–603. IEEE (2018)
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
Gao, J., Ge, R., Chen, K., Nevatia, R.: Motion-appearance co-memory networks for video question answering. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 6576–6585 (2018)
2018
Earlier work this paper cites.
Hendricks, L.A., Hu, R., Darrell, T., Akata, Z.: Grounding visual explanations. In: Proceedings of the European Conference on Computer Vision (ECCV). pp. 264–279 (2018)
2018
Earlier work this paper cites.
Liu, Y., Lu, Z., Li, J., Yang, T.: Hierarchically learned view-invariant representations for cross-view action recognition. IEEE Transactions on Circuits and Systems for Video Technology 29
2018
Earlier work this paper cites.
Liu, Y., Lu, Z., Li, J., Yang, T., Yao, C.: Global temporal representation based cnns for infrared action recognition. IEEE Signal Processing Letters 25
2018
Earlier work this paper cites.
Liu, Y., Lu, Z., Li, J., Yao, C., Deng, Y.: Transferable feature representation for visible-to-infrared cross-dataset human action recognition. Complexity 2018
2018
Earlier work this paper cites.
Chen, L., Zhang, H., Xiao, J., He, X., Pu, S., Chang, S.F.: Counterfactual critic multi-agent training for scene graph generation. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 4613–4623 (2019)
2019
Earlier work this paper cites.
Fan, C., Zhang, X., Zhang, S., Wang, W., Zhang, C., Huang, H.: Heterogeneous memory enhanced multimodal attention model for video question answering. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. 1999–2007 (2019)
2019
Earlier work this paper cites.
Fang, Z., Kong, S., Fowlkes, C., Yang, Y.: Modularized textual grounding for counterfactual resilience. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 6378–6388 (2019)
2019
Earlier work this paper cites.
Goyal, Y., Wu, Z., Ernst, J., Batra, D., Parikh, D., Lee, S.: Counterfactual visual explanations. In: International Conference on Machine Learning. pp. 2376–2384. PMLR (2019)
2019
Earlier work this paper cites.
Kanehira, A., Takemoto, K., Inayoshi, S., Harada, T.: Multimodal explanations by predicting counterfactuality in videos. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 8594–8602 (2019)
2019
Cited alongside, same era.
LE, H., SAHOO, D., CHEN, N.F.: Choi. multimodal transformer networks for end-to-end video-grounded dialogue systems.(2019). In: Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, Florence, Italy, 2019 July 28-August. vol. 2, pp. 5612–5623
2019
Cited alongside, same era.
Li, X., Song, J., Gao, L., Liu, X., Huang, W., He, X., Gan, C.: Beyond rnns: Positional self-attention with co-attention for video question answering. In: Proceedings of the AAAI Conference on Artificial Intelligence. vol. 33, pp. 8658–8665 (2019)
2019
Cited alongside, same era.
Liu, Y., Lu, Z., Li, J., Yang, T., Yao, C.: Deep image-to-video adaptation and fusion networks for action recognition. IEEE Transactions on Image Processing 29
Seo, A., Kang, G.C., Park, J., Zhang, B.T.: Attend what you need: Motion-appearance synergistic networks for video question answering. In: Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). pp. 6167–6177 (2021)
2021
Later among the works it cites.
Wang, J., Bao, B., Xu, C.: Dualvgr: A dual-visual graph reasoning unit for video question answering. IEEE Transactions on Multimedia (2021)
2021
Later among the works it cites.
Wang, T., Zhou, C., Sun, Q., Zhang, H.: Causal attention for unbiased visual recognition. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 3091–3100 (2021)
2021
Later among the works it cites.
Xu, L., Huang, H., Liu, J.: Sutd-trafficqa: A question answering benchmark and an efficient network for video reasoning over traffic events. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 9878–9888 (2021)
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2019
Cited alongside, same era.
Abbasnejad, E., Teney, D., Parvaneh, A., Shi, J., Hengel, A.v.d.: Counterfactual vision and language learning. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 10044–10054 (2020)
2020
Cited alongside, same era.
Besserve, M., Mehrjou, A., Sun, R., Schölkopf, B.: Counterfactuals uncover the modular structure of deep generative models. In: Eighth International Conference on Learning Representations (ICLR 2020) (2020)
2020
Cited alongside, same era.
Cao, D., Zeng, Y., Liu, M., He, X., Wang, M., Qin, Z.: Strong: Spatio-temporal reinforcement learning for cross-modal video moment localization. In: Proceedings of the 28th ACM International Conference on Multimedia. pp. 4162–4170 (2020)
2020
Cited alongside, same era.
Huang, D., Chen, P., Zeng, R., Du, Q., Tan, M., Gan, C.: Location-aware graph convolutional networks for video question answering. In: Proceedings of the AAAI Conference on Artificial Intelligence. vol. 34, pp. 11021–11028 (2020)
2020
Cited alongside, same era.
Jiang, J., Chen, Z., Lin, H., Zhao, X., Gao, Y.: Divide and conquer: Question-guided spatio-temporal contextual attention for video question answering. In: Proceedings of the AAAI Conference on Artificial Intelligence. vol. 34, pp. 11101–11108 (2020)
2020
Cited alongside, same era.
Jiang, P., Han, Y.: Reasoning with heterogeneous graph alignment for video question answering. In: Proceedings of the AAAI Conference on Artificial Intelligence. vol. 34, pp. 11109–11116 (2020)
2020
Cited alongside, same era.
JiayinCai, C., Shi, C., Li, L., Cheng, Y., Shan, Y.: Feature augmented memory with global attention network for videoqa. In: IJCAI. pp. 998–1004 (2020)
2020
Cited alongside, same era.
Le, T.M., Le, V., Venkatesh, S., Tran, T.: Hierarchical conditional relation networks for video question answering. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. 9972–9981 (2020)
2020
Cited alongside, same era.
2021
Later among the works it cites.
Yang, X., Zhang, H., Cai, J.: Deconfounded image captioning: A causal retrospect. IEEE Transactions on Pattern Analysis and Machine Intelligence (2021)
2021
Later among the works it cites.
Yang, X., Zhang, H., Qi, G., Cai, J.: Causal attention for vision-language tasks. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 9847–9857 (2021)
2021
Later among the works it cites.
Buch, S., Eyzaguirre, C., Gaidon, A., Wu, J., Fei-Fei, L., Niebles, J.C.: Revisiting the” video” in video-language understanding. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 2917–2927 (2022)
2022
Later among the works it cites.
Gao, L., Lei, Y., Zeng, P., Song, J., Wang, M., Shen, H.T.: Hierarchical representation network with auxiliary tasks for video captioning and video question answering. IEEE Transactions on Image Processing (2022)
2022
Later among the works it cites.
2022
Later among the works it cites.
Li, Y., Wang, X., Xiao, J., Ji, W., Chua, T.S.: Invariant grounding for video question answering. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 2928–2937 (2022)
2022
Later among the works it cites.
Liu, Y., Wang, K., Liu, L., Lan, H., Lin, L.: Tcgl: Temporal contrastive graph for self-supervised video representation learning. IEEE Transactions on Image Processing 31
2022
Later among the works it cites.
Liu, Y., Wei, Y.S., Yan, H., Li, G.B., Lin, L.: Causal reasoning meets visual representation learning: A prospective study. Machine Intelligence Research pp. 1–27 (2022)
2022
Later among the works it cites.
2022
Later among the works it cites.
Liu, Y., Zhang, X., Huang, F., Zhang, B., Li, Z.: Cross-attentional spatio-temporal semantic graph networks for video question answering. IEEE Transactions on Image Processing 31
2022
Later among the works it cites.
Ni, B., Peng, H., Chen, M., Zhang, S., Meng, G., Fu, J., Xiang, S., Ling, H.: Expanding language-image pretrained models for general video recognition. In: Computer Vision–ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part IV. pp. 1–18. Springer (2022)
2022
Later among the works it cites.
Ni, J., Sarbajna, R., Liu, Y., Ngu, A.H., Yan, Y.: Cross-modal knowledge distillation for vision-to-sensor action recognition. In: ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). pp. 4448–4452. IEEE (2022)
2022
Later among the works it cites.
Zhu, Y., Zhang, Y., Liu, L., Liu, Y., Li, G., Mao, M., Lin, L.: Hybrid-order representation learning for electricity theft detection. IEEE Transactions on Industrial Informatics (2022)
2022
Later among the works it cites.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
Liu, Y., Li, G., Lin, L.: Cross-modal causal relational reasoning for event-level visual question answering. IEEE Transactions on Pattern Analysis and Machine Intelligence (2023)
2023
Closest in time.
Liu, Y., Tan, Y., Lan, H.: Self-supervised contrastive learning for audio-visual action recognition. In: 2023 IEEE International Conference on Image Processing (ICIP). pp. 1000–1004. IEEE (2023)
2023
Closest in time.
2023
Closest in time.
Wang, K., Liu, L., Liu, Y., Li, G., Zhou, F., Lin, L.: Urban regional function guided traffic flow prediction. Information Sciences 634
2023
Closest in time.
Wei, Y., Liu, Y., Yan, H., Li, G., Lin, L.: Visual causal scene refinement for video question answering. In: Proceedings of the 31st ACM International Conference on Multimedia. p. 377–386. MM ’23, Association for Computing Machinery, New York, NY, USA (2023). https://doi.org/10.1145/3581783.3611873, https://doi.org/10.1145/3581783.3611873
2023
Closest in time.
2023
Closest in time.
Yan, H., Liu, Y., Wei, Y., Li, Z., Li, G., Lin, L.: Skeletonmae: graph-based masked autoencoder for skeleton sequence pre-training. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 5606–5618 (2023)
2023
Closest in time.