Fetching the paper…
Reading the bibliography…
Video question answering (VideoQA) is an essential task in vision-language understanding, which has attracted numerous research attention recently.
A. P. Dempster, N. M. Laird, and D. B. Rubin, “Maximum likelihood from incomplete data via the em algorithm,” Journal of the Royal Statistical Society: Series B (Methodological) , vol. 39, no. 1, pp. 1–22, 1977
1977
Earlier work this paper cites.
C. Fan, X. Zhang, S. Zhang, W. Wang, C. Zhang, and H. Huang, “Heterogeneous memory enhanced multimodal attention model for video question answering,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2019, pp. 1999–2007
2007
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in 2009 IEEE conference on computer vision and pattern recognition . Ieee, 2009, pp. 248–255
2009
Earlier work this paper cites.
V. Nair and G. E. Hinton, “Rectified linear units improve restricted boltzmann machines,” in Proceedings of the 27th International Conference on International Conference on Machine Learning , 2010, pp. 807–814
2010
Earlier work this paper cites.
M. Malinowski and M. Fritz, “A multi-world approach to question answering about real-world scenes based on uncertain input,” Advances in neural information processing systems , vol. 27, 2014
2014
Earlier work this paper cites.
J. Pennington, R. Socher, and C. D. Manning, “Glove: Global vectors for word representation,” in Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP) , 2014, pp. 1532–1543
2014
Earlier work this paper cites.
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov, “Dropout: a simple way to prevent neural networks from overfitting,” The journal of machine learning research , vol. 15, no. 1, pp. 1929–1958, 2014
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh, “Vqa: Visual question answering,” in Proceedings of the IEEE international conference on computer vision , 2015, pp. 2425–2433
2015
Earlier work this paper cites.
F. Caba Heilbron, V. Escorcia, B. Ghanem, and J. Carlos Niebles, “Activitynet: A large-scale video benchmark for human activity understanding,” in Proceedings of the ieee conference on computer vision and pattern recognition , 2015, pp. 961–970
2015
Earlier work this paper cites.
S. Sukhbaatar, J. Weston, R. Fergus et al. , “End-to-end memory networks,” Advances in neural information processing systems , vol. 28, 2015
2015
Earlier work this paper cites.
L. Yao, A. Torabi, K. Cho, N. Ballas, C. Pal, H. Larochelle, and A. Courville, “Describing videos by exploiting temporal structure,” in Proceedings of the IEEE international conference on computer vision , 2015, pp. 4507–4515
2015
Earlier work this paper cites.
D. Tran, L. Bourdev, R. Fergus, L. Torresani, and M. Paluri, “Learning spatiotemporal features with 3d convolutional networks,” in Proceedings of the IEEE international conference on computer vision , 2015, pp. 4489–4497
2015
Earlier work this paper cites.
S. Ren, K. He, R. Girshick, and J. Sun, “Faster r-cnn: Towards real-time object detection with region proposal networks,” Advances in neural information processing systems , vol. 28, 2015
2015
Earlier work this paper cites.
B. Thomee, D. A. Shamma, G. Friedland, B. Elizalde, K. Ni, D. Poland, D. Borth, and L.-J. Li, “Yfcc100m: The new data in multimedia research,” Communications of the ACM , vol. 59, no. 2, pp. 64–73, 2016
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 770–778
2016
Earlier work this paper cites.
Y. Jang, Y. Song, Y. Yu, Y. Kim, and G. Kim, “Tgif-qa: Toward spatio-temporal reasoning in visual question answering,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2017, pp. 2758–2766
2017
Earlier work this paper cites.
L. Zhu, Z. Xu, Y. Yang, and A. G. Hauptmann, “Uncovering the temporal context for video question answering,” International Journal of Computer Vision , vol. 124, no. 3, pp. 409–421, 2017
2017
Cited alongside, same era.
L. Anne Hendricks, O. Wang, E. Shechtman, J. Sivic, T. Darrell, and B. Russell, “Localizing moments in video with natural language,” in Proceedings of the IEEE international conference on computer vision , 2017, pp. 5803–5812
2017
Cited alongside, same era.
J. Gao, C. Sun, Z. Yang, and R. Nevatia, “Tall: Temporal activity localization via language query,” in Proceedings of the IEEE international conference on computer vision , 2017, pp. 5267–5275
2017
Cited alongside, same era.
D. Xu, Z. Zhao, J. Xiao, F. Wu, H. Zhang, X. He, and Y. Zhuang, “Video question answering via gradually refined attention over appearance and motion,” in Proceedings of the 25th ACM international conference on Multimedia , 2017, pp. 1645–1653
2017
Cited alongside, same era.
A. Yang, A. Miech, J. Sivic, I. Laptev, and C. Schmid, “Just ask: Learning to answer questions from millions of narrated videos,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 1686–1697
2019
Later among the works it cites.
T. M. Le, V. Le, S. Venkatesh, and T. Tran, “Hierarchical conditional relation networks for video question answering,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2020, pp. 9972–9981
2020
Later among the works it cites.
J. Jiang, Z. Chen, H. Lin, X. Zhao, and Y. Gao, “Divide and conquer: Question-guided spatio-temporal contextual attention for video question answering,” in Proceedings of the AAAI Conference on Artificial Intelligence , 2020, pp. 11 101–11 108
2020
Later among the works it cites.
M. Ma, S. Yoon, J. Kim, Y. Lee, S. Kang, and C. D. Yoo, “Vlanet: Video-language alignment network for weakly-supervised video moment retrieval,” in European Conference on Computer Vision . Springer, 2020, pp. 156–171
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems , vol. 30, 2017
2017
Cited alongside, same era.
S. Xie, R. Girshick, P. Dollár, Z. Tu, and K. He, “Aggregated residual transformations for deep neural networks,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2017, pp. 1492–1500
2017
Cited alongside, same era.
2017
Cited alongside, same era.
J. Gao, R. Ge, K. Chen, and R. Nevatia, “Motion-appearance co-memory networks for video question answering,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2018, pp. 6576–6585
2018
Cited alongside, same era.
Z. Yu, J. Yu, C. Xiang, J. Fan, and D. Tao, “Beyond bilinear: Generalized multimodal factorized high-order pooling for visual question answering,” IEEE transactions on neural networks and learning systems , vol. 29, no. 12, pp. 5947–5959, 2018
2018
Cited alongside, same era.
2018
Cited alongside, same era.
J. Lei, L. Yu, M. Bansal, and T. L. Berg, “Tvqa: Localized, compositional video question answering,” in EMNLP , 2018
2018
Cited alongside, same era.
S. Xie, C. Sun, J. Huang, Z. Tu, and K. Murphy, “Rethinking spatiotemporal feature learning: Speed-accuracy trade-offs in video classification,” in Proceedings of the European conference on computer vision (ECCV) , 2018, pp. 305–321
2018
Cited alongside, same era.
2020
Later among the works it cites.
P. Jiang and Y. Han, “Reasoning with heterogeneous graph alignment for video question answering,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 34, no. 07, 2020, pp. 11 109–11 116
2020
Later among the works it cites.
Y. Zhuang, D. Xu, X. Yan, W. Cheng, Z. Zhao, S. Pu, and J. Xiao, “Multichannel attention refinement for video question answering,” ACM Transactions on Multimedia Computing, Communications, and Applications (TOMM) , vol. 16, no. 1s, pp. 1–23, 2020
2020
Later among the works it cites.
T. Yu, J. Yu, Z. Yu, Q. Huang, and Q. Tian, “Long-term video question answering via multimodal hierarchical memory attentive networks,” IEEE Transactions on Circuits and Systems for Video Technology , vol. 31, no. 3, pp. 931–944, 2020
2020
Later among the works it cites.
D. Patel, R. Parikh, and Y. Shastri, “Recent advances in video question answering: A review of datasets and methods,” in International Conference on Pattern Recognition . Springer, 2021, pp. 339–356
2021
Later among the works it cites.
K. Khurana and U. Deshpande, “Video question-answering techniques, benchmark datasets and evaluation metrics leveraging video captioning: A comprehensive survey.” IEEE Access , 2021
2021
Later among the works it cites.
J. Wang, B.-K. Bao, and C. Xu, “Dualvgr: A dual-visual graph reasoning unit for video question answering,” IEEE Transactions on Multimedia , vol. 24, pp. 3369–3380, 2021
2021
Later among the works it cites.
J. Xiao, X. Shang, A. Yao, and T.-S. Chua, “Next-qa: Next phase of question-answering to explaining temporal actions,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2021, pp. 9777–9786
2021
Later among the works it cites.
Y. Wang, J. Deng, W. Zhou, and H. Li, “Weakly supervised temporal adjacent network for language grounding,” IEEE Transactions on Multimedia , 2021
2021
Later among the works it cites.
M. Grunde-McLaughlin, R. Krishna, and M. Agrawala, “Agqa: A benchmark for compositional spatio-temporal reasoning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2021, pp. 11 287–11 297
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
J. Huang, Y. Liu, S. Gong, and H. Jin, “Cross-sentence temporal and semantic relations in video activity localisation,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 7199–7208
2021
Later among the works it cites.
P. H. Seo, A. Nagrani, and C. Schmid, “Look before you speak: Visually contextualized utterances,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2021, pp. 16 877–16 887
2021
Later among the works it cites.
T. Qian, J. Chen, S. Chen, B. Wu, and Y.-G. Jiang, “Scene graph refinement network for visual question answering,” IEEE Transactions on Multimedia , 2022
2022
Closest in time.
Z. Xu, K. Wei, X. Yang, and C. Deng, “Point-supervised video temporal grounding,” IEEE Transactions on Multimedia , 2022
2022
Closest in time.