Fetching the paper…
Reading the bibliography…
Video Question Answering (Video QA) requires fine-grained understanding of both video and language modalities to answer the given questions.
Progressive Attention Memory Network for Movie Story Question Answering
Kim, J.; Ma, M.; Kim, K.; Kim, S.; and Yoo, C. D. 2019b · 1904
Earlier work this paper cites.
Tvqa+: Spatio-temporal grounding for video question answering
Lei, J.; Yu, L.; Berg, T. L.; and Bansal, M. 2019 · 1904
Earlier work this paper cites.
Learning to Reason with Relational Video Representation for Question Answering
Le, T. M.; Le, V.; Venkatesh, S.; and Tran, T. 2019 · 1907
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Liu, Y.; Ott, M.; Goyal, N.; Du, J.; Joshi, M.; Chen, D.; Levy, O.; Lewis, M.; Zettlemoyer, L.; and Stoyanov, V. 2019 · 1907
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C.; Shazeer, N.; Roberts, A.; Lee, K.; Narang, S.; Matena, M.; Zhou, Y.; Li, W.; and Liu, P. J. 2019 · 1910
Earlier work this paper cites.
What makes for good views for contrastive learning
Tian, Y.; Sun, C.; Poole, B.; Krishnan, D.; Schmid, C.; and Isola, P. 2020 · 2005
Earlier work this paper cites.
Dimensionality reduction by learning an invariant mapping
Hadsell, R.; Chopra, S.; and LeCun, Y. 2006 · 2006
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J.; Dong, W.; Socher, R.; Li, L.-J.; Li, K.; and Fei-Fei, L. 2009 · 2009
Earlier work this paper cites.
Discriminative unsupervised feature learning with convolutional neural networks
Dosovitskiy, A.; Springenberg, J. T.; Riedmiller, M.; and Brox, T. 2014 · 2014
Earlier work this paper cites.
Vqa: Visual question answering
Antol, S.; Agrawal, A.; Lu, J.; Mitchell, M.; Batra, D.; Lawrence Zitnick, C.; and Parikh, D. 2015 · 2015
Earlier work this paper cites.
Microsoft coco captions: Data collection and evaluation server
Chen, X.; Fang, H.; Lin, T.-Y.; Vedantam, R.; Gupta, S.; Dollár, P.; and Zitnick, C. L. 2015 · 2015
Earlier work this paper cites.
Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models
Plummer, B. A.; Wang, L.; Cervantes, C. M.; Caicedo, J. C.; Hockenmaier, J.; and Lazebnik, S. 2015 · 2015
Cited alongside, same era.
Faster r-cnn: Towards real-time object detection with region proposal networks
Ren, S.; He, K.; Girshick, R.; and Sun, J. 2015 · 2015
Cited alongside, same era.
Multimodal compact bilinear pooling for visual question answering and visual grounding
Fukui, A.; Park, D. H.; Yang, D.; Rohrbach, A.; Darrell, T.; and Rohrbach, M. 2016 · 2016
Cited alongside, same era.
Deep residual learning for image recognition
He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016 · 2016
Cited alongside, same era.
Ask, attend and answer: Exploring question-guided spatial attention for visual question answering
Xu, H.; and Saenko, K. 2016 · 2016
Cited alongside, same era.
Tvqa: Localized, compositional video question answering
Lei, J.; Yu, L.; Bansal, M.; and Berg, T. L. 2018 · 2018
Later among the works it cites.
Attentive moment retrieval in videos
Liu, M.; Wang, X.; Nie, L.; He, X.; Chen, B.; and Chua, T.-S. 2018 · 2018
Later among the works it cites.
End-to-end dense video captioning with masked transformer
Zhou, L.; Zhou, Y.; Corso, J. J.; Socher, R.; and Xiong, C. 2018 · 2018
Later among the works it cites.
Gaining extra supervision via multi-task learning for multi-modal video question answering
Kim, J.; Ma, M.; Kim, K.; Kim, S.; and Yoo, C. D. 2019a · 2019
Later among the works it cites.
A Simple Framework for Contrastive Learning of Visual Representations
Chen, T.; Kornblith, S.; Norouzi, M.; and Hinton, G. 2020 · 2020
Closest in time.
Character Matters: Video Story Understanding with Character-Aware Relations
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Making the V in VQA matter: Elevating the role of image understanding in Visual Question Answering
Goyal, Y.; Khot, T.; Summers-Stay, D.; Batra, D.; and Parikh, D. 2017 · 2017
Cited alongside, same era.
Tgif-qa: Toward spatio-temporal reasoning in visual question answering
Jang, Y.; Song, Y.; Yu, Y.; Kim, Y.; and Kim, G. 2017 · 2017
Cited alongside, same era.
Deepstory: Video story qa by deep embedded memory networks
Kim, K.-M.; Heo, M.-O.; Choi, S.-H.; and Zhang, B.-T. 2017 · 2017
Cited alongside, same era.
Dense-captioning events in videos
Krishna, R.; Hata, K.; Ren, F.; Fei-Fei, L.; and Carlos Niebles, J. 2017 · 2017
Cited alongside, same era.
A read-write memory network for movie story understanding
Na, S.; Lee, S.; Kim, J.; and Kim, G. 2017 · 2017
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2018 · 2018
Cited alongside, same era.
Dense Captioning with Joint Inference and Visual Context
Yang, L.; Tang, K.; Yang, J.; and Li, L.-J. 2016a
Cited in the paper.
Geng, S.; Zhang, J.; Fu, Z.; Gao, P.; Zhang, H.; and de Melo, G. 2020 · 2020
Closest in time.
Momentum contrast for unsupervised visual representation learning
He, K.; Fan, H.; Wu, Y.; Xie, S.; and Girshick, R. 2020 · 2020
Closest in time.
Supervised Contrastive Learning
Khosla, P.; Teterwak, P.; Wang, C.; Sarna, A.; Tian, Y.; Isola, P.; Maschinot, A.; Liu, C.; and Krishnan, D. 2020 · 2020
Closest in time.
Dense-Caption Matching and Frame-Selection Gating for Temporal Localization in VideoQA
Kim, H.; Tang, Z.; and Bansal, M. 2020 · 2020
Closest in time.
Modality Shifting Attention Network for Multi-Modal Video Question Answering
Kim, J.; Ma, M.; Pham, T.; Kim, K.; and Yoo, C. D. 2020 · 2020
Closest in time.
BERT representations for Video Question Answering
Yang, Z.; Garcia, N.; Chu, C.; Otani, M.; Nakashima, Y.; and Takemura, H. 2020 · 2020
Closest in time.