Fetching the paper…
Reading the bibliography…
This paper proposes a method to gain extra supervision via multi-task learning for multi-modal video question answering.
1905
Earlier work this paper cites.
R. Caruana, “Multitask learning: A knowledge-based source of inductive bias,” in Proceedings of the Tenth International Conference on Machine Learning (ICML) , 1993
1993
Earlier work this paper cites.
R. Collobert and J. Weston, “A unified architecture for natural language processing: Deep neural networks with multitask learning,” in Proceedings of the 25th International Conference on Machine Learning (ICML) , 2008
2008
Earlier work this paper cites.
Y. Bengio, J. Louradour, R. Collobert, and J. Weston, “Curriculum learning,” in Proceedings of the 26th Annual International Conference on Machine Learning (ICML) , 2009
2009
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2009
2009
Earlier work this paper cites.
J. Pennington, R. Socher, and C. D. Manning, “Glove: Global vectors for word representation.” in Conference on Empirical Methods in Natural Language Processing (EMNLP) , 2014
2014
Earlier work this paper cites.
J. Weston, S. Chopra, and A. Bordes, “Memory networks,” in International Conference on Learning Representations (ICLR) , 2015
2015
Earlier work this paper cites.
S. Sukhbaatar, J. Weston, R. Fergus et al. , “End-to-end memory networks,” in Advances in Neural Information Processing Systems (NIPS) , 2015
2015
Earlier work this paper cites.
M. Malinowski, M. Rohrbach, and M. Fritz, “Ask your neurons: A neural-based approach to answering questions about images,” in IEEE International Conference on Computer Vision (ICCV) , 2015
2015
Earlier work this paper cites.
S. Ren, K. He, R. Girshick, and J. Sun, “Faster r-cnn: Towards real-time object detection with region proposal networks,” in Advances in Neural Information Processing Systems (NIPS) , 2015
2015
Earlier work this paper cites.
E. Hoffer and N. Ailon, “Deep metric learning using triplet network,” in International Conference on Learning Representations Workshop Track (ICLR) , 2015
2015
Earlier work this paper cites.
Z. Yang, X. He, J. Gao, and A. Smola, “Stacked attention networks for image question answering,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2016
2016
Earlier work this paper cites.
M. Tapaswi, Y. Zhu, R. Stiefelhagen, A. Torralba, R. Urtasun, and S. Fidler, “Movieqa: Understanding stories in movies through question-answering,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2016
2016
Earlier work this paper cites.
L. Castrejón, Y. Aytar, C. Vondrick, H. Pirsiavash, and A. Torralba, “Learning aligned cross-modal representations from weakly aligned data,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2016
2016
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2016
2016
Cited alongside, same era.
H. Nam, J.-W. Ha, and J. Kim, “Dual attention networks for multimodal reasoning and matching,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2017
2017
Cited alongside, same era.
H. Ben-Younes, R. Cadène, N. Thome, and M. Cord, “Mutan: Multimodal tucker fusion for visual question answering,” in IEEE International Conference on Computer Vision (ICCV) , 2017
2017
Cited alongside, same era.
Y. Jang, Y. Song, Y. Yu, Y. Kim, and G. Kim, “Tgif-qa: Toward spatio-temporal reasoning in visual question answering,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2017
2017
Cited alongside, same era.
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma, M. S. Bernstein, and L. Fei-Fei, “Visual genome: Connecting language and vision using crowdsourced dense image annotations,” International Journal of Computer Vision , vol. 123, no. 1, pp. 32–73, May 2017
2017
Later among the works it cites.
A. F. H. H. Minjoon Seo, Aniruddha Kembhavi, “Bidirectional attention flow for machine comprehension,” in International Conference on Learning Representations (ICLR) , 2017
2017
Later among the works it cites.
S. Min, V. Zhong, R. Socher, and C. Xiong, “Efficient and robust question answering from minimal context over documents,” in Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (ACL) , 2018
2018
Later among the works it cites.
C. Xiong, V. Zhong, and R. Socher, “DCN+: Mixed objective and deep residual coattention for question answering,” in International Conference on Learning Representations (ICLR) , 2018
2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2017
Cited alongside, same era.
K. Kim, M. Heo, S. Choi, and B. Zhang, “Deepstory: Video story QA by deep embedded memory networks,” in International Joint Conference on Artificial Intelligence, (IJCAI) , 2017
2017
Cited alongside, same era.
A. Agrawal, J. Lu, S. Antol, M. Mitchell, C. L. Zitnick, D. Parikh, and D. Batra, “Vqa: Visual question answering,” Int. J. Comput. Vision , vol. 123, no. 1, pp. 4–31, May 2017
2017
Cited alongside, same era.
J. Gao, C. Sun, Z. Yang, and R. Nevatia, “Tall: Temporal activity localization via language query,” in IEEE International Conference on Computer Vision (ICCV) , 2017
2017
Cited alongside, same era.
J. Kim and C. D. Yoo, “Deep partial person re-identification via attention model,” in International Conference on Image Processing (ICIP) , 2017
2017
Cited alongside, same era.
T. Xiao, S. Li, B. Wang, L. Lin, and X. Wang, “Joint detection and identification feature learning for person search,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2017
2017
Cited alongside, same era.
A. Karpathy and L. Fei-Fei, “Deep visual-semantic alignments for generating image descriptions,” IEEE Trans. Pattern Anal. Mach. Intell. , vol. 39, no. 4, pp. 664–676, 2017
2017
Cited alongside, same era.
L. A. Hendricks, O. Wang, E. Shechtman, J. Sivic, T. Darrell, and B. C. Russell, “Localizing moments in video with natural language,” in IEEE International Conference on Computer Vision (ICCV) , 2017
2017
Cited alongside, same era.
Later among the works it cites.
P. Anderson, X. He, C. Buehler, D. Teney, M. Johnson, S. Gould, and L. Zhang, “Bottom-up and top-down attention for image captioning and visual question answering,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2018
2018
Later among the works it cites.
J. Gao, R. Ge, K. Chen, and R. Nevatia, “Motion-appearance co-memory networks for video question answering,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2018
2018
Later among the works it cites.
Y. Yu, J. Kim, and G. Kim, “A joint sequence fusion model for video question answering and retrieval,” in European Conference on Computer Vision (ECCV) , 2018
2018
Later among the works it cites.
J. Lei, L. Yu, M. Bansal, and T. L. Berg, “Tvqa: Localized, compositional video question answering,” in Conference on Empirical Methods in Natural Language Processing (EMNLP) , 2018
2018
Later among the works it cites.
J. Liang, L. Jiang, L. Cao, L.-J. Li, and A. Hauptmann, “Focal visual-text attention for visual question answering,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2018
2018
Later among the works it cites.
B. Wang, Y. Xu, Y. Han, and R. Hong, “Movie question answering: Remembering the textual cues for layered visual contents,” in AAAI Conference on Artificial Intelligence , 2018
2018
Later among the works it cites.
K.-M. Kim, S.-H. Choi, and B.-T. Zhang, “Multimodal dual attention memory for video story question answering,” in European Conference on Computer Vision (ECCV) , 2018
2018
Later among the works it cites.
Y. Li, N. Duan, B. Zhou, X. Chu, W. Ouyang, X. Wang, and M. Zhou, “Visual question generation as dual task of visual question answering,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2018
2018
Later among the works it cites.
Y. Yu, J. Kim, and G. Kim, “A joint sequence fusion model for video question answering and retrieval,” in European Conference on Computer Vision (ECCV) , 2018
2018
Later among the works it cites.
A. W. Yu, D. Dohan, Q. Le, T. Luong, R. Zhao, and K. Chen, “Fast and accurate reading comprehension by combining self-attention and convolution,” in International Conference on Learning Representations (ICLR) , 2018
2018
Later among the works it cites.