Fetching the paper…
Reading the bibliography…
Video description entails automatically generating coherent natural language sentences that narrate the content of a given video.
R. K. Srihari, “Automatic indexing and content-based retrieval of captioned images,” Computer
1995
Earlier work this paper cites.
A. Kojima, T. Tamura, and K. Fukunaga, “Natural language description of human activities from video images based on concept hierarchy of actions,” International Journal of Computer Vision (IJCV)
2002
Earlier work this paper cites.
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu, “Bleu: a method for automatic evaluation of machine translation,” in Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics (ACL)
2002
Earlier work this paper cites.
C.-Y. Lin, “Rouge: A package for automatic evaluation of summaries,” in Text Summarization Branches Out
2004
Earlier work this paper cites.
Y. Bengio, J. Louradour, R. Collobert, and J. Weston, “Curriculum learning,” in Proceedings of the International Conference on Machine Learning (ICML)
2009
Earlier work this paper cites.
2012
Earlier work this paper cites.
S. Ji, W. Xu, M. Yang, and K. Yu, “3d convolutional neural networks for human action recognition,” IEEE Transactions on Pattern Analysis and Machine Intelligence (PAMI)
2012
Earlier work this paper cites.
S. Guadarrama, N. Krishnamoorthy, G. Malkarnenkar, S. Venugopalan, R. Mooney, T. Darrell, and K. Saenko, “Youtube2text: Recognizing and describing arbitrary activities using semantic hierarchies and zero-shot recognition,” in Proceedings of the IEEE International Conference on Computer Vision (ICCV)
2013
Earlier work this paper cites.
P. Das, C. Xu, R. F. Doell, and J. J. Corso, “A thousand frames in just a few words: Lingual description of videos through latent topics and sparse object stitching,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
2013
Earlier work this paper cites.
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick, “Microsoft coco: Common objects in context,” in Proceedings of the European Conference on Computer Vision (ECCV)
2014
Earlier work this paper cites.
M. Denkowski and A. Lavie, “Meteor universal: Language specific translation evaluation for any target language,” in Proceedings of the Workshop on Statistical Machine Translation (STATMT)
2014
Earlier work this paper cites.
X. Chen and A. Gupta, “Webly supervised learning of convolutional networks,” in Proceedings of the IEEE International Conference on Computer Vision (ICCV)
2015
Earlier work this paper cites.
S. Venugopalan, H. Xu, J. Donahue, M. Rohrbach, R. Mooney, and K. Saenko, “Translating videos to natural language using deep recurrent neural networks,” in Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL)
2015
Earlier work this paper cites.
L. Yao, A. Torabi, K. Cho, N. Ballas, C. Pal, H. Larochelle, and A. Courville, “Describing videos by exploiting temporal structure,” in Proceedings of the IEEE International Conference on Computer Vision (ICCV)
2015
Earlier work this paper cites.
S. Venugopalan, M. Rohrbach, J. Donahue, R. Mooney, T. Darrell, and K. Saenko, “Sequence to sequence-video to text,” in Proceedings of the IEEE International Conference on Computer Vision (ICCV)
2015
Earlier work this paper cites.
S. Ren, K. He, R. Girshick, and J. Sun, “Faster r-cnn: Towards real-time object detection with region proposal networks,” in Proceedings of the Advances in Neural Information Processing Systems (NeurIPS)
2015
Earlier work this paper cites.
R. Vedantam, C. Lawrence Zitnick, and D. Parikh, “Cider: Consensus-based image description evaluation,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
2015
Earlier work this paper cites.
X. Wang and A. Gupta, “Unsupervised learning of visual representations using videos,” in Proceedings of the IEEE International Conference on Computer Vision (ICCV)
2015
Earlier work this paper cites.
H. Yu, J. Wang, Z. Huang, Y. Yang, and W. Xu, “Video paragraph captioning using hierarchical recurrent neural networks,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
2016
Earlier work this paper cites.
Y. Wei, X. Liang, Y. Chen, X. Shen, M.-M. Cheng, J. Feng, Y. Zhao, and S. Yan, “Stc: A simple to complex framework for weakly-supervised semantic segmentation,” IEEE Transactions on Pattern Analysis and Machine Intelligence (PAMI)
2016
Earlier work this paper cites.
M. Sachan and E. Xing, “Easy questions first? a case study on curriculum learning for question answering,” in Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (ACL)
2016
Earlier work this paper cites.
R. Krishna, K. Hata, F. Ren, L. Fei-Fei, and J. Carlos Niebles, “Dense-captioning events in videos,” in Proceedings of the IEEE International Conference on Computer Vision (ICCV)
2017
Cited alongside, same era.
P. Morerio, J. Cavazza, R. Volpi, R. Vidal, and V. Murino, “Curriculum dropout,” in Proceedings of the IEEE International Conference on Computer Vision (ICCV)
2017
Cited alongside, same era.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in Proceedings of the Advances in Neural Information Processing Systems (NeurIPS)
2017
Cited alongside, same era.
C. Hori, T. Hori, T.-Y. Lee, Z. Zhang, B. Harsham, J. R. Hershey, T. K. Marks, and K. Sumi, “Attention-based multimodal fusion for video description,” in Proceedings of the IEEE International Conference on Computer Vision (ICCV)
2017
Cited alongside, same era.
L. Zhou, Y. Kalantidis, X. Chen, J. J. Corso, and M. Rohrbach, “Grounded video description,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
2019
Later among the works it cites.
2019
Later among the works it cites.
B. Yuksel, P. Fazli, U. Mathur, V. Bisht, S. J. Kim, J. J. Lee, S. J. Jin, Y.-T. Siu, J. A. Miele, and I. Yoon, “Increasing video accessibility for visually impaired users with human-in-the-loop machine learning,” in Proceedings of the ACM SIGCHI Conference Extended Abstracts on Human Factors in Computing Systems (CHI)
2020
Later among the works it cites.
B. Yuksel, P. Fazli, U. Mathur, V. Bisht, S. J. Kim, J. J. Lee, S. J. Jin, Y.-T. Siu, J. A. Miele, and I. Yoon, “Human-in-the-loop machine learning to increase video accessibility for visually impaired and blind users,” in Proceedings of the ACM Conference on Designing Interactive Systems (DIS)
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2017
Cited alongside, same era.
2017
Cited alongside, same era.
R. Shetty, M. Rohrbach, L. Anne Hendricks, M. Fritz, and B. Schiele, “Speaking the same language: Matching machine to human captions by adversarial training,” in Proceedings of the IEEE International Conference on Computer Vision (ICCV)
2017
Cited alongside, same era.
C. Doersch and A. Zisserman, “Multi-task self-supervised visual learning,” in Proceedings of the IEEE International Conference on Computer Vision (ICCV)
2017
Cited alongside, same era.
S. Gella, M. Lewis, and M. Rohrbach, “A dataset for telling the stories of social media videos,” in Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP)
2018
Cited alongside, same era.
L. Zhou, C. Xu, and J. J. Corso, “Towards automatic learning of procedures from web instructional videos,” in Proceedings of the AAAI Conference on Artificial Intelligence (AAAI)
2018
Cited alongside, same era.
L. Zhou, Y. Zhou, J. J. Corso, R. Socher, and C. Xiong, “End-to-end dense video captioning with masked transformer,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
2018
Cited alongside, same era.
N. Xu, A.-A. Liu, Y. Wong, Y. Zhang, W. Nie, Y. Su, and M. Kankanhalli, “Dual-stream recurrent neural network for video captioning,” IEEE Transactions on Circuits and Systems for Video Technology (TCSVT)
2018
Cited alongside, same era.
2020
Later among the works it cites.
S. Narvekar, B. Peng, M. Leonetti, J. Sinapov, M. E. Taylor, and P. Stone, “Curriculum learning for reinforcement learning domains: A framework and survey,” The Journal of Machine Learning Research (JMLR)
2020
Later among the works it cites.
2020
Later among the works it cites.
V. Iashin and E. Rahtu, “Multi-modal dense video captioning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshop (CVPR Workshop)
2020
Later among the works it cites.
J. Lei, L. Yu, T. L. Berg, and M. Bansal, “Tvr: A large-scale dataset for video-subtitle moment retrieval,” in Proceedings of the European Conference on Computer Vision (ECCV)
2020
Later among the works it cites.
P. Soviany, C. Ardei, R. T. Ionescu, and M. Leordeanu, “Image difficulty curriculum for generative adversarial networks (cugan),” in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV)
2020
Later among the works it cites.
S. Ging, M. Zolfaghari, H. Pirsiavash, and T. Brox, “Coot: Cooperative hierarchical transformer for video-text representation learning,” in Proceedings of the Advances in Neural Information Processing Systems (NeurIPS)
2020
Later among the works it cites.
A. Bodi, P. Fazli, S. Ihorn, Y.-T. Siu, A. T. Scott, L. Narins, Y. Kant, A. Das, and I. Yoon, “Automated video description for blind and low vision users,” in Proceedings of the ACM SIGCHI Conference Extended Abstracts on Human Factors in Computing Systems (CHI)
2021
Later among the works it cites.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever, “Learning transferable visual models from natural language supervision,” in Proceedings of the International Conference on Machine Learning (ICML)
2021
Later among the works it cites.
C. S. Kanani, S. Saha, and P. Bhattacharyya, “Global object proposals for improving multi-sentence video descriptions,” in Proceedings of the International Joint Conference on Neural Networks (IJCNN)
2021
Later among the works it cites.
T. Wang, R. Zhang, Z. Lu, F. Zheng, R. Cheng, and P. Luo, “End-to-end dense video captioning with parallel decoding,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)
2021
Later among the works it cites.
Y. Song, S. Chen, and Q. Jin, “Towards diverse paragraph captioning for untrimmed videos,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
2021
Later among the works it cites.
J. Gou, B. Yu, S. J. Maybank, and D. Tao, “Knowledge distillation: A survey,” International Journal of Computer Vision
2021
Later among the works it cites.
N. Aafaq, A. S. Mian, N. Akhtar, W. Liu, and M. Shah, “Dense video captioning with early linguistic information fusion,” IEEE Transactions on Multimedia
2022
Later among the works it cites.
S. Li, B. Yang, and Y. Zou, “Adaptive curriculum learning for video captioning,” IEEE Access
2022
Later among the works it cites.
K. Yamazaki, K. Vo, Q. S. Truong, B. Raj, and N. Le, “Vltint: Visual-linguistic transformer-in-transformer for coherent video paragraph captioning,” in Proceedings of the AAAI Conference on Artificial Intelligence (AAAI)
2023
Closest in time.
S. Dong, T. Niu, X. Luo, W. Liu, and X. Xu, “Semantic embedding guided attention with explicit visual feature fusion for video captioning,” ACM Transactions on Multimedia Computing, Communications and Applications (TOMM)
2023
Closest in time.