Fetching the paper…
Reading the bibliography…
Video description involves the generation of the natural language description of actions, events, and objects in the video.
The case for case
Charles Fillmore. 1967 · 1967
Earlier work this paper cites.
Detection and tracking of point features
Carlo Tomasi and Takeo Kanade. 1991 · 1991
Earlier work this paper cites.
Association of motion verbs with vehicle movements extracted from dense optical flow fields. In European Conference on Computer Vision . Springer, 338–347
Henner Kollnig, H-H Nagel, and Michael Otte. 1994 · 1994
Earlier work this paper cites.
Good features to track. In 1994 Proceedings of IEEE conference on computer vision and pattern recognition . IEEE, 593–600
Jianbo Shi et al · 1994
Earlier work this paper cites.
The" inverse hollywood problem": From video to scripts and storyboards via causal analysis. In AAAI/IAAI . Citeseer, 132–137
Matthew Brand. 1997 · 1997
Earlier work this paper cites.
Generating natural language description of human behavior from video images. In Proceedings 15th International Conference on Pattern Recognition. ICPR-2000 , Vol. 4. IEEE, 728–731
Atsuhiro Kojima, Masao Izumi, Takeshi Tamura, and Kunio Fukunaga. 2000 · 2000
Earlier work this paper cites.
Monitoring human behavior from video taken in an office environment
Douglas Ayers and Mubarak Shah. 2001 · 2001
Earlier work this paper cites.
Natural language description of human activities from video images based on concept hierarchy of actions
Atsuhiro Kojima, Takeshi Tamura, and Kunio Fukunaga. 2002 · 2002
Earlier work this paper cites.
BLEU: a method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting on association for computational linguistics . Association for Computational Linguistics, 311–318
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
CASEˆ E: a hierarchical event representation for the analysis of videos. In AAAI . 263–268
Asaad Hakeem, Yaser Sheikh, and Mubarak Shah. 2004 · 2004
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries. In Text summarization branches out . 74–81
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Understanding inverse document frequency: on theoretical arguments for IDF
Stephen Robertson. 2004 · 2004
Earlier work this paper cites.
METEOR: An automatic metric for MT evaluation with improved correlation with human judgments. In Proceedings of the acl workshop on intrinsic and extrinsic evaluation measures for machine translation and/or summarization . 65–72
Satanjeev Banerjee and Alon Lavie. 2005 · 2005
Earlier work this paper cites.
Moses: Open source toolkit for statistical machine translation. In Proceedings of the 45th annual meeting of the ACL on interactive poster and demonstration sessions . Association for Computational Linguistics, 177–180
Philipp Koehn, Hieu Hoang, Alexandra Birch, Chris Callison-Burch, Marcello Federico, Nicola Bertoldi, Brooke Cowan, Wade Shen, Christine Moran, Richard Zens, et al · 2007
Earlier work this paper cites.
Natural language descriptions of human behavior from video sequences. In Annual Conference on Artificial Intelligence . Springer, 279–292
Carles Fernández Tena, Pau Baiget, Xavier Roca, and Jordi Gonzàlez. 2007 · 2007
Earlier work this paper cites.
Save: A framework for semantic annotation of visual events. In 2008 IEEE Computer Society Conference on Computer Vision and Pattern Recognition Workshops . IEEE, 1–8
Mun Wai Lee, Asaad Hakeem, Niels Haering, and Song-Chun Zhu. 2008 · 2008
Earlier work this paper cites.
The Corpus of Contemporary American English as the first reliable monitor corpus of English
Mark Davies. 2010 · 2010
Earlier work this paper cites.
Cascade object detection with deformable part models. In 2010 IEEE Computer society conference on computer vision and pattern recognition . IEEE, 2241–2248
Pedro F Felzenszwalb, Ross B Girshick, and David McAllester. 2010 · 2010
Earlier work this paper cites.
Collecting highly parallel data for paraphrase evaluation. In Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies-Volume 1 . Association for Computational Linguistics, 190–200
David L Chen and William B Dolan. 2011 · 2011
Earlier work this paper cites.
Human focused video description. In 2011 IEEE International Conference on Computer Vision Workshops (ICCV Workshops) . IEEE, 1480–1487
Muhammad Usman Ghani Khan, Lei Zhang, and Yoshihiko Gotoh. 2011 · 2011
Earlier work this paper cites.
Andrei Barbu, Alexander Bridge, Zachary Burchill, Dan Coroian, Sven Dickinson, Sanja Fidler, Aaron Michaux, Sam Mussman, Siddharth Narayanaswamy, Dhaval Salvi, et al · 2012
Earlier work this paper cites.
Automated textual descriptions for a wide range of video events with 48 human actions. In European Conference on Computer Vision . Springer, 372–380
Patrick Hanckmann, Klamer Schutte, and Gertjan J Burghouts. 2012 · 2012
Earlier work this paper cites.
A database for fine grained activity detection of cooking activities. In 2012 IEEE conference on computer vision and pattern recognition . IEEE, 1194–1201
Marcus Rohrbach, Sikandar Amin, Mykhaylo Andriluka, and Bernt Schiele. 2012a · 2012
Earlier work this paper cites.
A thousand frames in just a few words: Lingual description of videos through latent topics and sparse object stitching. In Proceedings of the IEEE conference on computer vision and pattern recognition . 2634–2641
Pradipto Das, Chenliang Xu, Richard F Doell, and Jason J Corso. 2013 · 2013
Earlier work this paper cites.
Grounding Action Descriptions in Videos
Michaela Regneri, Marcus Rohrbach, Dominikus Wetzel, Stefan Thater, Bernt Schiele, and Manfred Pinkal. 2013 · 2013
Earlier work this paper cites.
Translating video content to natural language descriptions. In Proceedings of the IEEE International Conference on Computer Vision . 433–440
Marcus Rohrbach, Wei Qiu, Ivan Titov, Stefan Thater, Manfred Pinkal, and Bernt Schiele. 2013 · 2013
Earlier work this paper cites.
Crowdsourcing transcription beyond mechanical turk. In First AAAI conference on human computation and crowdsourcing . Citeseer
Haofeng Zhou, Denys Baskov, and Matthew Lease. 2013 · 2013
Earlier work this paper cites.
Large-scale video classification with convolutional neural networks. In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition . 1725–1732
Andrej Karpathy, George Toderici, Sanketh Shetty, Thomas Leung, Rahul Sukthankar, and Li Fei-Fei. 2014 · 2014
Earlier work this paper cites.
Coherent multi-sentence video description with variable level of detail. In German conference on pattern recognition . Springer, 184–195
Anna Rohrbach, Marcus Rohrbach, Wei Qiu, Annemarie Friedrich, Manfred Pinkal, and Bernt Schiele. 2014 · 2014
Earlier work this paper cites.
Translating videos to natural language using deep recurrent neural networks
Subhashini Venugopalan, Huijuan Xu, Jeff Donahue, Marcus Rohrbach, Raymond Mooney, and Kate Saenko. 2014 · 2014
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al · 2015
Earlier work this paper cites.
Using descriptive video services to create a large data source for video annotation research
Atousa Torabi, Christopher Pal, Hugo Larochelle, and Aaron Courville. 2015 · 2015
Earlier work this paper cites.
Cider: Consensus-based image description evaluation. In Proceedings of the IEEE conference on computer vision and pattern recognition . 4566–4575
Ramakrishna Vedantam, C Lawrence Zitnick, and Devi Parikh. 2015 · 2015
Cited alongside, same era.
Sequence to sequence-video to text. In Proceedings of the IEEE international conference on computer vision . 4534–4542
Subhashini Venugopalan, Marcus Rohrbach, Jeffrey Donahue, Raymond Mooney, Trevor Darrell, and Kate Saenko. 2015 · 2015
Cited alongside, same era.
Spice: Semantic propositional image caption evaluation. In European Conference on Computer Vision . Springer, 382–398
Peter Anderson, Basura Fernando, Mark Johnson, and Stephen Gould. 2016 · 2016
Cited alongside, same era.
Early embedding and late reranking for video captioning. In Proceedings of the 24th ACM international conference on Multimedia . 1082–1086
Jianfeng Dong, Xirong Li, Weiyu Lan, Yujia Huo, and Cees GM Snoek. 2016 · 2016
Cited alongside, same era.
Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition . 770–778
Multimodal architecture for video captioning with memory networks and an attention mechanism
Wei Li, Dashan Guo, and Xiangzhong Fang. 2018a · 2018
Later among the works it cites.
Yolov3: An incremental improvement
Joseph Redmon and Ali Farhadi. 2018 · 2018
Later among the works it cites.
Interpretable video captioning via trajectory structured localization. In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition . 6829–6837
Xian Wu, Guanbin Li, Qingxing Cao, Qingge Ji, and Liang Lin. 2018 · 2018
Later among the works it cites.
Move forward and tell: A progressive generator of video descriptions. In Proceedings of the European Conference on Computer Vision (ECCV) . 468–483
Yilei Xiong, Bo Dai, and Dahua Lin. 2018 · 2018
Later among the works it cites.
End-to-end dense video captioning with masked transformer. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . 8739–8748
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Cited alongside, same era.
Beyond caption to narrative: Video captioning with multiple sentences. In 2016 IEEE International Conference on Image Processing (ICIP) . IEEE, 3364–3368
Andrew Shin, Katsunori Ohnishi, and Tatsuya Harada. 2016 · 2016
Cited alongside, same era.
Hollywood in homes: Crowdsourcing data collection for activity understanding. In European Conference on Computer Vision . Springer, 510–526
Gunnar A Sigurdsson, Gül Varol, Xiaolong Wang, Ali Farhadi, Ivan Laptev, and Abhinav Gupta. 2016 · 2016
Cited alongside, same era.
Improving lstm-based video description with linguistic knowledge mined from text
Subhashini Venugopalan, Lisa Anne Hendricks, Raymond Mooney, and Kate Saenko. 2016 · 2016
Cited alongside, same era.
Msr-vtt: A large video description dataset for bridging video and language. In Proceedings of the IEEE conference on computer vision and pattern recognition . 5288–5296
Jun Xu, Tao Mei, Ting Yao, and Yong Rui. 2016 · 2016
Cited alongside, same era.
Video paragraph captioning using hierarchical recurrent neural networks. In Proceedings of the IEEE conference on computer vision and pattern recognition . 4584–4593
Haonan Yu, Jiang Wang, Zhiheng Huang, Yi Yang, and Wei Xu. 2016 · 2016
Cited alongside, same era.
Generation for user generated videos. In European conference on computer vision . Springer, 609–625
Kuo-Hao Zeng, Tseng-Hung Chen, Juan Carlos Niebles, and Min Sun. 2016 · 2016
Cited alongside, same era.
Hierarchical boundary-aware neural encoder for video captioning. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . 1657–1666
Lorenzo Baraldi, Costantino Grana, and Rita Cucchiara. 2017 · 2017
Cited alongside, same era.
Luowei Zhou, Yingbo Zhou, Jason J Corso, Richard Socher, and Caiming Xiong. 2018 · 2018
Later among the works it cites.
Video description: A survey of methods, datasets, and evaluation metrics
Nayyer Aafaq, Ajmal Mian, Wei Liu, Syed Zulqarnain Gilani, and Mubarak Shah. 2019b · 2019
Later among the works it cites.
Manjot Bilkhu, Siyang Wang, and Tushar Dobhal. 2019 · 2019
Later among the works it cites.
Boundary Detector Encoder and Decoder with Soft Attention for Video Captioning. In Asia-Pacific Web (APWeb) and Web-Age Information Management (WAIM) Joint International Conference on Web and Big Data . Springer, 105–115
Tangming Chen, Qike Zhao, and Jingkuan Song. 2019 · 2019
Later among the works it cites.
Fused GRU with semantic-temporal attention for video captioning
Lianli Gao, Xuanhan Wang, Jingkuan Song, and Yang Liu. 2019 · 2019
Later among the works it cites.
A comprehensive survey of deep learning for image captioning
MD Zakir Hossain, Ferdous Sohel, Mohd Fairuz Shiratuddin, and Hamid Laga. 2019 · 2019
Later among the works it cites.
End-to-end video captioning with multitask reinforcement learning. In 2019 IEEE Winter Conference on Applications of Computer Vision (WACV) . IEEE, 339–348
Lijun Li and Boqing Gong. 2019 · 2019
Later among the works it cites.
SibNet: Sibling Convolutional Encoder for Video Captioning
S. Liu, Z. Ren, and J. Yuan. 2020 · 2019
Later among the works it cites.
End-to-End Video Captioning. In 2019 IEEE/CVF International Conference on Computer Vision Workshop (ICCVW) . 1474–1482
S. Olivastri, G. Singh, and F. Cuzzolin. 2019 · 2019
Later among the works it cites.
A novel automatic shot boundary detection algorithm: robust to illumination and motion effect
Alok Singh, Dalton Meitei Thounaojam, and Saptarshi Chakraborty. 2019 · 2019
Later among the works it cites.
VATEX: A large-scale, high-quality multilingual dataset for video-and-language research. In Proceedings of the IEEE International Conference on Computer Vision . 4581–4591
Xin Wang, Jiawei Wu, Junkun Chen, Lei Li, Yuan-Fang Wang, and William Yang Wang. 2019 · 2019
Later among the works it cites.
Joint event detection and description in continuous video streams. In 2019 IEEE Winter Conference on Applications of Computer Vision (WACV) . IEEE, 396–405
Huijuan Xu, Boyang Li, Vasili Ramanishka, Leonid Sigal, and Kate Saenko. 2019a · 2019
Later among the works it cites.
Semantic-filtered Soft-Split-Aware video captioning with audio-augmented feature
Yuecong Xu, Jianfei Yang, and Kezhi Mao. 2019b · 2019
Later among the works it cites.
Grounded video description. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . 6578–6587
Luowei Zhou, Yannis Kalantidis, Xinlei Chen, Jason J Corso, and Marcus Rohrbach. 2019 · 2019
Later among the works it cites.
Multi-View Features and Hybrid Reward Strategies for Vatex Video Captioning Challenge 2019
Xinxin Zhu, Longteng Guo, Peng Yao, Jing Liu, and Hanqing Lu. 2019 · 2019
Later among the works it cites.
Multi-modal Feature Fusion with Feature Attention for VATEX Captioning Challenge 2020
Ke Lin, Zhuoxin Gan, and Liwei Wang. 2020 · 2020
Closest in time.
Multi-Sentence Video Captioning using Content-oriented Beam Searching and Multi-stage Refining Algorithm
Masoomeh Nabati and Alireza Behrad. 2020a · 2020
Closest in time.
Video captioning using boosted and parallel Long Short-Term Memory networks
Masoomeh Nabati and Alireza Behrad. 2020b · 2020
Closest in time.
Understanding temporal structure for video captioning
Shagan Sah, Thang Nguyen, and Ray Ptucha. 2020 · 2020
Closest in time.
NITS-VC System for VATEX Video Captioning Challenge 2020
Alok Singh, Thoudam Doren Singh, and Sivaji Bandyopadhyay. 2020 · 2020
Closest in time.
Sequence in sequence for video captioning
Huiyun Wang, Chongyang Gao, and Yahong Han. 2020 · 2020
Closest in time.
Exploiting the local temporal information for video captioning
Ran Wei, Li Mi, Yaosi Hu, and Zhenzhong Chen. 2020 · 2020
Closest in time.
Video captioning with text-based dynamic attention and step-by-step learning
Huanhou Xiao and Jinglun Shi. 2020 · 2020
Closest in time.
Exploring diverse and fine-grained caption for video by incorporating convolutional architecture into LSTM-based model
Huanhou Xiao, Junwei Xu, and Jinglun Shi. 2020 · 2020
Closest in time.
Video Description Model Based on Temporal-Spatial and Channel Multi-Attention Mechanisms
Jie Xu, Haoliang Wei, Linke Li, Qiuru Fu, and Jinhong Guo. 2020 · 2020
Closest in time.
Object Relational Graph with Teacher-Recommended Learning for Video Captioning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 13278–13288
Ziqi Zhang, Yaya Shi, Chunfeng Yuan, Bing Li, Peijin Wang, Weiming Hu, and Zheng-Jun Zha. 2020 · 2020
Closest in time.
Video captioning with attention-based LSTM and semantic consistency
Lianli Gao, Zhao Guo, Hanwang Zhang, Xing Xu, and Heng Tao Shen. 2017 · 2055
Closest in time.