Fetching the paper…
Reading the bibliography…
Video description is the automatic generation of natural language sentences that describe the contents of a given video.
F. Nishida and S. Takamatsu. 1982. Japanese-English translation through internal expressions. In Proceedings of the 9th conference on Computational linguistics-Volume 1. Academia Praha, 271-276
1982
Earlier work this paper cites.
F. Nishida, S. Takamatsu, T. Tani, and T. Doi. 1988. Feedback of correcting information in post editing to a machine translation system. In Proceedings of the 12th conference on Computational linguistics-Volume 2. ACL, 476-481
1988
Earlier work this paper cites.
D. Koller, N. Heinze, and H. Nagel. 1991. Algorithmic characterization of vehicle trajectories from image sequences by motion verbs. In IEEE Computer Society Conference on CVPR. 90-95
1991
Earlier work this paper cites.
C. Tomasi and T. Kanade. 1991. Detection and tracking of point features
1991
Earlier work this paper cites.
C. Pollard and I. A. Sag. 1994. Head-driven phrase structure grammar. University of Chicago Press
1994
Earlier work this paper cites.
J. Shi and C. Tomasi. 1994. Good features to track. In IEEE CVPR
1994
Earlier work this paper cites.
A. F. Bobick and A. D. Wilson. 1997. A state-based approach to the representation and recognition of gesture. IEEE TPAMI 19, 12 (1997), 1325-1337
1997
Earlier work this paper cites.
M. Brand. 1997. The” Inverse hollywood problem”: from video to scripts and storyboards via causal analysis. In AAAI/IAAI. Citeseer, 132-137
1997
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber. 1997. Long Short-Term Memory. Neural Computation 9, 8 (1997), 1735-1780
1997
Earlier work this paper cites.
M. Schuster and K. K. Paliwal. 1997. Bidirectional Recurrent Neural Networks. IEEE Transactions on Signal Processing, Vol. 45, 11, 2673-2681
1997
Earlier work this paper cites.
C. Fellbaum. 1998. WordNet. Wiley Online Library
1998
Earlier work this paper cites.
C. S. Pinhanez and A. F. Bobick. 1998. Human action detection using pnf propagation of temporal constraints. In IEEE Computer Society Conference on CVPR
1998
Earlier work this paper cites.
D. G. Lowe. 1999. Object recognition from local scale-invariant features. In IEEE ICCV
1999
Earlier work this paper cites.
S. Hongeng, F. Brémond, and R. Nevatia. 2000. Bayesian framework for video surveillance application. In Pattern Recognition, 2000. Proceedings. 15th International Conference on, Vol. 1. IEEE, 164-170
2000
Earlier work this paper cites.
E. Reiter and R. Dale. 2000. Building natural language generation systems. Cambridge university press
2000
Earlier work this paper cites.
Y. Rubner, C. Tomasi, and L. J. Guibas. 2000. The earth mover’s distance as a metric for image retrieval. IJCV, Vol. 40, 2, 99-121
2000
Earlier work this paper cites.
P. Viola and M. Jones. 2001. Rapid object detection using a boosted cascade of simple features. In IEEE CVPR
2001
Earlier work this paper cites.
A. Kojima, T. Tamura, and K. Fukunaga. 2002. Natural language description of human activities from video images based on concept hierarchy of actions. IJCV 50, 2 (2002), 171-184
2002
Earlier work this paper cites.
P. Kuchi, P. Gabbur, P. S. Bhat, and S. S. David. 2002. Human face detection and tracking using skin color modeling and connected component operators. IETE Journal of Research 48, 3-4 (2002), 289–293
2002
Earlier work this paper cites.
D. Moore and I. Essa. 2002. Recognizing multitasked activities from video using stochastic context-free grammar. In AAAI/IAAI. 770-776
2002
Earlier work this paper cites.
K. Papineni, S. Roukos, T. Ward, and W. Zhu. 2002. BLEU: a method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting on ACL. 311-318
2002
Earlier work this paper cites.
K. Barnard, P. Duygulu, D. Forsyth, N. D. Freitas, D. M. Blei, and M. I. Jordan. 2003. Matching words and pictures. Journal of Machine Learning Research 3, Feb (2003), 1107-1135
2003
Earlier work this paper cites.
S. Gong and T. Xiang. 2003. Recognition of group activities using dynamic probabilistic networks. In IEEE ICCV
2003
Earlier work this paper cites.
A. Torralba, K. P. Murphy, W. T. Freeman, and M. A. Rubin. 2003. Context-based vision system for place and object recognition. In IEEE ICCV
2003
Earlier work this paper cites.
T. Berg, A. Berg, J. Edwards, M. Maire, R. White, Y. Teh, E. Learned-Miller, and D. A. Forsyth. 2004. Names and faces in the news. In IEEE CVPR
2004
Earlier work this paper cites.
A. Hakeem, Y. Sheikh, and M. Shah. 2004. CASE E
2004
Earlier work this paper cites.
C. Lin. 2004. Rouge: A package for automatic evaluation of summaries. in: Text Summarization Branches Out
2004
Earlier work this paper cites.
R. Nevatia, J. Hobbs, and B. Bolles. 2004. An ontology for video event representation. In CVPR Workshop. 119-119
2004
Earlier work this paper cites.
S. Robertson. 2004. Understanding inverse document frequency: on theoretical arguments for IDF. Journal of documentation, Vol. 60, 5, 503–520
2004
Earlier work this paper cites.
S. Banerjee and A. Lavie. 2005. METEOR: An automatic metric for MT evaluation with improved correlation with human judgments. ACL workshop on intrinsic and extrinsic evaluation measures for MT and/or summarization. 65-72
2005
Earlier work this paper cites.
N. Dalal and B. Triggs. 2005. Histograms of oriented gradients for human detection. In IEEE Computer Society Conference on CVPR
2005
Earlier work this paper cites.
D. Roy. 2005. Semiotic schemas: A framework for grounding language in action and perception. Artificial Intelligence, Vol. 167(1-2), 170-205
2005
Earlier work this paper cites.
D. Roy and E. Reiter. 2005. Connecting Language to the World. Artificial Intelligence, Vol. 167(1-2), 1-12
2005
Earlier work this paper cites.
N. Dalal, B. Triggs, and C. Schmid. 2006. Human detection using oriented histograms of flow and appearance. In IEEE ECCV
2006
Earlier work this paper cites.
P. Koehn, H. Hoang, A. Birch, C. Callison-Burch, M. Federico, N. Bertoldi, B. Cowan, WadeShen, C. Moran, R. Zens, et al. 2007. Moses: Open source toolkit for statistical machine translation. In Proceedings of the 45th annual meeting of the ACL on interactive poster and demonstration sessions.ACL, 177-180
2007
Earlier work this paper cites.
P. Felzenszwalb, D. McAllester, and D. Ramanan. 2008. A discriminatively trained, multiscale, deformable part model. In IEEE CVPR
2008
Earlier work this paper cites.
M. W. Lee, A. Hakeem, N. Haering, and S. Zhu. 2008. Save: A framework for semantic annotation of visual events. In IEEE Computer Society Conference on CVPR Workshops. 1-8
2008
Earlier work this paper cites.
R. Chaudhry, A. Ravichandran, G. Hager, R. Vidal. 2009. Histograms of oriented optical flow and Binet-Cauchy kernels on nonlinear dynamical systems for the recognition of human actions, CVPR 2009
2009
Earlier work this paper cites.
J. Deng, K. Li, M. Do, H. Su, and L. Fei-Fei. 2009. Construction and analysis of a large scale image ontology. Vision Sciences Society 186, 2 (2009)
2009
Earlier work this paper cites.
I. Maglogiannis, D. Vouyioukas, and C. Aggelopoulos. 2009. Face detection and recognition of natural human emotion using Markov random fields. Personal and Ubiquitous Computing, Vol. 13, 1, 95-101
2009
Earlier work this paper cites.
H. Wang, M. M. Ullah, A. Klaser, I. Laptev, and C. Schmid. 2009. Evaluation of local spatio-temporal features for action recognition. In BMVC 2009-British Machine Vision Conference. BMVA Press, 124-1
2009
Earlier work this paper cites.
D. Chen, W. Dolan, S. Raghavan, T. Huynh, and R. Mooney. 2010. Collecting highly parallel data for paraphrase evaluation. In JAIR: - Volume 37. ACL, 397-435
2010
Earlier work this paper cites.
M. Everingham, L. V. Gool, C. K. I. Williams, J. Winn, and A. Zisserman. 2010. The pascal visual object classes (voc) challenge. IJCV 88, 2 (2010), 303-338
2010
Earlier work this paper cites.
A. Farhadi, M. Hejrati, M. A. Sadeghi, P. Young, C. Rashtchian, J. Hockenmaier, and D. Forsyth. 2010. Every picture tells a story: Generating sentences from images. In IEEE ECCV
2010
Earlier work this paper cites.
P. F. Felzenszwalb, R. B. Girshick, and D. McAllester. 2010. Cascade object detection with deformable part models. In IEEE CVPR
2010
Earlier work this paper cites.
P. F. Felzenszwalb, R. B. Girshick, D. McAllester, and D. Ramanan. 2010. Object detection with discriminatively trained part-based models. IEEE TPAMI 32, 9 (2010), 1627-1645
2010
Earlier work this paper cites.
W. Kim, J. Park, and C. Kim. 2010. A novel method for efficient indoor–outdoor image classification. Journal of Signal Processing Systems 61, 3 (2010), 251-258
2010
Earlier work this paper cites.
C. Matuszek, D. Fox, and K. Koscher. 2010. Following directions using statistical machine translation. In 5th ACM/IEEE International Conference on Human-Robot Interaction (HRI)
2010
Earlier work this paper cites.
D. Chen and W. Dolan. 2011. Collecting highly parallel data for paraphrase evaluation. In ACL: Human Language Technologies-Volume 1. ACL, 190-200
2011
Earlier work this paper cites.
M. U. G. Khan, L. Zhang, and Y. Gotoh. 2011. Human focused video description. In IEEE International Conference on Computer Vision Workshops (ICCV Workshops)
2011
Earlier work this paper cites.
G. Kulkarni, V. Premraj, S. Dhar, S. Li, Y. Choi, A. C. Berg, and T. L. Berg. 2011. Baby talk: Understanding and generating image descriptions. In IEEE CVPR
2011
Earlier work this paper cites.
S. Li, G. Kulkarni, T. L. Berg, A. C. Berg, and Y. Choi. 2011. Composing simple image descriptions using web-scale n-grams. In Conference on Computational Natural Language Learning (CNLL)
2011
Earlier work this paper cites.
S. Tellex, T. Kollar, S. Dickerson, M. R. Walter, Ashis Gopal Banerjee, Seth J Teller, and Nicholas Roy. 2011. Understanding Natural Language Commands for Robotic Navigation and Mobile Manipulation. In AAAI
2011
Earlier work this paper cites.
2012
Earlier work this paper cites.
G. Burghouts, H. Bouma, R. D. Hollander, S V. D. Broek, and K. Schutte. 2012. Recognition of 48 human behaviors from video. In Int. Symp. Optronics in Defense and Security, OPTRO
2012
Earlier work this paper cites.
D. Ding, F. Metze, S. Rawat, P. F. Schulam, S. Burger, E. Younessian, L. Bao, M. G. Christel, and A. Hauptmann. 2012. Beyond audio and video retrieval: towards multimedia summarization. In 2nd ACM International Conference on Multimedia Retrieval (ICMR)
2012
Earlier work this paper cites.
P. Hanckmann, K. Schutte, and G. J. Burghouts. 2012. Automated textual descriptions for a wide range of video events with 48 human actions. In IEEE ECCV
2012
Earlier work this paper cites.
M. U. G. Khan and Y. Gotoh. 2012. Describing video contents in natural language. In Workshop on Innovative Hybrid Approaches to the Processing of Textual Data. ACL, 27-35
2012
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton. 2012. ImageNet classification with deep convolutional neural networks. In Advances in Neural Information Processing Systems. 1097-1105
2012
Earlier work this paper cites.
M. Rohrbach, S. Amin, M. Andriluka, and B. Schiele. 2012. A database for fine grained activity detection of cooking activities. In IEEE CVPR
2012
Earlier work this paper cites.
M. Rohrbach, M. Regneri, M. Andriluka, S. Amin, M. Pinkal, and B. Schiele. 2012. Script data for attribute-based recognition of composite activities. In IEEE ECCV
2012
Earlier work this paper cites.
S. Zhu and D. Mumford. 2007. A stochastic grammar of images. Foundations and Trends in Computer Graphics and Vision, Vol. 2, 4, 259-362
2012
Earlier work this paper cites.
P. Das, C. Xu, R. F. Doell, and J. J. Corso. 2013. A thousand frames in just a few words: Lingual description of videos through latent topics and sparse object stitching. In IEEE CVPR
2013
Earlier work this paper cites.
A. Graves, A. Mohamed, and G. Hinton. 2013. Speech recognition with deep recurrent neural networks. In IEEE Int Conference on Acoustics, Speech and Signal Processing (ICASSP). 6645-6649
2013
Earlier work this paper cites.
S. Guadarrama, N. Krishnamoorthy, G. Malkarnenkar, S. Venugopalan, R. Mooney, T. Darrell, and K. Saenko. 2013. Recognizing and describing activities using semantic hierarchies and zero-shot recognition. In IEEE ICCV
2013
Cited alongside, same era.
S. Guadarrama, L. Riano, D. Golland, D. Go, Y. Jia, D. Klein, P. Abbeel, T. Darrell, et al. 2013. Grounding spatial relations for human-robot interaction. In Intelligent Robots and Systems (IROS). 1640-1647
2013
Cited alongside, same era.
L. Han, A. L. Kashyap, T. Finin, J. Mayfield, and J. Weese. 2013. UMBC-EBIQUITY-CORE: semantic textual similarity systems. In Second Joint Conference on Lexical and Computational Semantics (* SEM), Volume 1: Proceedings of the Main Conference and the Shared Task: Semantic Textual Similarity, Vol. 1. 44-52
2013
Cited alongside, same era.
N. Krishnamoorthy, G. Malkarnenkar, R. J. Mooney, K. Saenko, and S. Guadarrama. 2013. Generating Natural-Language Video Descriptions Using Text-Mined Knowledge. In AAAI, Vol. 1. 2
Y. Liu and Z. Shi. 2016. Boosting video description generation by explicitly translating from frame-level captions. In Proceedings of the 2016 ACM on Multimedia Conference. ACM, 631-634
2016
Later among the works it cites.
P. Pan, Z. Xu, Y. Yang, F. Wu, and Y. Zhuang. 2016. Hierarchical recurrent neural encoder for video representation with application to captioning. In IEEE CVPR
2016
Later among the works it cites.
Y. Pan, T. Mei, T. Yao, H. Li, and Y. Rui. 2016. Jointly modeling embedding and translation to bridge video and language. In IEEE CVPR
2016
Later among the works it cites.
V. Ramanishka, A. Das, D. H. Park, S. Venugopalan, L. A. Hendricks, M. Rohrbach, and K. Saenko. 2016. Multimodal video description. In Proceedings of ACM on Multimedia Conference. ACM, 1092-1096
2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2013
Cited alongside, same era.
T. Mikolov, I. Sutskever, K. Chen, G. S. Corrado, and J. Dean. 2013. Distributed representations of words and phrases and their compositionality. In Advances in Neural Information Processing Systems. 3111-3119
2013
Cited alongside, same era.
2013
Cited alongside, same era.
M. Regneri, M. Rohrbach, D. Wetzel, S. Thater, B. Schiele, and M. Pinkal. 2013. Grounding action descriptions in videos. Transactions of the Association for Computational Linguistics, Vol. 1, 25-36
2013
Cited alongside, same era.
M. Rohrbach, W. Qiu, I. Titov, S. Thater, M. Pinkal, and B. Schiele. 2013. Translating video content to natural language descriptions. In IEEE ICCV
2013
Cited alongside, same era.
H. Yu and J. M. Siskind. 2013. Grounded Language Learning from Video Sentences. In ACL(1). 53-63
2013
Cited alongside, same era.
Casting words transcription service, 2014. http://castingwords.com/
2014
Cited alongside, same era.
2014
Cited alongside, same era.
2014
Cited alongside, same era.
2016
Later among the works it cites.
A. Shin, K. Ohnishi, and T. Harada. 2016. Beyond caption to narrative: Video captioning with multiple sentences. In IEEE International Conference on Image Processing (ICIP)
2016
Later among the works it cites.
G. A. Sigurdsson, G. Varol, X. Wang, A. Farhadi, I. Laptev, and A. Gupta. 2016. Hollywood in homes: Crowdsourcing data collection for activity understanding. In IEEE ECCV
2016
Later among the works it cites.
2016
Later among the works it cites.
J. K. Wang and R. Gaizauskas. 2016. Cross-validating Image Description Datasets and Evaluation Metrics. In Proceedings of 10th Language Resources and Evaluation Conference. European Language Resources Association, 3059-3066
2016
Later among the works it cites.
Q. Wu, P. Wang, C. Shen, A. Dick, A. Hengel. 2016. Ask Me Anything: Free-form Visual Question Answering Based on Knowledge from External Sources. In IEEE CVPR
2016
Later among the works it cites.
J. Xu, T. Mei, T. Yao, and Y. Rui. 2016. MSR-VTT: A large video description dataset for bridging video and language. In IEEE CVPR
2016
Later among the works it cites.
H. Yu, J. Wang, Z. Huang, Y. Yang, and W. Xu. 2016. Video paragraph captioning using hierarchical recurrent neural networks. In IEEE CVPR
2016
Later among the works it cites.
2016
Later among the works it cites.
K. Zeng, T. Chen, J. C. Niebles, and M. Sun. 2016. Title Generation for User Generated Videos. In IEEE ECCV
2016
Later among the works it cites.
Activity Net Captions Challenge, Evaluations, 2017. https://github.com/ranjaykrishna/densevid_eval
2017
Later among the works it cites.
Activity Net Captions Challenge, Results, 2017. http://activity-net.org/challenges/2017/evaluation.html
2017
Later among the works it cites.
Activity Net Captions Challenge, Task5: Dense-Captioning Events in Videos, 2017. http://activity-net.org/challenges/2017/index.html
2017
Later among the works it cites.
Activity Net Challenge, 2017. http://activity-net.org/challenges/2017/index.html
2017
Later among the works it cites.
Language in Vision, 2017. https://www.sciencedirect.com/journal/computer-vision-and-image-understanding/vol/163
2017
Later among the works it cites.
The Large Scale Movie Description Challenge (LSMDC) Online Evaluations, 2017. https://competitions.codalab.org/competitions/6121#learn_the_details-evaluation
2017
Later among the works it cites.
The Large Scale Movie Description Challenge (LSMDC) Online Results, 2017. https://competitions.codalab.org/competitions/6121#results
2017
Later among the works it cites.
L. Baraldi, C. Grana, and R. Cucchiara. 2017. Hierarchical Boundary-Aware Neural Encoder for Video Captioning. In IEEE CVPR
2017
Later among the works it cites.
O. Bojar, R. Chatterjee, C. Federmann, Y. Graham, B. Haddow, S. Huang, M. Huck, P. Koehn, Q. Liu, V. Logacheva, et al. 2017. 2nd Conference on Machine Translation. 169-214
2017
Later among the works it cites.
A. Das, S. Kottur, K. Gupta, A. Singh, D. Yadav, J. M. F. Moura, D. Parikh, and D. Batra. 2017. Visual Dialog. In IEEE CVPR
2017
Later among the works it cites.
Z. Gan, C. Gan, X. He, Y. Pu, K. Tran, J. Gao, L. Carin, and L. Deng. 2017. Semantic Compositional Networks for visual captioning. In IEEE CVPR
2017
Later among the works it cites.
A. George, B. Asad, F. Jonathan, J. David, D. Andrew, M. Willie, M. Martial, S. Alan, G. Yvette, and K. Wessel. 2017. TRECVID 2017: Evaluating Ad-hoc and Instance Video Search, Events Detection, Video Captioning, and Hyperlinking. In Proceedings of TRECVID 2017
2017
Later among the works it cites.
2017
Later among the works it cites.
2017
Later among the works it cites.
Y. Graham, T. Baldwin, A. Moffat, and J. Zobel. 2017. Can machine translation systems be evaluated by the crowd alone. Natural Language Engineering 23, 1 (2017), 3-30
2017
Later among the works it cites.
J. Johnson, B. Hariharan, L. V. D. Maaten, L. Fei-Fei, C. L. Zitnick, R. Girshick. 2017. CLEVR: A Diagnostic Dataset for Compositional Language and Elementary Visual Reasoning. In IEEE CVPR
2017
Later among the works it cites.
2017
Later among the works it cites.
Y. Pan, T. Yao, H. Li, and T. Mei. 2017. Video Captioning With Transferred Semantic Attributes. In IEEE CVPR
2017
Later among the works it cites.
2017
Later among the works it cites.
2017
Later among the works it cites.
2017
Later among the works it cites.
2017
Later among the works it cites.
A. Rohrbach, A. Torabi, M. Rohrbach, N. Tandon, C. Pal, H. Larochelle, A. Courville, and B. Schiele. 2017. Movie description. IJCV, Vol. 123, 1, 94-120
2017
Later among the works it cites.
Z. Shen, J. Li, Z. Su, M. Li, Y. Chen, Y. Jiang, and X. Xue. 2017. Weakly Supervised Dense Video Captioning. In IEEE CVPR
2017
Later among the works it cites.
2017
Later among the works it cites.
2017
Later among the works it cites.
T. Yao, Y. Li, Z. Qiu, F. Long, Y. Pan, D. Li, and T. Mei. 2017. MSR Asia MSM at ActivityNet Challenge 2017: Trimmed Action Recognition, Temporal Action Proposals and Dense-Captioning Events in Videos
2017
Later among the works it cites.
Y. Yu, J. Choi, Y. Kim, K. Yoo, S. Lee, and G. Kim. 2017. Supervising Neural Attention Models for Video Captioning by Human Gaze Data. In IEEE CVPR
2017
Later among the works it cites.
X. Zhang, K. Gao, Y. Zhang, D. Zhang, J. Li, and Q. Tian. 2017. Task-Driven Dynamic Fusion: Reducing Ambiguity in Video Description. In IEEE CVPR
2017
Later among the works it cites.
B. Andrei, M. Tao, N. Siddharth, Z. Quanshi, S. Nishant, L. Jiebo, and S. Rahul. 2018. A Workshop on Language and Vision at CVPR 2018. http://languageandvision.com/
2018
Closest in time.
2018
Closest in time.
S. Gella, M. Lewis, and M. Rohrbach. 2018. A Dataset for Telling the Stories of Social Media Videos. In Proc of the 2018 Conference on Empirical Methods in Natural Language Processing. 968–974
2018
Closest in time.
D. Harwath, A. Recasens, D. Suris, G. Chuang, A. Torralba, and J. Glass. 2018. Jointly Discovering Visual Objects and Spoken Words from Raw Sensory Input. In IEEE ECCV
2018
Closest in time.
Drew A. Hudson, Christopher D. Manning. 2018. Compositional Attention Networks for Machine Reasoning. In ICLR
2018
Closest in time.
J. Kim, A. Rohrbach, T. Darrell, J. Canny, Z. Akata. 2018. Textual Explanations for Self-Driving Vehicles, ECCV 2018
2018
Closest in time.
2018
Closest in time.
Y. Li, T. Yao, Y. Pan, H. Chao, and T. Mei. 2018. Jointly Localizing and Describing Events for Dense Video Captioning. In IEEE CVPR
2018
Closest in time.
M. Margaret, M. Ishan, H. Ting-Hao, and F. Frank. 2018. Story Telling Workshop and Visual Story Telling Challenge at NAACL 2018
2018
Closest in time.
A. Owens, A. A. Efros. 2018. Audio-Visual Scene Analysis with Self-Supervised Multisensory Features. In IEEE ECCV
2018
Closest in time.
2018
Closest in time.
B. Wang, L. Ma, W. Zhang, and W. Liu. 2018. Reconstruction Network for Video Captioning. In IEEE CVPR
2018
Closest in time.
J. Wang, W. Jiang, L. Ma, W. Liu, and Y. Xu. 2018. Bidirectional Attentive Fusion with Context Gating for Dense Video Captioning. In IEEE CVPR
2018
Closest in time.
J. Wang, W. Wang, Y. Huang, L. Wang, T. Tan. 2018. M3: Multimodal Memory Modelling for Video Captioning. CVPR
2018
Closest in time.
X. Wu, G. Li, Q. Cao, Q. Ji, and L. Lin. 2018. Interpretable Video Captioning via Trajectory Structured Localization. In IEEE CVPR
2018
Closest in time.
2018
Closest in time.
2018
Closest in time.
L. Zhou, C. Xu, and J. J. Corso. 2018. Towards automatic learning of procedures from web instructional videos. In Thirty-Second AAAI Conference on Artificial Intelligence
2018
Closest in time.
L. Zhou, Y. Zhou, J. J. Corso, R. Socher, and C. Xiong. 2018. End-to-End Dense Video Captioning with Masked Transformer. In IEEE CVPR
2018
Closest in time.
N. Aafaq, N. Akhtar, W. Liu, S. Z. Gilani and A. Mian. 2019. Spatio-Temporal Dynamics and Semantic Attribute Enriched Visual Encoding for Video Captioning. In IEEE CVPR
2019
Closest in time.