Fetching the paper…
Reading the bibliography…
In this paper, we propose to guide the video caption generation with Part-of-Speech (POS) information, based on a gated fusion of multiple representations of input videos.
Natural language description of human activities from video images based on concept hierarchy of actions
Atsuhiro Kojima, Takeshi Tamura, and Kunio Fukunaga · 2002
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
Feature-rich part-of-speech tagging with a cyclic dependency network
Kristina Toutanova, Dan Klein, Christopher D Manning, and Yoram Singer · 2003
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Chin-Yew Lin · 2004
Earlier work this paper cites.
Meteor: An automatic metric for mt evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie · 2005
Earlier work this paper cites.
Collecting highly parallel data for paraphrase evaluation
David L Chen and William B Dolan · 2011
Earlier work this paper cites.
Adadelta: an adaptive learning rate method
Matthew D Zeiler · 2012
Earlier work this paper cites.
Youtube2text: Recognizing and describing arbitrary activities using semantic hierarchies and zero-shot recognition
Sergio Guadarrama, Niveda Krishnamoorthy, Girish Malkarnenkar, Subhashini Venugopalan, Raymond Mooney, Trevor Darrell, and Kate Saenko · 2013
Earlier work this paper cites.
Tv-l1 optical flow estimation
Javier Sánchez Pérez, Enric Meinhardt-Llopis, and Gabriele Facciolo · 2013
Earlier work this paper cites.
Translating video content to natural language descriptions
Marcus Rohrbach, Wei Qiu, Ivan Titov, Stefan Thater, Manfred Pinkal, and Bernt Schiele · 2013
Earlier work this paper cites.
Coherent multi-sentence video description with variable level of detail
Anna Rohrbach, Marcus Rohrbach, Wei Qiu, Annemarie Friedrich, Manfred Pinkal, and Bernt Schiele · 2014
Earlier work this paper cites.
Microsoft coco captions: Data collection and evaluation server
Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Dollár, and C Lawrence Zitnick · 2015
Earlier work this paper cites.
Long-term recurrent convolutional networks for visual recognition and description
Jeffrey Donahue, Lisa Anne Hendricks, Sergio Guadarrama, Marcus Rohrbach, Subhashini Venugopalan, Kate Saenko, and Trevor Darrell · 2015
Earlier work this paper cites.
Multimodal convolutional neural networks for matching image and sentence
Lin Ma, Zhengdong Lu, Lifeng Shang, and Hang Li · 2015
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al · 2015
Earlier work this paper cites.
Learning spatiotemporal features with 3d convolutional networks
Du Tran, Lubomir Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri · 2015
Earlier work this paper cites.
Cider: Consensus-based image description evaluation
Ramakrishna Vedantam, C Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
Sequence to sequence-video to text
Subhashini Venugopalan, Marcus Rohrbach, Jeffrey Donahue, Raymond Mooney, Trevor Darrell, and Kate Saenko · 2015
Earlier work this paper cites.
Translating videos to natural language using deep recurrent neural networks
Subhashini Venugopalan, Huijuan Xu, Jeff Donahue, Marcus Rohrbach, Raymond Mooney, and Kate Saenko · 2015
Earlier work this paper cites.
Jointly modeling deep video and compositional text to bridge vision and language in a unified framework
Ran Xu, Caiming Xiong, Wei Chen, and Jason J Corso · 2015
Cited alongside, same era.
Describing videos by exploiting temporal structure
Li Yao, Atousa Torabi, Kyunghyun Cho, Nicolas Ballas, Christopher Pal, Hugo Larochelle, and Aaron Courville · 2015
Cited alongside, same era.
Multimodal compact bilinear pooling for visual question answering and visual grounding
Akira Fukui, Dong Huk Park, Daylen Yang, Anna Rohrbach, Trevor Darrell, and Marcus Rohrbach · 2016
Cited alongside, same era.
Describing videos using multi-modal fusion
Qin Jin, Jia Chen, Shizhe Chen, Yifan Xiong, and Alexander Hauptmann · 2016
Cited alongside, same era.
Learning to answer questions from image using convolutional neural network
Lin Ma, Zhengdong Lu, and Hang Li · 2016
Cited alongside, same era.
Jointly modeling embedding and translation to bridge video and language
Regularizing rnns for caption generation by reconstructing the past with the present
Xinpeng Chen, Lin Ma, Wenhao Jiang, Jian Yao, and Wei Liu · 2018
Later among the works it cites.
Less is more: Picking informative frames for video captioning
Yangyu Chen, Shuhui Wang, Weigang Zhang, and Qingming Huang · 2018
Later among the works it cites.
Embodied question answering
Abhishek Das, Samyak Datta, Georgia Gkioxari, Stefan Lee, Devi Parikh, and Dhruv Batra · 2018
Later among the works it cites.
Recurrent fusion network for image captioning
Wenhao Jiang, Lin Ma, Yu-Gang Jiang, Wei Liu, and Tong Zhang · 2018
Later among the works it cites.
Sibnet: Sibling convolutional encoder for video captioning
Sheng Liu, Zhou Ren, and Junsong Yuan · 2018
Later among the works it cites.
Quantization-based hashing: a general framework for scalable image and video retrieval
Jingkuan Song, Lianli Gao, Li Liu, Xiaofeng Zhu, and Nicu Sebe · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yingwei Pan, Tao Mei, Ting Yao, Houqiang Li, and Yong Rui · 2016
Cited alongside, same era.
Multimodal video description
Vasili Ramanishka, Abir Das, Dong Huk Park, Subhashini Venugopalan, Lisa Anne Hendricks, Marcus Rohrbach, and Kate Saenko · 2016
Cited alongside, same era.
Frame-and segment-level features and candidate pool evaluation for video caption generation
Rakshith Shetty and Jorma Laaksonen · 2016
Cited alongside, same era.
Msr-vtt: A large video description dataset for bridging video and language
Jun Xu, Tao Mei, Ting Yao, and Yong Rui · 2016
Cited alongside, same era.
Video paragraph captioning using hierarchical recurrent neural networks
Haonan Yu, Jiang Wang, Zhiheng Huang, Yi Yang, and Wei Xu · 2016
Cited alongside, same era.
Quo vadis, action recognition? a new model and the kinetics dataset
Joao Carreira and Andrew Zisserman · 2017
Cited alongside, same era.
Video captioning with guidance of multimodal latent topics
Shizhe Chen, Jia Chen, Qin Jin, and Alexander Hauptmann · 2017
Cited alongside, same era.
Later among the works it cites.
Reconstruction network for video captioning
Bairui Wang, Lin Ma, Wei Zhang, and Wei Liu · 2018
Later among the works it cites.
Bidirectional attentive fusion with context gating for dense video captioning
Jingwen Wang, Wenhao Jiang, Lin Ma, Wei Liu, and Yong Xu · 2018
Later among the works it cites.
M3: Multimodal memory modelling for video captioning
Junbo Wang, Wei Wang, Yan Huang, Liang Wang, and Tieniu Tan · 2018
Later among the works it cites.
A survey on learning to hash
Jingdong Wang, Ting Zhang, Nicu Sebe, Jingkuang Song, and Heng Tao Shen · 2018
Later among the works it cites.
Interpretable video captioning via trajectory structured localization
Xian Wu, Guanbin Li, Qingxing Cao, Qingge Ji, and Liang Lin · 2018
Later among the works it cites.
Motion guided spatial attention for video captioning
Shaoxiang Chen and Yu-Gang Jiang · 2019
Closest in time.
Diverse and controllable image captioning with part-of-speech guidance
Aditya Deshpande, Jyoti Aneja, Liwei Wang, Alexander Schwing, and David A Forsyth · 2019
Closest in time.
Unsupervised image captioning
Yang Feng, Lin Ma, Wei Liu, and Jiebo Luo · 2019
Closest in time.
Image caption generation with part of speech guidance
Xinwei He, Baoguang Shi, Xiang Bai, Gui-Song Xia, Zhaoxiang Zhang, and Weisheng Dong · 2019
Closest in time.
Matching image and sentence with multi-faceted representations
Lin Ma, Wenhao Jiang, Zequn Jie, Yu-Gang Jiang, and Wei Liu · 2019
Closest in time.
Hierarchical photo-scene encoder for album storytelling
Bairui Wang, Lin Ma, Wei Zhang, Wenhao Jiang, and Feng Zhang · 2019
Closest in time.
Reconstruct and represent video contents for captioning via reinforcement learning
Wei Zhang, Bairui Wang, Lin Ma, and Wei Liu · 2019
Closest in time.