Fetching the paper…
Reading the bibliography…
Taking full advantage of the information from both vision and language is critical for the video captioning task.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Chin-Yew Lin · 2004
Earlier work this paper cites.
Collecting highly parallel data for paraphrase evaluation
David L. Chen and William B. Dolan · 2011
Earlier work this paper cites.
Meteor universal: Language specific translation evaluation for any target language
Michael J. Denkowski and Alon Lavie · 2014
Earlier work this paper cites.
On using monolingual corpora in neural machine translation
Çaglar Gülçehre, Orhan Firat, Kelvin Xu, Kyunghyun Cho, Loïc Barrault, Huei-Chi Lin, Fethi Bougares, Holger Schwenk, and Yoshua Bengio · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael S. Bernstein, Alexander C. Berg, and Fei-Fei Li · 2015
Earlier work this paper cites.
Cider: Consensus-based image description evaluation
Ramakrishna Vedantam, C. Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
Sequence to sequence-video to text
Subhashini Venugopalan, Marcus Rohrbach, Jeffrey Donahue, Raymond Mooney, Trevor Darrell, and Kate Saenko · 2015
Earlier work this paper cites.
Translating videos to natural language using deep recurrent neural networks
Subhashini Venugopalan, Huijuan Xu, Jeff Donahue, Marcus Rohrbach, Raymond J. Mooney, and Kate Saenko · 2015
Earlier work this paper cites.
Describing videos by exploiting temporal structure
Li Yao, Atousa Torabi, Kyunghyun Cho, Nicolas Ballas, Christopher J. Pal, Hugo Larochelle, and Aaron C. Courville · 2015
Earlier work this paper cites.
Hierarchical recurrent neural encoder for video representation with application to captioning
Pingbo Pan, Zhongwen Xu, Yi Yang, Fei Wu, and Yueting Zhuang · 2016
Earlier work this paper cites.
Jointly modeling embedding and translation to bridge video and language
Yingwei Pan, Tao Mei, Ting Yao, Houqiang Li, and Yong Rui · 2016
Earlier work this paper cites.
MSR-VTT: A large video description dataset for bridging video and language
Jun Xu, Tao Mei, Ting Yao, and Yong Rui · 2016
Earlier work this paper cites.
Video paragraph captioning using hierarchical recurrent neural networks
Haonan Yu, Jiang Wang, Zhiheng Huang, Yi Yang, and Wei Xu · 2016
Earlier work this paper cites.
The kinetics human action video dataset
Will Kay, João Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, Mustafa Suleyman, and Andrew Zisserman · 2017
Earlier work this paper cites.
Semi-supervised classification with graph convolutional networks
Thomas N. Kipf and Max Welling · 2017
Cited alongside, same era.
Inception-v4, inception-resnet and the impact of residual connections on learning
Christian Szegedy, Sergey Ioffe, Vincent Vanhoucke, and Alexander A. Alemi · 2017
Cited alongside, same era.
Petar Velickovic, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Liò, and Yoshua Bengio · 2017
Cited alongside, same era.
Aggregated residual transformations for deep neural networks
Saining Xie, Ross B. Girshick, Piotr Dollár, Zhuowen Tu, and Kaiming He · 2017
Cited alongside, same era.
Catching the temporal regions-of-interest for video captioning
Ziwei Yang, Yahong Han, and Zheng Wang · 2017
Cited alongside, same era.
Less is more: Picking informative frames for video captioning
Non-local neural networks
Xiaolong Wang, Ross B. Girshick, Abhinav Gupta, and Kaiming He · 2018
Later among the works it cites.
Videos as space-time region graphs
Xiaolong Wang and Abhinav Gupta · 2018
Later among the works it cites.
Exploring visual relationship for image captioning
Ting Yao, Yingwei Pan, Yehao Li, and Tao Mei · 2018
Later among the works it cites.
Neural motifs: Scene graph parsing with global context
Rowan Zellers, Mark Yatskar, Sam Thomson, and Yejin Choi · 2018
Later among the works it cites.
Spatio-temporal dynamics and semantic attribute enriched visual encoding for video captioning
Nayyer Aafaq, Naveed Akhtar, Wei Liu, Syed Zulqarnain Gilani, and Ajmal Mian · 2019
Later among the works it cites.
MMDetection: Open mmlab detection toolbox and benchmark
Kai Chen, Jiaqi Wang, Jiangmiao Pang, Yuhang Cao, Yu Xiong, Xiaoxiao Li, Shuyang Sun, Wansen Feng, Ziwei Liu, Jiarui Xu, Zheng Zhang, Dazhi Cheng, Chenchen Zhu, Tianheng Cheng, Qijie Zhao, Buyu Li, Xin Lu, Rui Zhu, Yue Wu, Jifeng Dai, Jingdong Wang, Jianping Shi, Wanli Ouyang, Chen Change Loy, and Dahua Lin · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yangyu Chen, Shuhui Wang, Weigang Zhang, and Qingming Huang · 2018
Cited alongside, same era.
Not all words are equal: Video-specific information loss for video captioning
Jiarong Dong, Ke Gao, Xiaokai Chen, Junbo Guo, Juan Cao, and Yongdong Zhang · 2018
Cited alongside, same era.
Can spatiotemporal 3d cnns retrace the history of 2d cnns and imagenet?
Kensho Hara, Hirokatsu Kataoka, and Yutaka Satoh · 2018
Cited alongside, same era.
An analysis of incorporating an external language model into a sequence-to-sequence model
Anjuli Kannan, Yonghui Wu, Patrick Nguyen, Tara N. Sainath, Zhijeng Chen, and Rohit Prabhavalkar · 2018
Cited alongside, same era.
Sibnet: Sibling convolutional encoder for video captioning
Sheng Liu, Zhou Ren, and Junsong Yuan · 2018
Cited alongside, same era.
Out of the box: Reasoning with graph convolution nets for factual visual question answering
Medhini Narasimhan, Svetlana Lazebnik, and Alexander G. Schwing · 2018
Cited alongside, same era.
Learning conditioned graph structures for interpretable visual question answering
Will Norcliffe-Brown, Stathis Vafeias, and Sarah Parisot · 2018
Cited alongside, same era.
Later among the works it cites.
Motion guided spatial attention for video captioning
Shaoxiang Chen and Yu-Gang Jiang · 2019
Later among the works it cites.
BERT: pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Later among the works it cites.
Joint syntax representation learning and visual cue translation for video captioning
Jingyi Hou, Xinxiao Wu, Wentian Zhao, Jiebo Luo, and Yunde Jia · 2019
Later among the works it cites.
Hierarchical global-local temporal modeling for video captioning
Yaosi Hu, Zhenzhong Chen, Zheng-Jun Zha, and Feng Wu · 2019
Later among the works it cites.
Relation-aware graph attention network for visual question answering
Linjie Li, Zhe Gan, Yu Cheng, and Jingjing Liu · 2019
Later among the works it cites.
Memory-attended recurrent network for video captioning
Wenjie Pei, Jiyuan Zhang, Xiangrong Wang, Lei Ke, Xiaoyong Shen, and Yu-Wing Tai · 2019
Later among the works it cites.
Controllable video captioning with pos sequence guidance based on gated fusion network
Bairui Wang, Lin Ma, Wei Zhang, Wenhao Jiang, Jingwen Wang, and Wei Liu · 2019
Later among the works it cites.
Vatex: A large-scale, high-quality multilingual dataset for video-and-language research
Xin Wang, Jiawei Wu, Junkun Chen, Lei Li, Yuan-Fang Wang, and William Yang Wang · 2019
Later among the works it cites.
Auto-encoding scene graphs for image captioning
Xu Yang, Kaihua Tang, Hanwang Zhang, and Jianfei Cai · 2019
Later among the works it cites.
Object-aware aggregation with bidirectional temporal graph for video captioning
Junchao Zhang and Yuxin Peng · 2019
Later among the works it cites.