Fetching the paper…
Reading the bibliography…
Generating natural language descriptions for videos, i.e., video captioning, essentially requires step-by-step reasoning along the generation process.
Natural language description of human activities from video images based on concept hierarchy of actions
Atsuhiro Kojima, Takeshi Tamura, and Kunio Fukunaga · 2002
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin · 2004
Earlier work this paper cites.
Meteor: An automatic metric for mt evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie · 2005
Earlier work this paper cites.
Collecting highly parallel data for paraphrase evaluation
David L Chen and William B Dolan · 2011
Earlier work this paper cites.
Youtube2text: Recognizing and describing arbitrary activities using semantic hierarchies and zero-shot recognition
Sergio Guadarrama, Niveda Krishnamoorthy, Girish Malkarnenkar, Subhashini Venugopalan, Raymond Mooney, Trevor Darrell, and Kate Saenko · 2013
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2015
Earlier work this paper cites.
Long-term recurrent convolutional networks for visual recognition and description
Jeffrey Donahue, Lisa Anne Hendricks, Sergio Guadarrama, Marcus Rohrbach, Subhashini Venugopalan, Kate Saenko, and Trevor Darrell · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael S. Bernstein, Alexander C. Berg, and Fei-Fei Li · 2015
Earlier work this paper cites.
Cider: Consensus-based image description evaluation
Ramakrishna Vedantam, C Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
Sequence to sequence-video to text
Subhashini Venugopalan, Marcus Rohrbach, Jeffrey Donahue, Raymond Mooney, Trevor Darrell, and Kate Saenko · 2015
Earlier work this paper cites.
Describing videos by exploiting temporal structure
Li Yao, Atousa Torabi, Kyunghyun Cho, Nicolas Ballas, Christopher Pal, Hugo Larochelle, and Aaron Courville · 2015
Earlier work this paper cites.
Neural module networks
Jacob Andreas, Marcus Rohrbach, Trevor Darrell, and Dan Klein · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Faster r-cnn: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun · 2016
Cited alongside, same era.
Msr-vtt: A large video description dataset for bridging video and language
Jun Xu, Tao Mei, Ting Yao, and Yong Rui · 2016
Cited alongside, same era.
Quo vadis, action recognition? a new model and the kinetics dataset
Joao Carreira and Andrew Zisserman · 2017
Cited alongside, same era.
Learning to reason: End-to-end module networks for visual question answering
Ronghang Hu, Jacob Andreas, Marcus Rohrbach, Trevor Darrell, and Kate Saenko · 2017
Cited alongside, same era.
Categorical reparameterization with gumbel-softmax
Eric Jang, Shixiang Gu, and Ben Poole · 2017
Context-aware visual policy network for sequence-level image captioning
Daqing Liu, Zheng-Jun Zha, Hanwang Zhang, Yongdong Zhang, and Feng Wu · 2018
Later among the works it cites.
Reconstruction network for video captioning
Bairui Wang, Lin Ma, Wei Zhang, and Wei Liu · 2018
Later among the works it cites.
Motion guided spatial attention for video captioning
Shaoxiang Chen and Yu-Gang Jiang · 2019
Later among the works it cites.
Learning to compose and reason with language tree structures for visual grounding
Richang Hong, Daqing Liu, Xiaoyu Mo, Xiangnan He, and Hanwang Zhang · 2019
Later among the works it cites.
Joint syntax representation learning and visual cue translation for video captioning
Jingyi Hou, Xinxiao Wu, Wentian Zhao, Jiebo Luo, and Yunde Jia · 2019
Later among the works it cites.
Learning to assemble neural module tree networks for visual grounding
Daqing Liu, Hanwang Zhang, Feng Wu, and Zheng-Jun Zha · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
The kinetics human action video dataset
Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, Mustafa Suleyman, and Andrew Zisserman · 2017
Cited alongside, same era.
Mam-rnn: Multi-level attention model based rnn for video captioning
Xuelong Li, Bin Zhao, Xiaoqiang Lu, et al · 2017
Cited alongside, same era.
Video captioning with transferred semantic attributes
Yingwei Pan, Ting Yao, Houqiang Li, and Tao Mei · 2017
Cited alongside, same era.
Inception-v4, inception-resnet and the impact of residual connections on learning
Christian Szegedy, Sergey Ioffe, Vincent Vanhoucke, and Alexander A Alemi · 2017
Cited alongside, same era.
Bottom-up and top-down attention for image captioning and visual question answering
Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang · 2018
Cited alongside, same era.
Using syntax to ground referring expressions in natural images
Volkan Cirik, Taylor Berg-Kirkpatrick, and Louis-Philippe Morency · 2018
Cited alongside, same era.
Later among the works it cites.
Memory-attended recurrent network for video captioning
Wenjie Pei, Jiyuan Zhang, Xiangrong Wang, Lei Ke, Xiaoyong Shen, and Yu-Wing Tai · 2019
Later among the works it cites.
Image captioning with compositional neural module networks
Junjiao Tian and Jean Oh · 2019
Later among the works it cites.
Controllable video captioning with pos sequence guidance based on gated fusion network
Bairui Wang, Lin Ma, Wei Zhang, Wenhao Jiang, Jingwen Wang, and Wei Liu · 2019
Later among the works it cites.
Making history matter: History-advantage sequence training for visual dialog
T. Yang, Z. Zha, and H. Zhang · 2019
Later among the works it cites.
Learning to collocate neural modules for image captioning
Xu Yang, Hanwang Zhang, and Jianfei Cai · 2019
Later among the works it cites.
Context-aware visual policy network for fine-grained image captioning
Z. Zha, D. Liu, H. Zhang, Y. Zhang, and F. Wu · 2019
Later among the works it cites.
Object-aware aggregation with bidirectional temporal graph for video captioning
Junchao Zhang and Yuxin Peng · 2019
Later among the works it cites.