Fetching the paper…
Reading the bibliography…
While significant progress has been made in the image captioning task, video description is still in its infancy due to the complex nature of video data.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Actor-critic algorithms
Vijay R Konda and John N Tsitsiklis · 2000
Earlier work this paper cites.
Natural language description of human activities from video images based on concept hierarchy of actions
A. Kojima, T. Tamura, and K. Fukunaga · 2002
Earlier work this paper cites.
BLEU: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei jing Zhu · 2002
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Generative adversarial nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Meteor universal: Language specific translation evaluation for any target language
Michael Denkowski Alon Lavie · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher Manning · 2014
Earlier work this paper cites.
Coherent multi-sentence video description with variable level of detail
Anna Rohrbach, Marcus Rohrbach, Wei Qiu, Annemarie Friedrich, Manfred Pinkal, and Bernt Schiele · 2014
Earlier work this paper cites.
Activitynet: A large-scale video benchmark for human activity understanding
Fabian Caba Heilbron, Victor Escorcia, Bernard Ghanem, and Juan Carlos Niebles · 2015
Earlier work this paper cites.
Long-term recurrent convolutional networks for visual recognition and description
Jeff Donahue, Lisa Anne Hendricks, Sergio Guadarrama, Marcus Rohrbach, Subhashini Venugopalan, Kate Saenko, and Trevor Darrell · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun · 2015
Earlier work this paper cites.
The long-short story of movie description
Anna Rohrbach, Marcus Rohrbach, and Bernt Schiele · 2015
Earlier work this paper cites.
Cider: Consensus-based image description evaluation
Ramakrishna Vedantam, C Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
Sequence to sequence – video to text
Subhashini Venugopalan, Marcus Rohrbach, Jeff Donahue, Raymond Mooney, Trevor Darrell, and Kate Saenko · 2015
Earlier work this paper cites.
Describing videos by exploiting temporal structure
Li Yao, Atousa Torabi, Kyunghyun Cho, Nicolas Ballas, Christopher Pal, Hugo Larochelle, and Aaron Courville · 2015
Earlier work this paper cites.
Reasoning about pragmatics with neural listeners and speakers
Jacob Andreas and Dan Klein · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Categorical reparameterization with gumbel-softmax
Eric Jang, Shixiang Gu, and Ben Poole · 2016
Earlier work this paper cites.
Re-evaluating automatic metrics for image captioning
Mert Kilickaya, Aykut Erdem, Nazli Ikizler-Cinbis, and Erkut Erdem · 2016
Earlier work this paper cites.
Hierarchical recurrent neural encoder for video representation with application to captioning
Pingbo Pan, Zhongwen Xu, Yi Yang, Fei Wu, and Yueting Zhuang · 2016
Earlier work this paper cites.
Sequence level training with recurrent neural networks
Marc’Aurelio Ranzato, Sumit Chopra, Michael Auli, and Wojciech Zaremba · 2016
Earlier work this paper cites.
Generative adversarial text to image synthesis
Scott Reed, Zeynep Akata, Xinchen Yan, Lajanugen Logeswaran, Bernt Schiele, and Honglak Lee · 2016
Earlier work this paper cites.
Frame- and segment-level features and candidate pool evaluation for video caption generation
Rakshith Shetty and Jorma Laaksonen · 2016
Earlier work this paper cites.
Beyond caption to narrative: Video captioning with multiple sentences
Andrew Shin, Katsunori Ohnishi, and Tatsuya Harada · 2016
Cited alongside, same era.
Temporal segment networks: Towards good practices for deep action recognition
Limin Wang, Yuanjun Xiong, Zhe Wang, Yu Qiao, Dahua Lin, Xiaoou Tang, and Luc Van Gool · 2016
Cited alongside, same era.
Video paragraph captioning using hierarchical recurrent neural networks
Haonan Yu, Jiang Wang, Zhiheng Huang, Yi Yang, and Wei Xu · 2016
Cited alongside, same era.
Spatio-temporal attention models for grounded video captioning
Mihai Zanfir, Elisabeta Marinoiu, and Cristian Sminchisescu · 2016
Cited alongside, same era.
Maximum-likelihood augmented discrete generative adversarial networks
Tong Che, Yanran Li, Ruixiang Zhang, R Devon Hjelm, Wenjie Li, Yangqiu Song, and Yoshua Bengio · 2017
Cited alongside, same era.
Massimo Caccia, Lucas Caccia, William Fedus, Hugo Larochelle, Joelle Pineau, and Laurent Charlin · 2018
Closest in time.
Improving image captioning with conditional generative adversarial nets
Chen Chen, Shuai Mu, Wanpeng Xiao, Zexiong Ye, Liesi Wu, Fuming Ma, and Qi Ju · 2018
Closest in time.
Less is more: Picking informative frames for video captioning
Yangyu Chen, Shuhui Wang, Weigang Zhang, and Qingming Huang · 2018
Closest in time.
Learning to evaluate image captioning
Yin Cui, Guandao Yang, Andreas Veit, Xun Huang, and Serge Belongie · 2018
Closest in time.
A dataset for telling the stories of social media videos
Spandana Gella, Mike Lewis, and Marcus Rohrbach · 2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Towards diverse and natural image descriptions via a conditional gan
Bo Dai, Sanja Fidler, Raquel Urtasun, and Dahua Lin · 2017
Cited alongside, same era.
Long-term recurrent convolutional networks for visual recognition and description
Jeff Donahue, Lisa Anne Hendricks, Sergio Guadarrama, Marcus Rohrbach, Subhashini Venugopalan, Kate Saenko, and Trevor Darrell · 2017
Cited alongside, same era.
Image-to-image translation with conditional adversarial networks
Phillip Isola, Jun-Yan Zhu, Tinghui Zhou, and Alexei A Efros · 2017
Cited alongside, same era.
The kinetics human action video dataset
Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, et al · 2017
Cited alongside, same era.
Hadamard product for low-rank bilinear pooling
Jin-Hwa Kim, Kyoung Woon On, Woosang Lim, Jeonghee Kim, JungWoo Ha, and Byoung-Tak Zhang · 2017
Cited alongside, same era.
Dense-captioning events in videos
Ranjay Krishna, Kenji Hata, Frederic Ren, Li Fei-Fei, and Juan Carlos Niebles · 2017
Cited alongside, same era.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, et al · 2017
Cited alongside, same era.
Improving reinforcement learning based image captioning with natural language prior
Tszhang Guo, Shiyu Chang, Mo Yu, and Kun Bai · 2018
Closest in time.
Can spatiotemporal 3d cnns retrace the history of 2d cnns and imagenet
Kensho Hara, Hirokatsu Kataoka, and Yutaka Satoh · 2018
Closest in time.
Grounding visual explanations
Lisa Anne Hendricks, Ronghang Hu, Trevor Darrell, and Zeynep Akata · 2018
Closest in time.
Learning to write with cooperative discriminators
Ari Holtzman, Jan Buys, Maxwell Forbes, Antoine Bosselut, David Golub, and Yejin Choi · 2018
Closest in time.
End-to-end video captioning with multitask reinforcement learning
Lijun Li and Boqing Gong · 2018
Closest in time.
Jointly localizing and describing events for dense video captioning
Yehao Li, Ting Yao, Yingwei Pan, Hongyang Chao, and Tao Mei · 2018
Closest in time.
Improved image captioning with adversarial semantic alignment
Igor Melnyk, Tom Sercu, Pierre L Dognin, Jarret Ross, and Youssef Mroueh · 2018
Closest in time.
Learning a text-video embedding from incomplete and heterogeneous data
Antoine Miech, Ivan Laptev, and Josef Sivic · 2018
Closest in time.
Object hallucination in image captioning
Anna Rohrbach, Lisa Anne Hendricks, Kaylee Burns, Trevor Darrell, and Kate Saenko · 2018
Closest in time.
Learning-based composite metrics for improved caption evaluation
Naeha Sharif, Lyndon White, Mohammed Bennamoun, and Syed Afaq Ali Shah · 2018
Closest in time.
Reconstruction network for video captioning
Bairui Wang, Lin Ma, Wei Zhang, and Wei Liu · 2018
Closest in time.
Show, reward and tell: Automatic generation of narrative paragraph from photo stream by adversarial training
Jing Wang, Jianlong Fu, Jinhui Tang, Zechao Li, and Tao Mei · 2018
Closest in time.
Bidirectional attentive fusion with context gating for dense video captioning
Jingwen Wang, Wenhao Jiang, Lin Ma, Wei Liu, and Yong Xu · 2018
Closest in time.
Video-to-video synthesis
Ting-Chun Wang, Ming-Yu Liu, Jun-Yan Zhu, Guilin Liu, Andrew Tao, Jan Kautz, and Bryan Catanzaro · 2018
Closest in time.
No metrics are perfect: Adversarial reward learning for visual storytelling
Xin Wang, Wenhu Chen, Yuan-Fang Wang, and William Yang Wang · 2018
Closest in time.
Video captioning via hierarchical reinforcement learning
Xin Wang, Wenhu Chen, Jiawei Wu, Yuan-Fang Wang, and William Yang Wang · 2018
Closest in time.
Move forward and tell: A progressive generator of video descriptions
Yilei Xiong, Bo Dai, and Dahua Lin · 2018
Closest in time.
Fine-grained video captioning for sports narrative
Huanyu Yu, Shuo Cheng, Bingbing Ni, Minsi Wang, Jian Zhang, and Xiaokang Yang · 2018
Closest in time.
Towards automatic learning of procedures from web instructional videos
Luowei Zhou, Chenliang Xu, and Jason J Corso · 2018
Closest in time.
End-to-end dense video captioning with masked transformer
Luowei Zhou, Yingbo Zhou, Jason J Corso, Richard Socher, and Caiming Xiong · 2018
Closest in time.