Fetching the paper…
Reading the bibliography…
Video captioning combines video understanding and language generation.
Constant-time machine translation with conditional masked language models
Marjan Ghazvininejad, Omer Levy, Yinhan Liu, and Luke Zettlemoyer. 2019 · 1904
Earlier work this paper cites.
The attention system of the human brain
Michael I Posner and Steven E Petersen. 1990 · 1990
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Control of goal-directed and stimulus-driven attention in the brain
Maurizio Corbetta and Gordon L Shulman. 2002 · 2002
Earlier work this paper cites.
BLEU: a Method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Automatic evaluation of summaries using n-gram co-occurrence statistics
Chin-Yew Lin and Eduard H. Hovy. 2003 · 2003
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
METEOR: An automatic metric for MT evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie. 2005 · 2005
Earlier work this paper cites.
Object-based auditory and visual attention
Barbara G Shinn-Cunningham. 2008 · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Fei-Fei Li. 2009 · 2009
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton. 2012 · 2012
Earlier work this paper cites.
Youtube2text: Recognizing and describing arbitrary activities using semantic hierarchies and zero-shot recognition
Sergio Guadarrama, Niveda Krishnamoorthy, Girish Malkarnenkar, Subhashini Venugopalan, Raymond J. Mooney, Trevor Darrell, and Kate Saenko. 2013 · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba. 2014 · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
Microsoft COCO captions: Data collection and evaluation server
Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Dollár, and C. Lawrence Zitnick. 2015 · 2015
Earlier work this paper cites.
Learning spatiotemporal features with 3d convolutional networks
Du Tran, Lubomir D. Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri. 2015 · 2015
Earlier work this paper cites.
Cider: Consensus-based image description evaluation
Ramakrishna Vedantam, C. Lawrence Zitnick, and Devi Parikh. 2015 · 2015
Earlier work this paper cites.
Translating videos to natural language using deep recurrent neural networks
Subhashini Venugopalan, Huijuan Xu, Jeff Donahue, Marcus Rohrbach, Raymond J. Mooney, and Kate Saenko. 2015 · 2015
Earlier work this paper cites.
Show and tell: A neural image caption generator
Oriol Vinyals, Alexander Toshev, Samy Bengio, and Dumitru Erhan. 2015 · 2015
Earlier work this paper cites.
Show, attend and tell: Neural image caption generation with visual attention
Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhudinov, Rich Zemel, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
Describing videos by exploiting temporal structure
Li Yao, Atousa Torabi, Kyunghyun Cho, Nicolas Ballas, Christopher J. Pal, Hugo Larochelle, and Aaron C. Courville. 2015 · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Cited alongside, same era.
Sequence-level knowledge distillation
Yoon Kim and Alexander M. Rush. 2016 · 2016
Cited alongside, same era.
How blind people interact with visual content on social networking services
Violeta Voykinska, Shiri Azenkot, Shaomei Wu, and Gilly Leshed. 2016 · 2016
Cited alongside, same era.
What value do explicit high level concepts have in vision to language problems?
Qi Wu, Chunhua Shen, Lingqiao Liu, Anthony R. Dick, and Anton van den Hengel. 2016 · 2016
Cited alongside, same era.
MSR-VTT: A large video description dataset for bridging video and language
Jun Xu, Tao Mei, Ting Yao, and Yong Rui. 2016 · 2016
Cited alongside, same era.
Visual dialog
Abhishek Das, Satwik Kottur, Khushi Gupta, Avi Singh, Deshraj Yadav, José M. F. Moura, Devi Parikh, and Dhruv Batra. 2017 · 2017
Non-autoregressive neural machine translation with enhanced decoder input
Junliang Guo, Xu Tan, Di He, Tao Qin, Linli Xu, and Tie-Yan Liu. 2019 · 2019
Later among the works it cites.
Memory-attended recurrent network for video captioning
Wenjie Pei, Jiyuan Zhang, Xiangrong Wang, Lei Ke, Xiaoyong Shen, and Yu-Wing Tai. 2019 · 2019
Later among the works it cites.
Retrieving sequential information for non-autoregressive neural machine translation
Chenze Shao, Yang Feng, Jinchao Zhang, Fandong Meng, Xilin Chen, and Jie Zhou. 2019 · 2019
Later among the works it cites.
Object-aware aggregation with bidirectional temporal graph for video captioning
Junchao Zhang and Yuxin Peng. 2019 · 2019
Later among the works it cites.
Intention oriented image captions with guiding objects
Yue Zheng, Yali Li, and Shengjin Wang. 2019 · 2019
Later among the works it cites.
Video description: A survey of methods, datasets, and evaluation metrics
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
The kinetics human action video dataset
Will Kay, João Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, Mustafa Suleyman, and Andrew Zisserman. 2017 · 2017
Cited alongside, same era.
Knowing when to look: Adaptive attention via a visual sentinel for image captioning
Jiasen Lu, Caiming Xiong, Devi Parikh, and Richard Socher. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Learning multimodal attention LSTM networks for video captioning
Jun Xu, Ting Yao, Yongdong Zhang, and Tao Mei. 2017 · 2017
Cited alongside, same era.
Bottom-up and top-down attention for image captioning and VQA
Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang. 2018 · 2018
Cited alongside, same era.
Less is more: Picking informative frames for video captioning
Yangyu Chen, Shuhui Wang, Weigang Zhang, and Qingming Huang. 2018 · 2018
Cited alongside, same era.
Nayyer Aafaq, Ajmal Mian, Wei Liu, Syed Zulqarnain Gilani, and Mubarak Shah. 2020 · 2020
Later among the works it cites.
Say as you wish: Fine-grained control of image caption generation with abstract scene graphs
Shizhe Chen, Qin Jin, Peng Wang, and Qi Wu. 2020 · 2020
Later among the works it cites.
Aligned cross entropy for non-autoregressive machine translation
Marjan Ghazvininejad, Vladimir Karpukhin, Luke Zettlemoyer, and Omer Levy. 2020 · 2020
Later among the works it cites.
Non-autoregressive machine translation with disentangled context transformer
Jungo Kasai, James Cross, Marjan Ghazvininejad, and Jiatao Gu. 2020 · 2020
Later among the works it cites.
Spatio-temporal graph for video captioning with knowledge distillation
Boxiao Pan, Haoye Cai, De-An Huang, Kuan-Hui Lee, Adrien Gaidon, Ehsan Adeli, and Juan Carlos Niebles. 2020 · 2020
Later among the works it cites.
A study of non-autoregressive model for sequence generation
Yi Ren, Jinglin Liu, Xu Tan, Zhou Zhao, Sheng Zhao, and Tie-Yan Liu. 2020 · 2020
Later among the works it cites.
STAT: spatial-temporal attention mechanism for video captioning
Chenggang Yan, Yunbin Tu, Xingzheng Wang, Yongbing Zhang, Xinhong Hao, Yongdong Zhang, and Qionghai Dai. 2020 · 2020
Later among the works it cites.
Controllable video captioning with an exemplar sentence
Yitian Yuan, Lin Ma, Jingwen Wang, and Wenwu Zhu. 2020 · 2020
Later among the works it cites.
Object relational graph with teacher-recommended learning for video captioning
Ziqi Zhang, Yaya Shi, Chunfeng Yuan, Bing Li, Peijin Wang, Weiming Hu, and Zheng-Jun Zha. 2020 · 2020
Later among the works it cites.
Syntax-aware action targeting for video captioning
Qi Zheng, Chaoyue Wang, and Dacheng Tao. 2020 · 2020
Later among the works it cites.
Multi-task learning with shared encoder for non-autoregressive machine translation
Yongchang Hao, Shilin He, Wenxiang Jiao, Zhaopeng Tu, Michael R. Lyu, and Xing Wang. 2021 · 2021
Closest in time.
Can latent alignments improve autoregressive machine translation?
Adi Haviv, Lior Vassertail, and Omer Levy. 2021 · 2021
Closest in time.
Improving video captioning with temporal composition of a visual-syntactic embedding
Jesus Perez-Martin, Benjamin Bustos, and Jorge Perez. 2021 · 2021
Closest in time.
Semantic grouping network for video captioning
Hobin Ryu, Sunghun Kang, Haeyong Kang, and Chang D. Yoo. 2021 · 2021
Closest in time.
Non-autoregressive coarse-to-fine video captioning
Bang Yang, Yuexian Zou, Fenglin Liu, and Can Zhang. 2021 · 2021
Closest in time.