Fetching the paper…
Reading the bibliography…
Building correspondences across different modalities, such as video and language, has recently become critical in many visual recognition applications, such as video captioning.
Paradigms of Artificial Intelligence Programming: Case Studies in Common Lisp
P. Norvig · 1992
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. J. J. Schmidhuber · 1997
Earlier work this paper cites.
Bleu: a 521 method for automatic evaluation of machine translation
K. Papineni, S. Roukos, T. Ward, and W. Zhu · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
C. Lin · 2004
Earlier work this paper cites.
Meteor: An automatic metric for mt evaluation with improved correlation with human judgements
S. Banerjee and A. Lavie · 2005
Earlier work this paper cites.
Collecting highly parallel data for paraphrase evaluation
D. L. Chen and W. B. Dolan · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
Youtube2text: Recognizing and describing arbitrary activities using semantic hierarchies and zero-shot recognition
S. Guadarrama, N. Krishnamoorthy, G. Malkarnenkar, S. Venugopalan, R. Mooney, T. Darrell, and K. Saenko · 2013
Earlier work this paper cites.
Learning phrase representations using rnn encoder-decoder for statistical machine translation
K. Cho, B. van Merrienboer, C. Gulcehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y. Bengio · 2014
Earlier work this paper cites.
Coherent multi-sentence video description with variable level of detail
A. Rohrbach, M. Rohrbach, W. Qiu, A. Friedrich, M. Pinkal, and B. Schiele · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
I. Sutskever, O. Vinyals, and Q. V. Le · 2014
Earlier work this paper cites.
From captions to visual concepts and back
H. Fang, S. Gupta, F. Iandola, R. Srivastava, L. Deng, P.Dollar, J. Gao, X. He, M. Mitchell, J. C. Platt, C. L. Zitnick, and G. Zweig · 2015
Earlier work this paper cites.
Fully convolutional networks for semantic segmentation
Jonathan Long, Evan Shelhamer, and Trevor Darrell · 2015
Cited alongside, same era.
Faster r-cnn: Towards real-time object detection with region proposal networks
S. Ren, K. He, R. Girshick, and J. Sun · 2015
Cited alongside, same era.
Imagenet large scale visual recognition challenge
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Karpathy, M. Bernstein, A. C. Berg, and L. Fei-Fei · 2015
Cited alongside, same era.
Going deeper with convolutions
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich · 2015
Cited alongside, same era.
Cider: consensus- 535 based image description evaluation
R. Vedantam, C. Lawrence, and D. Parikh · 2015
Cited alongside, same era.
Sequence to sequence—video to text
S. Venugopalan, M. Rohrbach, J. Donahue, R. Mooney, T. Darrell, and K. Saenko · 2015
Msr-vtt: A large video description dataset for bridging video and language
J. Xu, T. Mei, T. Yao, and Y. Rui · 2016
Later among the works it cites.
Video paragraph captioning using hierarchical recurrent neural networks
H. Yu, J. Wang, Z. Huang, Y. Yang, and W. Xu · 2016
Later among the works it cites.
Semantic compositional networks for visual captioning
Z. Gan, C. Gan, X. He, Y. Pu, K. Tran, J. Gao, L. Carin, and L. Deng · 2017
Later among the works it cites.
Knowing when to look: Adaptive attention via a visual sentinel for image captioning
J. Lu, C. Xiong, D. Parikh, and R. Socher · 2017
Later among the works it cites.
To create what you tell: Generating videos from captions
Y. Pan, Z. Qiu, T. Yao, H. Li, and T. Mei · 2017
Later among the works it cites.
Video captioning with transferred semantic attributes
Y. Pan, T. Yao, H. Li, and T. Mei · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Translating videos to natural language using deep recurrent neural network
S. Venugopalan, H. Xu, J. Donahue, M. Rohrbach, R. Mooney, and K. Saenko · 2015
Cited alongside, same era.
Show, attend and tell: Neural image caption generation with visual attention
K. Xu, J. Ba, R. Kiros, K. Cho, A. Courville, R. Salakhutdinov, R. Zemel, and Y. Bengio · 2015
Cited alongside, same era.
Jointly modeling deep video and compositional text to bridge vision and language in a unified framework
R. Xu, C. Xiong, W. Chen, and J. J. Corso · 2015
Cited alongside, same era.
Describing videos by exploiting temporal structure
L. Yao, A. Torabi, K. Cho, N. Ballas, C. Pal, H. Larochelle, and A. Courville · 2015
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
Jointly modeling embedding and translation to bridge video and language
Y. Pan, T. Mei, T. Yao, H. Li, , and Y. Rui · 2016
Cited alongside, same era.
Later among the works it cites.
Inception-v4, inception-resnet and the impact of residual connections on learning
C. Szegedy, S. Ioffe, V. Vanhoucke, and A. A. Alemi · 2017
Later among the works it cites.
Learning multimodal attention lstm networks for video captioning
J. Xu, T. Yao, Y. Zhang, and T. Mei · 2017
Later among the works it cites.
Task-driven dynamic fusion: Reducing ambiguity in video description
X. Zhang, K. Gao, Y. Zhang, D. Zhang, J. Li, and Q. Tian · 2017
Later among the works it cites.
Less is more: Picking informative frames for video captioning
Y. Chen, S. Wang, W. Zhang, and Q. Huang · 2018
Later among the works it cites.
Reconstruction network for video captioning
B. Wang, L. Ma, W. Zhang, and W. Liu · 2018
Later among the works it cites.
Deep learning for video classification and captioning
Z. Wu, T. Yao, Y. Fu, and Y. Jiang · 2018
Later among the works it cites.