Fetching the paper…
Reading the bibliography…
Models based on deep convolutional networks have dominated recent image interpretation tasks; we investigate whether models which are also recurrent, or "temporally deep", are effective for tasks involving sequences, visual and otherwise.
D. E. Rumelhart, G. E. Hinton, and R. J. Williams, “Learning internal representations by error propagation,” DTIC Document, Tech. Rep., 1985
1985
Earlier work this paper cites.
R. J. Williams and D. Zipser, “A learning algorithm for continually running fully recurrent neural networks,” in Neural Computation , 1989
1989
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” in Neural Computation . MIT Press, 1997
1997
Earlier work this paper cites.
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu, “BLEU: a method for automatic evaluation of machine translation,” in ACL , 2002
2002
Earlier work this paper cites.
T. Brox, A. Bruhn, N. Papenberg, and J. Weickert, “High accuracy optical flow estimation based on a theory for warping,” in ECCV , 2004
2004
Earlier work this paper cites.
C.-Y. Lin, “Rouge: A package for automatic evaluation of summaries,” in Text Summarization Branches Out: Proceedings of the ACL-04 Workshop , 2004
2004
Earlier work this paper cites.
S. Banerjee and A. Lavie, “METEOR: An automatic metric for MT evaluation with improved correlation with human judgments,” in Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization , 2005
2005
Earlier work this paper cites.
P. Koehn, H. Hoang, A. Birch, C. Callison-Burch, M. Federico, N. Bertoldi, B. Cowan, W. Shen, C. Moran, R. Zens, C. Dyer, O. Bojar, A. Constantin, and E. Herbst, “Moses: Open source toolkit for statistical machine translation,” in ACL , 2007
2007
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “ImageNet: A large-scale hierarchical image database,” in CVPR , 2009
2009
Earlier work this paper cites.
M. Baccouche, F. Mamalet, C. Wolf, C. Garcia, and A. Baskurt, “Action classification in soccer videos with long short-term memory recurrent neural networks,” in International Conference on Artificial Neural Networks (ICANN) , 2010
2010
Earlier work this paper cites.
A. Farhadi, M. Hejrati, M. Sadeghi, P. Young, C. Rashtchian, J. Hockenmaier, and D. Forsyth, “Every picture tells a story: Generating sentences from images,” in ECCV , 2010
2010
Earlier work this paper cites.
M. Baccouche, F. Mamalet, C. Wolf, C. Garcia, and A. Baskurt, “Sequential deep learning for human action recognition,” in Human Behavior Understanding , 2011
2011
Earlier work this paper cites.
I. Sutskever, J. Martens, and G. E. Hinton, “Generating text with recurrent neural networks,” in ICML , 2011
2011
Earlier work this paper cites.
H. Kuehne, H. Jhuang, E. Garrote, T. Poggio, and T. Serre, “HMDB: a large video database for human motion recognition,” in ICCV , 2011
2011
Earlier work this paper cites.
G. Kulkarni, V. Premraj, S. Dhar, S. Li, Y. Choi, A. C. Berg, and T. L. Berg, “Baby talk: Understanding and generating simple image descriptions,” in CVPR , 2011
2011
Earlier work this paper cites.
Y. Yang, C. L. Teo, H. Daumé III, and Y. Aloimonos, “Corpus-guided sentence generation of natural images,” in EMNLP , 2011
2011
Earlier work this paper cites.
M. U. G. Khan, L. Zhang, and Y. Gotoh, “Human focused video description,” in ICCV Workshops , 2011
2011
Earlier work this paper cites.
C. C. Tan, Y.-G. Jiang, and C.-W. Ngo, “Towards textually describing complex video contents with audio-visual concept classifiers,” in ACM MM , 2011
2011
Earlier work this paper cites.
O. Vinyals, S. V. Ravuri, and D. Povey, “Revisiting recurrent neural networks for robust ASR,” in ICASSP , 2012
2012
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “ImageNet classification with deep convolutional neural networks,” in NIPS , 2012
2012
Earlier work this paper cites.
K. Soomro, A. R. Zamir, and M. Shah, “UCF101: A dataset of 101 human actions classes from videos in the wild,” CRCV-TR-12-01, Tech. Rep., 2012
2012
Earlier work this paper cites.
M. Mitchell, X. Han, J. Dodge, A. Mensch, A. Goyal, A. Berg, K. Yamaguchi, T. Berg, K. Stratos, and H. Daumé III, “Midge: Generating image descriptions from computer vision detections,” in Proceedings of the 13th Conference of the European Chapter of the Association for Computational Linguistics , 2012
2012
Earlier work this paper cites.
P. Kuznetsova, V. Ordonez, A. C. Berg, T. L. Berg, and Y. Choi, “Collective generation of natural image descriptions,” in ACL , 2012
2012
Earlier work this paper cites.
A. Barbu, A. Bridge, Z. Burchill, D. Coroian, S. Dickinson, S. Fidler, A. Michaux, S. Mussman, S. Narayanaswamy, D. Salvi, L. Schmidt, J. Shangguan, J. M. Siskind, J. Waggoner, S. Wang, J. Wei, Y. Yin, and Z. Zhang, “Video in sentences out,” in The Conference on Uncertainty in Artificial Intelligence (UAI) , 2012
2012
Earlier work this paper cites.
S. Ji, W. Xu, M. Yang, and K. Yu, “3D convolutional neural networks for human action recognition,” in IEEE Trans. Pattern Anal. Mach. Intell. , 2013
2013
Earlier work this paper cites.
M. Rohrbach, W. Qiu, I. Titov, S. Thater, M. Pinkal, and B. Schiele, “Translating video content to natural language descriptions,” in ICCV , 2013
2013
Cited alongside, same era.
2013
Cited alongside, same era.
P. Y. Micah Hodosh and J. Hockenmaier, “Framing image description as a ranking task: Data, models and evaluation metrics,” in JAIR , vol. 47, 2013
2013
Cited alongside, same era.
A. Frome, G. S. Corrado, J. Shlens, S. Bengio, J. Dean, T. Mikolov et al. , “Devise: A deep visual-semantic embedding model,” in NIPS , 2013
2013
Cited alongside, same era.
H. Wang, A. Kläser, C. Schmid, and C. Liu, “Dense trajectories and motion boundary descriptors for action recognition,” in IJCV , 2013
2013
J. Thomason, S. Venugopalan, S. Guadarrama, K. Saenko, and R. J. Mooney, “Integrating language and vision to generate natural language descriptions of videos in the wild,” in International Conference on Computational Linguistics (COLING) , 2014
2014
Closest in time.
H. Sak, O. Vinyals, G. Heigold, A. Senior, E. McDermott, R. Monga, and M. Mao, “Sequence discriminative distributed training of long short-term memory recurrent neural networks,” in Interspeech , 2014
2014
Closest in time.
K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,” in ICLR , 2015
2015
Closest in time.
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich, “Going deeper with convolutions,” in CVPR , 2015
2015
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
H. Wang and C. Schmid, “Action recognition with improved trajectories,” in ICCV , 2013
2013
Cited alongside, same era.
S. Guadarrama, N. Krishnamoorthy, G. Malkarnenkar, S. Venugopalan, R. Mooney, T. Darrell, and K. Saenko, “YouTube2Text: Recognizing and describing arbitrary activities using semantic hierarchies and zero-shoot recognition,” in ICCV , 2013
2013
Cited alongside, same era.
P. Das, C. Xu, R. Doell, and J. Corso, “Thousand frames in just a few words: Lingual description of videos through latent topics and sparse object stitching,” in CVPR , 2013
2013
Cited alongside, same era.
A. Karpathy, G. Toderici, S. Shetty, T. Leung, R. Sukthankar, and L. Fei-Fei, “Large-scale video classification with convolutional neural networks,” in CVPR , 2014
2014
Cited alongside, same era.
K. Simonyan and A. Zisserman, “Two-stream convolutional networks for action recognition in videos,” in NIPS , 2014
2014
Cited alongside, same era.
A. Graves and N. Jaitly, “Towards end-to-end speech recognition with recurrent neural networks,” in ICML , 2014
2014
Cited alongside, same era.
I. Sutskever, O. Vinyals, and Q. V. Le, “Sequence to sequence learning with neural networks,” in NIPS , 2014
2014
Cited alongside, same era.
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. C. Berg, and L. Fei-Fei, “ImageNet Large Scale Visual Recognition Challenge,” in IJCV , vol. 115, no. 3, 2015
2015
Closest in time.
J. Mao, W. Xu, Y. Yang, J. Wang, and A. Yuille, “Deep captioning with multimodal recurrent neural networks (m-RNN),” in ICLR , 2015
2015
Closest in time.
R. Kiros, R. Salakhuditnov, and R. S. Zemel, “Unifying visual-semantic embeddings with multimodal neural language models,” in TACL , 2015
2015
Closest in time.
R. Vedantam, C. L. Zitnick, and D. Parikh, “CIDEr: Consensus-based image description evaluation,” in CVPR , 2015
2015
Closest in time.
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan, “Show and tell: A neural image caption generator,” in CVPR , 2015
2015
Closest in time.
J. Devlin, H. Cheng, H. Fang, S. Gupta, L. Deng, X. He, G. Zweig, and M. Mitchell, “Language models for image captioning: The quirks and what works,” in ACL , 2015
2015
Closest in time.
J. Mao, W. Xu, Y. Yang, J. Wang, Z. Huang, and A. Yuille, “Learning like a child: Fast novel visual concept learning from sentence descriptions of images,” in ICCV , 2015
2015
Closest in time.
H. Fang, S. Gupta, F. Iandola, R. Srivastava, L. Deng, P. Dollár, J. Gao, X. He, M. Mitchell, J. Platt et al. , “From captions to visual concepts and back,” in CVPR , 2015
2015
Closest in time.
2015
Closest in time.
J. Donahue, L. A. Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell, “Long-term recurrent convolutional networks for visual recognition and description,” in CVPR , 2015
2015
Closest in time.
K. Xu, J. Ba, R. Kiros, A. Courville, R. Salakhutdinov, R. Zemel, and Y. Bengio, “Show, attend and tell: Neural image caption generation with visual attention,” in ICML , 2015
2015
Closest in time.
A. Karpathy and L. Fei-Fei, “Deep visual-semantic alignments for generating image descriptions,” in CVPR , 2015
2015
Closest in time.
J. Y.-H. Ng, M. Hausknecht, S. Vijayanarasimhan, O. Vinyals, R. Monga, and G. Toderici, “Beyond short snippets: Deep networks for video classification,” in CVPR , 2015
2015
Closest in time.
2015
Closest in time.
S. Venugopalan, H. Xu, J. Donahue, M. Rohrbach, R. Mooney, and K. Saenko, “Translating videos to natural language using deep recurrent neural networks,” in NAACL , 2015
2015
Closest in time.
S. Venugopalan, M. Rohrbach, J. Donahue, R. Mooney, T. Darrell, and K. Saenko, “Sequence to sequence–video to text,” in ICCV , 2015
2015
Closest in time.
L. Yao, A. Torabi, K. Cho, N. Ballas, C. Pal, H. Larochelle, and A. Courville, “Describing videos by exploiting temporal structure,” in CVPR , vol. 1050, 2015
2015
Closest in time.
2015
Closest in time.
L. A. Hendricks, S. Venugopalan, M. Rohrbach, R. Mooney, K. Saenko, and T. Darrell, “Deep compositional captioning: Describing novel object categories without paired training data,” in CVPR , 2016
2016
Closest in time.