Fetching the paper…
Reading the bibliography…
We propose a high-level concept word detector that can be integrated with any video-to-language models.
Long Short-Term Memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Bidirectional Recurrent Neural Networks
M. Schuster and K. K. Paliwal · 1997
Earlier work this paper cites.
BLEU: A Method for Automatic Evaluation of Machine Translation
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu · 2002
Earlier work this paper cites.
ROUGE: A Package for Automatic Evaluation of Summaries
C.-Y. Lin · 2004
Earlier work this paper cites.
METEOR: An Automatic Metric for MT Evaluation with Improved Correlation with Human Judgments
S. Banerjee and A. Lavie · 2005
Earlier work this paper cites.
Natural Language Processing with Python
S. Bird, E. Loper, and E. Klein · 2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
X. Glorot and Y. Bengio · 2010
Earlier work this paper cites.
Collecting Highly Parallel Data for Paraphrase Evaluation
D. L. Chen and W. B. Dolan · 2011
Earlier work this paper cites.
A Thousand Frames in Just a Few Words: Lingual Description of Videos through Latent Topics and Sparse Object Stitching
P. Das, C. Xu, R. F. Doell, and J. J. Corso · 2013
Earlier work this paper cites.
Maxout networks
I. J. Goodfellow, D. Warde-farley, M. Mirza, A. Courville, and Y. Bengio · 2013
Earlier work this paper cites.
YouTube2Text: Recognizing and Describing Arbitrary Activities Using Semantic Hierarchies and Zero-shot Recognition
S. Guadarrama, N. Krishnamoorthy, G. Malkarnenkar, S. Venugopalan, R. Mooney, T. Darrell, and K. Saenko · 2013
Earlier work this paper cites.
Distributed Representations of Words and Phrases and their Compositionality
T. Mikolov, I. Sutskever, K. Chen, G. Corrado, and J. Dean · 2013
Earlier work this paper cites.
Translating Video Content to Natural Language Descriptions
M. Rohrbach, W. Qiu, I. Titov, S. Thater, M. Pinkal, and B. Schiele · 2013
Earlier work this paper cites.
Coherent Multi-Sentence Video Description with Variable Level of Detail
A. Rohrbach, M. Rohrbach, W. Qiu, A. Friedrich, M. Pinkal, and B. Schiele · 2014
Cited alongside, same era.
Dropout: A Simple Way to Prevent Neural Networks from Overfitting
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov · 2014
Cited alongside, same era.
Long-term Recurrent Convolutional Networks for Visual Recognition and Description
J. Donahue, L. A. Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell · 2015
Cited alongside, same era.
From Captions to Visual Concepts and Back
H. Fang, S. Gupta, F. Iandola, R. K. Srivastava, L. Deng, P. Dollar, J. Gao, X. He, M. Mitchell, J. C. Platt, C. Lawrence Zitnick, and G. Zweig · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
D. Kingma and J. Ba · 2015
Cited alongside, same era.
The Long-Short Story of Movie Description
A. Rohrbach, M. Rohrbach, and B. Schiele · 2015
J. L. Ba, J. R. Kiros, and G. E. Hinton · 2016
Closest in time.
Multimodal Compact Bilinear Pooling for Visual Question Answering and Visual Grounding
A. Fukui, D. Park, A. Rohrbach, T. Darrel, and M. Rohrbach · 2016
Closest in time.
Deep Residual Learning for Image Recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Closest in time.
Temporal tessellation for video annotation and summarization
D. Kaufman, G. Levi, T. Hassner, and L. Wolf · 2016
Closest in time.
Video fill in the blank with merging lstms
A. Mazaheri, D. Zhang, and M. Shah · 2016
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
ImageNet Large Scale Visual Recognition Challenge
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. C. Berg, and L. Fei-Fei · 2015
Cited alongside, same era.
CIDEr: Consensus-based Image Description Evaluation
R. Vedantam, C. L. Zitnick, and D. Parikh · 2015
Cited alongside, same era.
Sequence to Sequence - Video to Text
S. Venugopalan, R. Marcus, D. Jeffrey, M. Raymond, D. Trevor, and S. Kate · 2015
Cited alongside, same era.
Translating Videos to Natural Language Using Deep Recurrent Neural Networks
S. Venugopalan, H. Xu, J. Donahue, M. Rohrbach, R. Mooney, and K. Saenko · 2015
Cited alongside, same era.
Show, Attend and Tell: Neural Image Caption Generation with Visual Attention
K. Xu, J. Ba, R. Kiros, A. Courville, R. Salakhutdinov, R. Zemel, and Y. Bengio · 2015
Cited alongside, same era.
Jointly Modeling Deep Video and Compositional Text to Bridge Vision and Language in a Unified Framework
R. Xu, C. Xiong, W. Chen, and J. J. Corso · 2015
Cited alongside, same era.
MovieQA: Understanding Stories in Movies through Question-Answering
M. Tapaswi, Y. Zhu, R. Stiefelhagen, A. Torralba, R. Urtasun, and S. Fidler · 2016
Closest in time.
Learning Language-Visual Embedding for Movie Understanding with Natural-Language
A. Torabi, N. Tandon, and L. Sigal · 2016
Closest in time.
Captioning Images with Diverse Objects
S. Venugopalan, L. A. Hendricks, M. Rohrbach, R. Mooney, T. Darrell, and K. Saenko · 2016
Closest in time.
Q. Wu, C. Shen, A. v. d. Hengel, P. Wang, and A. Dick · 2016
Closest in time.
What value do explicit high level concepts have in vision to language problems?
Q. Wu, C. Shen, L. Liu, A. Dick, and A. van den Hengel · 2016
Closest in time.
Image captioning with semantic attention
Q. You, H. Jin, Z. Wang, C. Fang, and J. Luo · 2016
Closest in time.
Video Paragraph Captioning Using Hierarchical Recurrent Neural Networks
H. Yu, J. Wang, Z. Huang, Y. Yang, and W. Xu · 2016
Closest in time.
Movie Description
A. Rohrbach, A. Torabi, M. Rohrbach, N. Tandon, C. Pal, H. Larochelle, A. Courville, and B. Schiele · 2017
Closest in time.