Fetching the paper…
Reading the bibliography…
This paper strives to find the sentence best describing the content of an image or video.
A discriminative kernel-based approach to rank images from text queries
D. Grangier and S. Bengio · 2008
Earlier work this paper cites.
Polynomial semantic indexing
B. Bai, J. Weston, D. Grangier, R. Collobert, K. Sadamasa, Y. Qi, C. Cortes, and M. Mohri · 2009
Earlier work this paper cites.
Collecting highly parallel data for paraphrase evaluation
D. L. Chen and W. B. Dolan · 2011
Earlier work this paper cites.
ImageNet classification using deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. Hinton · 2012
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
T. Tieleman and G. Hinton · 2012
Earlier work this paper cites.
Recent developments in openSMILE, the Munich open-source multimedia feature extractor
F. Eyben, F. Weninger, F. Gross, and B. Schuller · 2013
Earlier work this paper cites.
DeViSE: A deep visual-semantic embedding model
A. Frome, G. Corrado, J. Shlens, S. Bengio, J. Dean, M. Ranzato, and T. Mikolov · 2013
Earlier work this paper cites.
Framing image description as a ranking task: Data, models and evaluation metrics
M. Hodosh, P. Young, and J. Hockenmaier · 2013
Earlier work this paper cites.
Learning deep structured semantic models for web search using clickthrough data
P.-S. Huang, X. He, J. Gao, L. Deng, A. Acero, and L. Heck · 2013
Earlier work this paper cites.
Efficient estimation of word representations in vector space
T. Mikolov, K. Chen, G. Corrado, and J. Dean · 2013
Earlier work this paper cites.
Improving image-sentence embeddings using large weakly annotated photo collections
Y. Gong, L. Wang, M. Hodosh, J. Hockenmaier, and S. Lazebnik · 2014
Earlier work this paper cites.
Caffe: Convolutional architecture for fast feature embedding
Y. Jia, E. Shelhamer, J. Donahue, S. Karayev, J. Long, R. Girshick, S. Guadarrama, and T. Darrell · 2014
Earlier work this paper cites.
CNN features off-the-shelf: An astounding baseline for recognition
A. Razavian, H. Azizpour, J. Sullivan, and S. Carlsson · 2014
Earlier work this paper cites.
Overfeat: Integrated recognition, localization and detection using convolutional networks
P. Sermanet, D. Eigen, X. Zhang, M. Mathieu, R. Fergus, and Y. LeCun · 2014
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
P. Young, A. Lai, M. Hodosh, and J. Hockenmaier · 2014
Earlier work this paper cites.
Predicting deep zero-shot convolutional neural networks using textual descriptions
J. L. Ba, K. Swersky, S. Fidler, and R. Salakhutdinov · 2015
Cited alongside, same era.
From captions to visual concepts and back
H. Fang, S. Gupta, F. Iandola, R. Srivastava, L. Deng, P. Dollar, J. Gao, X. He, M. Mitchell, J. Platt, L. Zitnick, and G. Zweig · 2015
Cited alongside, same era.
Objects2action: Classifying and localizing actions without any video example
M. Jain, J. C. van Gemert, T. Mensink, and C. G. M. Snoek · 2015
Cited alongside, same era.
Deep visual-semantic alignments for generating image descriptions
A. Karpathy and L. Fei-Fei · 2015
Cited alongside, same era.
Fisher vectors derived from hybrid gaussian-laplacian mixture models for image annotation
B. Klein, G. Lev, G. Sadeh, and L. Wolf · 2015
Cited alongside, same era.
Zero-shot image tagging by hierarchical semantic embedding
Show and tell: A neural image caption generator
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan · 2015
Later among the works it cites.
Jointly modeling deep video and compositional text to bridge vision and language in a unified framework
R. Xu, C. Xiong, W. Chen, and J. J. Corso · 2015
Later among the works it cites.
Describing videos by exploiting temporal structure
L. Yao, A. Torabi, K. Cho, N. Ballas, C. Pal, H. Larochelle, and A. Courville · 2015
Later among the works it cites.
EventNet: A large scale structured concept library for complex event detection in video
G. Ye, Y. Li, H. Xu, D. Liu, and S.-F. Chang · 2015
Later among the works it cites.
Trecvid 2016: Evaluating video search, video event detection, localization, and hyperlinking
G. Awad et al · 2016
Closest in time.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
X. Li, S. Liao, W. Lan, X. Du, and G. Yang · 2015
Cited alongside, same era.
Multimodal convolutional neural networks for matching image and sentence
L. Ma, Z. Lu, L. Shang, and H. Li · 2015
Cited alongside, same era.
Deep captioning with multimodal recurrent neural networks (m-RNN)
J. Mao, W. Xu, Y. Yang, J. Wang, Z. Huang, and A. Yuille · 2015
Cited alongside, same era.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
B. Plummer, L. Wang, C. Cervantes, J. Caicedo, J. Hockenmaier, and S. Lazebnik · 2015
Cited alongside, same era.
Imagenet large scale visual recognition challenge
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. Berg, and L. Fei-Fei · 2015
Cited alongside, same era.
Very Deep Convolutional Networks for Large-Scale Image Recognition
K. Simonyan and A. Zisserman · 2015
Cited alongside, same era.
Going deeper with convolutions
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich · 2015
Cited alongside, same era.
Closest in time.
Visual word2vec (vis-w2v): Learning visually grounded word embeddings using abstract scenes
S. Kottur, R. Vedantam, J. M. Moura, and D. Parikh · 2016
Closest in time.
The ImageNet shuffle: Reorganized pre-training for video event detection
P. Mettes, D. Koelma, and C. Snoek · 2016
Closest in time.
Learning joint representations of videos and sentences with web image search
M. Otani, Y. Nakashima, E. Rahtu, J. Heikkilä, and N. Yokoya · 2016
Closest in time.
Jointly modeling embedding and translation to bridge video and language
Y. Pan, T. Mei, T. Yao, H. Li, and Y. Rui · 2016
Closest in time.
Image captioning with deep bidirectional LSTMs
C. Wang, H. Yang, C. Bartz, and C. Meinel · 2016
Closest in time.
Deep learning for video classification and captioning
Z. Wu, T. Yao, Y. Fu, and Y.-G. Jiang · 2016
Closest in time.
MSR-VTT: A large video description dataset for bridging video and language
J. Xu, T. Mei, T. Yao, and Y. Rui · 2016
Closest in time.
Video paragraph captioning using hierarchical recurrent neural networks
H. Yu, J. Wang, Z. Huang, Y. Yang, and W. Xu · 2016
Closest in time.