Fetching the paper…
Reading the bibliography…
We describe a novel cross-modal embedding space for actions, named Action2Vec, which combines linguistic cues from class labels with spatio-temporal features derived from video clips.
Rank correlation methods
M. G. Kendall · 1955
Earlier work this paper cites.
Verbs semantics and lexical selection
Z. Wu and M. Palmer · 1994
Earlier work this paper cites.
Wordnet: a lexical database for english
G. A. Miller · 1995
Earlier work this paper cites.
A study on similarity and relatedness using distributional and wordnet-based approaches
E. Agirre, E. Alfonseca, K. Hall, J. Kravalova, M. Paşca, and A. Soroa · 2009
Earlier work this paper cites.
Ucf101: A dataset of 101 human actions classes from videos in the wild
K. Soomro, A. R. Zamir, and M. Shah · 2012
Earlier work this paper cites.
Hmdb51: A large video database for human motion recognition
H. Kuehne, H. Jhuang, R. Stiefelhagen, and T. Serre · 2013
Earlier work this paper cites.
Efficient estimation of word representations in vector space
T. Mikolov, K. Chen, G. Corrado, and J. Dean · 2013
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
T. Mikolov, I. Sutskever, K. Chen, G. S. Corrado, and J. Dean · 2013
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
D. Bahdanau, K. Cho, and Y. Bengio · 2014
Earlier work this paper cites.
Large-scale video classification with convolutional neural networks
A. Karpathy, G. Toderici, S. Shetty, T. Leung, R. Sukthankar, and L. Fei-Fei · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Earlier work this paper cites.
Unifying visual-semantic embeddings with multimodal neural language models
R. Kiros, R. Salakhutdinov, and R. S. Zemel · 2014
Earlier work this paper cites.
Attribute-based classification for zero-shot visual object categorization
C. H. Lampert, H. Nickisch, and S. Harmeling · 2014
Earlier work this paper cites.
Dependency-based word embeddings
O. Levy and Y. Goldberg · 2014
Cited alongside, same era.
Glove: Global vectors for word representation
J. Pennington, R. Socher, and C. Manning · 2014
Cited alongside, same era.
Two-stream convolutional networks for action recognition in videos
K. Simonyan and A. Zisserman · 2014
Cited alongside, same era.
Vqa: Visual question answering
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. Lawrence Zitnick, and D. Parikh · 2015
Cited alongside, same era.
Deep visual-semantic alignments for generating image descriptions
A. Karpathy and L. Fei-Fei · 2015
Cited alongside, same era.
Unsupervised domain adaptation for zero-shot learning
E. Kodirov, T. Xiang, Z. Fu, and S. Gong · 2015
Cited alongside, same era.
Analogy-based detection of morphological and semantic relations with word embeddings: what works and what doesn’t
A. Gladkova, A. Drozd, and S. Matsuoka · 2016
Later among the works it cites.
Hierarchical recurrent neural encoder for video representation with application to captioning
P. Pan, Z. Xu, Y. Yang, F. Wu, and Y. Zhuang · 2016
Later among the works it cites.
Jointly modeling embedding and translation to bridge video and language
Y. Pan, T. Mei, T. Yao, H. Li, and Y. Rui · 2016
Later among the works it cites.
Learning deep structure-preserving image-text embeddings
L. Wang, Y. Li, and S. Lazebnik · 2016
Later among the works it cites.
Multi-task zero-shot action recognition with prioritised data augmentation
X. Xu, T. M. Hospedales, and S. Gong · 2016
Later among the works it cites.
Image captioning with semantic attention
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning spatiotemporal features with 3d convolutional networks
D. Tran, L. Bourdev, R. Fergus, L. Torresani, and M. Paluri · 2015
Cited alongside, same era.
Show and tell: A neural image caption generator
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan · 2015
Cited alongside, same era.
Semantic embedding space for zero-shot action recognition
X. Xu, T. Hospedales, and S. Gong · 2015
Cited alongside, same era.
Word2visualvec: Image and video to sentence matching by visual feature prediction
J. Dong, X. Li, and C. G. Snoek · 2016
Cited alongside, same era.
Convolutional two-stream network fusion for video action recognition
C. Feichtenhofer, A. Pinz, and A. Zisserman · 2016
Cited alongside, same era.
Learning attributes equals multi-source domain generalization
C. Gan · 2016
Cited alongside, same era.
Q. You, H. Jin, Z. Wang, C. Fang, and J. Luo · 2016
Later among the works it cites.
Bottom-up and top-down attention for image captioning and vqa
P. Anderson, X. He, C. Buehler, D. Teney, M. Johnson, S. Gould, and L. Zhang · 2017
Later among the works it cites.
Quo vadis, action recognition? a new model and the kinetics dataset
J. Carreira and A. Zisserman · 2017
Later among the works it cites.
The kinetics human action video dataset
W. Kay, J. Carreira, K. Simonyan, B. Zhang, C. Hillier, S. Vijayanarasimhan, F. Viola, T. Green, T. Back, P. Natsev, et al · 2017
Later among the works it cites.
Learning cross-modal embeddings for cooking recipes and food images
A. Salvador, N. Hynes, Y. Aytar, J. Marin, F. Ofli, I. Weber, and A. Torralba · 2017
Later among the works it cites.
Transductive zero-shot action recognition by word-vector embedding
X. Xu, T. Hospedales, and S. Gong · 2017
Later among the works it cites.
Text2shape: Generating shapes from natural language by learning joint embeddings
K. Chen, C. B. Choy, M. Savva, A. X. Chang, T. Funkhouser, and S. Savarese · 2018
Later among the works it cites.