Fetching the paper…
Reading the bibliography…
State-of-the-art methods for zero-shot visual recognition formulate learning as a joint embedding problem of images and side information.
Distributional structure
Z. Harris · 1954
Earlier work this paper cites.
Wordnet: a lexical database for English
G. A. Miller · 1995
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Automated flower classification over a large number of classes
M.-E. Nilsback and A. Zisserman · 2008
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Zero-shot learning with semantic output codes
M. Palatucci, D. Pomerleau, G. Hinton, and T. Mitchell · 2009
Earlier work this paper cites.
Label embedding trees for large multi-class tasks
S. Bengio, J. Weston, and D. Grangier · 2010
Earlier work this paper cites.
Caltech-UCSD Birds 200
P. Welinder, S. Branson, T. Mita, C. Wah, F. Schroff, S. Belongie, and P. Perona · 2010
Earlier work this paper cites.
Large scale image annotation: Learning to rank with joint word-image embeddings
J. Weston, S. Bengio, and N. Usunier · 2010
Earlier work this paper cites.
Baby talk: understanding and generating simple image descriptions
G. Kulkarni, V. Premraj, S. Dhar, S. Li, Y. choi, A. Berg, and T. Berg · 2011
Earlier work this paper cites.
Multimodal deep learning
J. Ngiam, A. Khosla, M. Kim, J. Nam, H. Lee, and A. Y. Ng · 2011
Earlier work this paper cites.
Im2Text: Describing images using 1 million captioned photographs
V. Ordonez, G. Kulkarni, and T. Berg · 2011
Earlier work this paper cites.
Evaluating knowledge transfer and zero-shot learning in a large-scale setting
M. Rohrbach, M. Stark, and B. Schiele · 2011
Earlier work this paper cites.
Discovering localized attributes for fine-grained recognition
K. Duan, D. Parikh, D. J. Crandall, and K. Grauman · 2012
Earlier work this paper cites.
ImageNet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
Zero-shot video retrieval using content and concepts
J. Dalton, J. Allan, and P. Mirajkar · 2013
Earlier work this paper cites.
Fine-grained crowdsourcing for fine-grained recognition
J. Deng, J. Krause, and L. Fei-Fei · 2013
Earlier work this paper cites.
Write a classifier: Zero-shot learning using purely textual descriptions
M. Elhoseiny, B. Saleh, and A. Elgammal · 2013
Earlier work this paper cites.
Devise: A deep visual-semantic embedding model
A. Frome, G. S. Corrado, J. Shlens, S. Bengio, J. Dean, and T. Mikolov · 2013
Cited alongside, same era.
Distributed representations of words and phrases and their compositionality
T. Mikolov, I. Sutskever, K. Chen, G. S. Corrado, and J. Dean · 2013
Cited alongside, same era.
Zero-shot learning by convex combination of semantic embeddings
M. Norouzi, T. Mikolov, S. Bengio, Y. Singer, J. Shlens, A. Frome, G. Corrado, and J. Dean · 2013
Cited alongside, same era.
Zero-shot learning through cross-modal transfer
R. Socher, M. Ganjoo, H. Sridhar, O. Bastani, C. Manning, and A. Ng · 2013
Cited alongside, same era.
Decaf: A deep convolutional activation feature for generic visual recognition
J. Donahue, Y. Jia, O. Vinyals, J. Hoffman, N. Zhang, E. Tzeng, and T. Darrell · 2014
Cited alongside, same era.
Predicting deep zero-shot convolutional neural networks using textual descriptions
J. Ba, K. Swersky, S. Fidler, and R. Salakhutdinov · 2015
Later among the works it cites.
Long-term recurrent convolutional networks for visual recognition and description
J. Donahue, L. A. Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell · 2015
Later among the works it cites.
Transductive multi-view zero-shot learning
Y. Fu, T. M. Hospedales, T. Xiang, and S. Gong · 2015
Later among the works it cites.
Learning hypergraph-regularized attribute predictors
S. Huang, M. Elhoseiny, A. Elgammal, and D. Yang · 2015
Later among the works it cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
S. Ioffe and C. Szegedy · 2015
Later among the works it cites.
Deep visual-semantic alignments for generating image descriptions
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Fu, T. M. Hospedales, T. Xiang, Z. Fu, and S. Gong · 2014
Cited alongside, same era.
Composite concept discovery for zero-shot video event detection
A. Habibian, T. Mensink, and C. G. Snoek · 2014
Cited alongside, same era.
Attribute-based classification for zero-shot visual object categorization
C. Lampert, H. Nickisch, and S. Harmeling · 2014
Cited alongside, same era.
Microsoft COCO: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Cited alongside, same era.
Glove: Global vectors for word representation
J. Pennington, R. Socher, and C. D. Manning · 2014
Cited alongside, same era.
Improved multimodal deep learning with variation of information
K. Sohn, W. Shang, and H. Lee · 2014
Cited alongside, same era.
Multimodal learning with deep boltzmann machines
N. Srivastava and R. Salakhutdinov · 2014
Cited alongside, same era.
A. Karpathy and F. Li · 2015
Later among the works it cites.
Ranking and retrieval of image sequences from multiple paragraph queries
G. Kim, S. Moon, and L. Sigal · 2015
Later among the works it cites.
Deep captioning with multimodal recurrent neural networks (M-RNN)
J. Mao, W. Xu, Y. Yang, J. Wang, and A. Yuille · 2015
Later among the works it cites.
Facebook’s ai can caption photos for the blind on its own, October 2015
C. Metz · 2015
Later among the works it cites.
Beyond short snippets: Deep networks for video classification
J. Y.-H. Ng, M. Hausknecht, S. Vijayanarasimhan, O. Vinyals, R. Monga, and G. Toderici · 2015
Later among the works it cites.
ImageNet Large Scale Visual Recognition Challenge
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. C. Berg, and L. Fei-Fei · 2015
Later among the works it cites.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2015
Later among the works it cites.
Going deeper with convolutions
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich · 2015
Later among the works it cites.
Show and tell: A neural image caption generator
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan · 2015
Later among the works it cites.
Show, attend and tell: Neural image caption generation with visual attention
K. Xu, J. Ba, R. Kiros, A. Courville, R. Salakhutdinov, R. Zemel, and Y. Bengio · 2015
Later among the works it cites.
Character-level convolutional networks for text classification
X. Zhang, J. Zhao, and Y. LeCun · 2015
Later among the works it cites.
Conditional random fields as recurrent neural networks
S. Zheng, S. Jayasumana, B. Romera-Paredes, V. Vineet, Z. Su, D. Du, C. Huang, and P. H. Torr · 2015
Later among the works it cites.