Fetching the paper…
Reading the bibliography…
Scaling up visual category recognition to large numbers of classes remains challenging.
Distributional structure
Z. Harris · 1954
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Support vector machines for multiple-instance learning
S. Andrews, I. Tsochantaridis, and T. Hofmann · 2002
Earlier work this paper cites.
Learning words from sights and sounds: A computational model
D. K. Roy and A. P. Pentland · 2002
Earlier work this paper cites.
Generating typed dependency parses from phrase structure parses
M.-C. De Marneffe, B. MacCartney, and C. Manning · 2006
Earlier work this paper cites.
Weakly supervised scale-invariant learning of models for visual recognition
R. Fergus, P. Perona, and A. Zisserman · 2007
Earlier work this paper cites.
Learning visual attributes
V. Ferrari and A. Zisserman · 2007
Earlier work this paper cites.
Robust object detection with interleaved categorization and segmentation
B. Leibe, A. Leonardis, and B. Schiele · 2008
Earlier work this paper cites.
Attribute and simile classifiers for face verification
N. Kumar, A. C. Berg, P. N. Belhumeur, and S. K. Nayar · 2009
Earlier work this paper cites.
Joint learning of visual attributes, object classes and visual saliency
G. Wang and D. Forsyth · 2009
Earlier work this paper cites.
Label embedding trees for large multi-class tasks
S. Bengio, J. Weston, and D. Grangier · 2010
Earlier work this paper cites.
Attribute-centric recognition for cross-category generalization
A. Farhadi, I. Endres, and D. Hoiem · 2010
Earlier work this paper cites.
Object detection with discriminatively trained part-based models
P. F. Felzenszwalb, R. B. Girshick, D. McAllester, and D. Ramanan · 2010
Earlier work this paper cites.
A discriminative latent model of object classes and attributes
Y. Wang and G. Mori · 2010
Earlier work this paper cites.
Caltech-UCSD Birds 200
P. Welinder, S. Branson, T. Mita, C. Wah, F. Schroff, S. Belongie, and P. Perona · 2010
Earlier work this paper cites.
Large scale image annotation: Learning to rank with joint word-image embeddings
J. Weston, S. Bengio, and N. Usunier · 2010
Earlier work this paper cites.
Attribute-based transfer learning for object categorization with zero or one training example
X. Yu and Y. Aloimonos · 2010
Earlier work this paper cites.
Relative attributes
D. Parikh and K. Grauman · 2011
Earlier work this paper cites.
Evaluating knowledge transfer and zero-shot learning in a large-scale setting
M. Rohrbach, M. Stark, and B.Schiele · 2011
Earlier work this paper cites.
Wsabie: Scaling up to large vocabulary image annotation
J. Weston, S. Bengio, and N. Usunier · 2011
Cited alongside, same era.
Detecting actions, poses, and objects with relational phraselets
C. Desai and D. Ramanan · 2012
Cited alongside, same era.
Discovering localized attributes for fine-grained recognition
K. Duan, D. Parikh, D. J. Crandall, and K. Grauman · 2012
Cited alongside, same era.
Online incremental attribute-based zero-shot learning
P. Kankuekul, A. Kawewong, S. Tangruamsub, and O. Hasegawa · 2012
Cited alongside, same era.
Face detection, pose estimation, and landmark localization in the wild
X. Zhu and D. Ramanan · 2012
Cited alongside, same era.
Devise: A deep visual-semantic embedding model
A. Frome, G. S. Corrado, J. Shlens, S. Bengio, J. Dean, and T. Mikolov · 2013
Cited alongside, same era.
Beyond pascal: A benchmark for 3D object detection in the wild
Y. Xiang, R. Mottaghi, and S. Savarese · 2014
Later among the works it cites.
Part-based R-CNNs for fine-grained category detection
N. Zhang, J. Donahue, R. Girshick, and T. Darrell · 2014
Later among the works it cites.
Panda: Pose aligned networks for deep attribute modeling
N. Zhang, M. Paluri, M. Ranzato, T. Darrell, and L. Bourdev · 2014
Later among the works it cites.
Learning deep features for scene recognition using places database
B. Zhou, A. Lapedriza, J. Xiao, A. Torralba, and A. Oliva · 2014
Later among the works it cites.
Label-embedding for image classification
Z. Akata, F. Perronnin, Z. Harchaoui, and C. Schmid · 2015
Later among the works it cites.
Evaluation of Output Embeddings for Fine-Grained Image Classification
Z. Akata, S. Reed, D. Walter, H. Lee, and B. Schiele · 2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Attribute-based classification for zero-shot visual object categorization
C. Lampert, H. Nickisch, and S. Harmeling · 2013
Cited alongside, same era.
Efficient estimation of word representations in vector space
T. Mikolov, K. Chen, G. Corrado, and J. Dean · 2013
Cited alongside, same era.
Distributed representations of words and phrases and their compositionality
T. Mikolov, I. Sutskever, K. Chen, G. S. Corrado, and J. Dean · 2013
Cited alongside, same era.
Linguistic regularities in continuous space word representations
T. Mikolov, W.-t. Yih, and G. Zweig · 2013
Cited alongside, same era.
Zero-shot learning by convex combination of semantic embeddings
M. Norouzi, T. Mikolov, S. Bengio, Y. Singer, J. Shlens, A. Frome, G. Corrado, and J. Dean · 2013
Cited alongside, same era.
Articulated human detection with flexible mixtures of parts
Y. Yang and D. Ramanan · 2013
Cited alongside, same era.
Predicting deep zero-shot convolutional neural networks using textual descriptions
J. Ba, K. Swersky, S. Fidler, and R. Salakhutdinov · 2015
Later among the works it cites.
P-cnn: Pose-based cnn features for action recognition
G. Cheron, I. Laptev, and C. Schmid · 2015
Later among the works it cites.
Long-term recurrent convolutional networks for visual recognition and description
J. Donahue, L. A. Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell · 2015
Later among the works it cites.
Are you talking to a machine? dataset and methods for multilingual image question answering
H. Gao, J. Mao, J. Zhou, Z. Huang, L. Wang, and W. Xu · 2015
Later among the works it cites.
Deep visual-semantic alignments for generating image descriptions
A. Karpathy and F. Li · 2015
Later among the works it cites.
Fisher vectors derived from hybrid gaussian-laplacian mixture models for image annotation
B. Klein, G. Lev, G. Sadeh, and L. Wolf · 2015
Later among the works it cites.
Ask your neurons: A neural-based approach to answering questions about images
M. Malinowski, M. Rohrbach, and M. Fritz · 2015
Later among the works it cites.
Deep captioning with multimodal recurrent neural networks (m-rnn)
J. Mao, W. Xu, Y. Yang, J. Wang, Z. Huang, and A. L. Yuille · 2015
Later among the works it cites.
Person recognition in personal photo collections
S. Oh, R. Benenson, M. Fritz, and B. Shiele · 2015
Later among the works it cites.
Image question answering: A visual semantic embedding model and a new dataset
M. Ren, R. Kiros, and R. Zemel · 2015
Later among the works it cites.
Show and tell: A neural image caption generator
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan · 2015
Later among the works it cites.
Object detectors emerge in deep scene cnns
B. Zhou, A. Khosla, À. Lapedriza, A. Oliva, and A. Torralba · 2015
Later among the works it cites.