Fetching the paper…
Reading the bibliography…
People can recognize scenes across many different modalities beyond natural images.
Signature verification using a “siamese” time delay neural network
J. Bromley, J. W. Bentz, L. Bottou, I. Guyon, Y. LeCun, C. Moore, E. Säckinger, and R. Shah · 1993
Earlier work this paper cites.
Canonical correlation analysis: An overview with application to learning methods
D. R. Hardoon, S. Szedmak, and J. Shawe-Taylor · 2004
Earlier work this paper cites.
Contextual models for object detection using boosted random fields
K. Murphy and W. Freeman · 2004
Earlier work this paper cites.
Cross-generalization: Learning novel classes from a single example by feature replacement
E. Bart and S. Ullman · 2005
Earlier work this paper cites.
Object classification from a single example utilizing class relevance metrics
M. Fink · 2005
Earlier work this paper cites.
Geometric context from a single image
D. Hoiem, A. A. Efros, and M. Hebert · 2005
Earlier work this paper cites.
One-shot learning of object categories
L. Fei-Fei, R. Fergus, and P. Perona · 2006
Earlier work this paper cites.
Curriculum learning
Y. Bengio, J. Louradour, R. Collobert, and J. Weston · 2009
Earlier work this paper cites.
Learning to detect unseen object classes by between-class attribute transfer
C. H. Lampert, H. Nickisch, and S. Harmeling · 2009
Earlier work this paper cites.
Zero-shot learning with semantic output codes
M. Palatucci, D. Pomerleau, G. E. Hinton, and T. M. Mitchell · 2009
Earlier work this paper cites.
A new approach to cross-modal multimedia retrieval
N. Rasiwasia, J. Costa Pereira, E. Coviello, G. Doyle, G. R. Lanckriet, R. Levy, and N. Vasconcelos · 2010
Earlier work this paper cites.
Adapting visual category models to new domains
K. Saenko, B. Kulis, M. Fritz, and T. Darrell · 2010
Earlier work this paper cites.
Sun database: Large-scale scene recognition from abbey to zoo
J. Xiao, J. Hays, K. Ehinger, A. Oliva, A. Torralba, et al · 2010
Earlier work this paper cites.
Sketch-based image retrieval: Benchmark and bag-of-features descriptors
M. Eitz, K. Hildebrand, T. Boubekeur, and M. Alexa · 2011
Earlier work this paper cites.
Domain adaptation for object recognition: An unsupervised approach
R. Gopalan, R. Li, and R. Chellappa · 2011
Earlier work this paper cites.
Learning cross-modality similarity for multinomial data
Y. Jia, M. Salzmann, and T. Darrell · 2011
Earlier work this paper cites.
Multimodal deep learning
J. Ngiam, A. Khosla, M. Kim, J. Nam, H. Lee, and A. Y. Ng · 2011
Cited alongside, same era.
Unbiased look at dataset bias
A. Torralba, A. Efros, et al · 2011
Cited alongside, same era.
What makes a good detector?–structured priors for learning from few examples
T. Gao, M. Stark, and D. Koller · 2012
Cited alongside, same era.
Undoing the damage of dataset bias
A. Khosla, T. Zhou, T. Malisiewicz, A. A. Efros, and A. Torralba · 2012
Cited alongside, same era.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Cited alongside, same era.
Decaf: A deep convolutional activation feature for generic visual recognition
J. Donahue, Y. Jia, O. Vinyals, J. Hoffman, N. Zhang, E. Tzeng, and T. Darrell · 2013
Unifying visual-semantic embeddings with multimodal neural language models
R. Kiros, R. Salakhutdinov, and R. S. Zemel · 2014
Later among the works it cites.
Cluster canonical correlation analysis
N. Rasiwasia, D. Mahajan, V. Mahadevan, and G. Aggarwal · 2014
Later among the works it cites.
Object detectors emerge in deep scene cnns
B. Zhou, A. Khosla, A. Lapedriza, A. Oliva, and A. Torralba · 2014
Later among the works it cites.
Learning deep features for scene recognition using places database
B. Zhou, A. Lapedriza, J. Xiao, A. Torralba, and A. Oliva · 2014
Later among the works it cites.
Part level transfer regularization for enhancing exemplar svms
Y. Aytar and A. Zisserman · 2015
Later among the works it cites.
Predicting deep zero-shot convolutional neural networks using textual descriptions
J. Ba, K. Swersky, S. Fidler, and R. Salakhutdinov · 2015
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Write a classifier: Zero-shot learning using purely textual descriptions
M. Elhoseiny, B. Saleh, and A. Elgammal · 2013
Cited alongside, same era.
Devise: A deep visual-semantic embedding model
A. Frome, G. S. Corrado, J. Shlens, S. Bengio, J. Dean, T. Mikolov, et al · 2013
Cited alongside, same era.
Zero-shot learning by convex combination of semantic embeddings
M. Norouzi, T. Mikolov, S. Bengio, Y. Singer, J. Shlens, A. Frome, G. S. Corrado, and J. Dean · 2013
Cited alongside, same era.
Zero-shot learning through cross-modal transfer
R. Socher, M. Ganjoo, C. D. Manning, and A. Ng · 2013
Cited alongside, same era.
Hoggles: Visualizing object detection features
C. Vondrick, A. Khosla, T. Malisiewicz, and A. Torralba · 2013
Cited alongside, same era.
Bringing semantics into focus using visual abstraction
C. L. Zitnick and D. Parikh · 2013
Cited alongside, same era.
Later among the works it cites.
Inverting convolutional networks with convolutional networks
A. Dosovitskiy and T. Brox · 2015
Later among the works it cites.
Skip-thought vectors
R. Kiros, Y. Zhu, R. Salakhutdinov, R. S. Zemel, A. Torralba, R. Urtasun, and S. Fidler · 2015
Later among the works it cites.
Learning transferable features with deep adaptation networks
M. Long and J. Wang · 2015
Later among the works it cites.
A. Owens, P. Isola, J. McDermott, A. Torralba, E. H. Adelson, and W. T. Freeman · 2015
Later among the works it cites.
Multi-label cross-modal retrieval
V. Ranjan, N. Rasiwasia, and C. Jawahar · 2015
Later among the works it cites.
Simultaneous deep transfer across domains and tasks
E. Tzeng, J. Hoffman, T. Darrell, and K. Saenko · 2015
Later among the works it cites.
Learning visual biases from human imagination
C. Vondrick, H. Pirsiavash, A. Oliva, and A. Torralba · 2015
Later among the works it cites.
Sketch-based 3d shape retrieval using convolutional neural networks
F. Wang, L. Kang, and Y. Li · 2015
Later among the works it cites.
Show, attend and tell: Neural image caption generation with visual attention
K. Xu, J. Ba, R. Kiros, A. Courville, R. Salakhutdinov, R. Zemel, and Y. Bengio · 2015
Later among the works it cites.