Fetching the paper…
Reading the bibliography…
We study scalable and uniform understanding of facts in images.
Zipf, G.K.: The psycho-biology of language. (1935)
1935
Earlier work this paper cites.
Deng, J., Dong, W., Socher, R., Li, L.J., Li, K., Fei-Fei, L.: Imagenet: A large-scale hierarchical image database. In: CVPR. IEEE (2009)
2009
Earlier work this paper cites.
Farhadi, A., Endres, I., Hoiem, D., Forsyth, D.: Describing objects by their attributes. In: CVPR (2009)
2009
Earlier work this paper cites.
Lampert, C.H., Nickisch, H., Harmeling, S.: Learning to detect unseen object classes by between-class attribute transfer. In: CVPR (2009)
2009
Earlier work this paper cites.
Muja, M., Lowe, D.: Flann-fast library for approximate nearest neighbors user manual. Computer Science Department, University of British Columbia, Vancouver, BC, Canada (2009)
2009
Earlier work this paper cites.
Palatucci, M., Pomerleau, D., Hinton, G.E., Mitchell, T.M.: Zero-shot learning with semantic output codes. In: NIPS (2009)
2009
Earlier work this paper cites.
Yao, B., Fei-Fei, L.: Grouplet: A structured image representation for recognizing human and object interactions. In: CVPR (2010)
2010
Earlier work this paper cites.
Sadeghi, M.A., Farhadi, A.: Recognition using visual phrases. In: CVPR (2011)
2011
Earlier work this paper cites.
Salakhutdinov, R., Torralba, A., Tenenbaum, J.: Learning to share visual appearance for multiclass object detection. In: Computer Vision and Pattern Recognition (CVPR), 2011 IEEE Conference on. pp. 1481–1488. IEEE (2011)
2011
Earlier work this paper cites.
Yao, B., Jiang, X., Khosla, A., Lin, A.L., Guibas, L., Fei-Fei, L.: Human action recognition by learning bases of action attributes and parts. In: ICCV (2011)
2011
Earlier work this paper cites.
Everingham, M., Van Gool, L., Williams, C.K.I., Winn, J., Zisserman, A.: The PASCAL Visual Object Classes Challenge 2012 (VOC2012) Results. http://www.pascal-network.org/challenges/VOC/voc2012/workshop/index.html
2012
Earlier work this paper cites.
Krizhevsky, A., Sutskever, I., Hinton, G.E.: Imagenet classification with deep convolutional neural networks. In: NIPS (2012)
2012
Earlier work this paper cites.
Akata, Z., Perronnin, F., Harchaoui, Z., Schmid, C.: Label-embedding for attribute-based classification. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 819–826 (2013)
2013
Earlier work this paper cites.
Elhoseiny, M., Saleh, B., Elgammal, A.: Write a classifier: Zero-shot learning using purely textual descriptions. In: ICCV (2013)
2013
Earlier work this paper cites.
Frome, A., Corrado, G.S., Shlens, J., Bengio, S., Dean, J., Mikolov, T., et al.: Devise: A deep visual-semantic embedding model. In: NIPS (2013)
2013
Earlier work this paper cites.
Mikolov, T., Sutskever, I., Chen, K., Corrado, G.S., Dean, J.: Distributed representations of words and phrases and their compositionality. In: NIPS (2013)
2013
Cited alongside, same era.
Socher, R., Ganjoo, M., Sridhar, H., Bastani, O., Manning, C.D., Ng, A.Y.: Zero shot learning through cross-modal transfer. In: NIPS (2013)
2013
Cited alongside, same era.
Antol, S., Zitnick, C.L., Parikh, D.: Zero-Shot Learning via Visual Abstraction. In: ECCV (2014)
2014
Cited alongside, same era.
Antol, S., Zitnick, C.L., Parikh, D.: Zero-shot learning via visual abstraction. In: ECCV (2014)
2014
Cited alongside, same era.
Chen, C.Y., Grauman, K.: Inferring analogous attributes. In: CVPR (2014)
2014
Cited alongside, same era.
2015
Closest in time.
Gkioxari, G., Malik, J.: Finding action tubes. In: CVPR (2015)
2015
Closest in time.
Gupta, A.: Sports Dataset. http://www.cs.cmu.edu/~abhinavg/Downloads.html (2009), [Online; accessed 15-July-2015]
2015
Closest in time.
Johnson, J., Krishna, R., Stark, M., Li, L.J., Shamma, D., Bernstein, M., Fei-Fei, L.: Image retrieval using scene graphs. In: CVPR (2015)
2015
Closest in time.
Kiros, J.R.: Image-sentence tacl15 implementation. https://github.com/ryankiros/visual-semantic-embedding (2015), [Online; accessed 19-Nov-2015]
2015
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Gong, Y., Ke, Q., Isard, M., Lazebnik, S.: A multi-view embedding space for modeling internet images, tags, and their semantics. International journal of computer vision 106(2), 210–233 (2014)
2014
Cited alongside, same era.
Jia, Y., Shelhamer, E., Donahue, J., Karayev, S., Long, J., Girshick, R., Guadarrama, S., Darrell, T.: Caffe: Convolutional architecture for fast feature embedding. In: Proceedings of the ACM International Conference on Multimedia. pp. 675–678. ACM (2014)
2014
Cited alongside, same era.
Karpathy, A., Joulin, A., Li, F.F.F.: Deep fragment embeddings for bidirectional image sentence mapping. In: Advances in neural information processing systems. pp. 1889–1897 (2014)
2014
Cited alongside, same era.
Norouzi, M., Mikolov, T., Bengio, S., Singer, Y., Shlens, J., Frome, A., Corrado, G.S., Dean, J.: Zero-shot learning by convex combination of semantic embeddings. In: ICLR (2014)
2014
Cited alongside, same era.
Pennington, J., Socher, R., Manning, C.D.: Glove: Global vectors for word representation. EMNLP (2014)
2014
Cited alongside, same era.
Zhang, N., Paluri, M., Ranzato, M., Darrell, T., Bourdev, L.: Panda: Pose aligned networks for deep attribute modeling. In: Computer Vision and Pattern Recognition (CVPR), 2014 IEEE Conference on. pp. 1637–1644. IEEE (2014)
2014
Cited alongside, same era.
Zhou, B., Lapedriza, A., Xiao, J., Torralba, A., Oliva, A.: Learning deep features for scene recognition using places database. In: NIPS (2014)
2014
Cited alongside, same era.
Kiros, R., Salakhutdinov, R., Zemel, R.S.: Unifying visual-semantic embeddings with multimodal neural language models. TACL (2015)
2015
Closest in time.
Malinowski, M., Rohrbach, M., Fritz, M.: Ask your neurons: A neural-based approach to answering questions about images. In: ICCV (2015)
2015
Closest in time.
Mao, J., Xu, W., Yang, Y., Wang, J., Yuille, A.: Deep captioning with multimodal recurrent neural networks (m-rnn). ICLR (2015)
2015
Closest in time.
Ren, M., Kiros, R., Zemel, R.: Exploring models and data for image question answering. In: NIPS (2015)
2015
Closest in time.
Romera-Paredes, B., Torr, P.: An embarrassingly simple approach to zero-shot learning. In: Proceedings of The 32nd International Conference on Machine Learning. pp. 2152–2161 (2015)
2015
Closest in time.
Simonyan, K., Zisserman, A.: Very deep convolutional networks for large-scale image recognition. In: ICLR (2015)
2015
Closest in time.
Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S., Anguelov, D., Erhan, D., Vanhoucke, V., Rabinovich, A.: Going deeper with convolutions (June 2015)
2015
Closest in time.
Vinyals, O., Toshev, A., Bengio, S., Erhan, D.: Show and tell: A neural image caption generator (2015)
2015
Closest in time.
Xu, K., Ba, J., Kiros, R., Courville, A., Salakhutdinov, R., Zemel, R., Bengio, Y.: Show, attend and tell: Neural image caption generation with visual attention. In: ICML (2015)
2015
Closest in time.
Mohamed Elhoseiny, Scott Cohen, W.C.B.P.A.E.: Automatic annotation of structured facts in images. In: Arxiv (2016)
2016
Closest in time.