Fetching the paper…
Reading the bibliography…
Over the last two decades we have witnessed strong progress on modeling visual object classes, scenes and attributes that have significantly contributed to automated image understanding.
A computational analysis of the apprehension of spatial relations
G. D. Logan and D. D. Sadler · 1996
Earlier work this paper cites.
Grounding spatial language in perception: an empirical and computational investigation
T. Regier and L. A. Carlson · 2001
Earlier work this paper cites.
Miles: Multiple-instance learning via embedded instance selection
Y. Chen, J. Bi, and J. Z. Wang · 2006
Earlier work this paper cites.
Generating typed dependency parses from phrase structure parses
M.-C. De Marneffe, B. MacCartney, and C. D. Manning · 2006
Earlier work this paper cites.
Beyond bags of features: Spatial pyramid matching for recognizing natural scene categories
S. Lazebnik, C. Schmid, and J. Ponce · 2006
Earlier work this paper cites.
Spatial reference in linguistic human-robot interaction: Iterative, empirically supported development of a model of projective relations
R. Moratz and T. Tenbrink · 2006
Earlier work this paper cites.
Situated dialogue and spatial organization: What, where… and why
G.-J. M. Kruijff, H. Zender, P. Jensfelt, and H. I. Christensen · 2007
Earlier work this paper cites.
Pascal 2008 results, 2008
M. Everingham, L. Van Gool, C. Williams, J. Winn, and A. Zisserman · 2008
Earlier work this paper cites.
Linear spatial pyramid matching using sparse coding for image classification
J. Yang, K. Yu, Y. Gong, and T. Huang · 2009
Earlier work this paper cites.
Exploiting hierarchical context on a large database of object categories
M. J. Choi, J. J. Lim, A. Torralba, and A. S. Willsky · 2010
Earlier work this paper cites.
Object detection with discriminatively trained part based models
P. F. Felzenszwalb, R. B. Girshick, D. McAllester, and D. Ramanan · 2010
Earlier work this paper cites.
A game-theoretic approach to generating spatial descriptions
D. Golland, P. Liang, and D. Klein · 2010
Cited alongside, same era.
Collecting image annotations using amazon’s mechanical turk
C. Rashtchian, P. Young, M. Hodosh, and J. Hockenmaier · 2010
Cited alongside, same era.
Grounding spatial language for video search
S. Tellex, T. Kollar, G. Shaw, N. Roy, and D. Roy · 2010
Cited alongside, same era.
Deep sparse rectifier networks
X. Glorot, A. Bordes, and Y. Bengio · 2011
Cited alongside, same era.
Image ranking and retrieval based on multi-attribute queries
B. Siddiquie, R. S. Feris, and L. S. Davis · 2011
Cited alongside, same era.
Understanding natural language commands for robotic navigation and mobile manipulation
S. Tellex, T. Kollar, S. Dickerson, M. R. Walter, A. G. Banerjee, S. Teller, and N. Roy · 2011
Cited alongside, same era.
Grounding spatial relations for human-robot interaction
S. Guadarrama, L. Riano, D. Golland, D. Gouhring, Y. Jia, D. Klein, P. Abbeel, and T. Darrell · 2013
Later among the works it cites.
Framing image description as a ranking task: Data, models and evaluation metrics
M. Hodosh, P. Young, and J. Hockenmaier · 2013
Later among the works it cites.
Learning smooth pooling regions for visual recognition
M. Malinowski and M. Fritz · 2013
Later among the works it cites.
Rich feature hierarchies for accurate object detection and semantic segmentation
R. Girshick, J. Donahue, T. Darrell, and J. Malik · 2014
Closest in time.
Deep fragment embeddings for bidirectional image sentence mapping
A. Karpathy, A. Joulin, and L. Fei-Fei · 2014
Closest in time.
What are you talking about? text-to-image coreference
C. Kong, D. Lin, M. Bansal, R. Urtasun, and S. Fidler · 2014
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Improving word representations via global context and multiple word prototypes
E. H. Huang, R. Socher, C. D. Manning, and A. Y. Ng · 2012
Cited alongside, same era.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. Hinton · 2012
Cited alongside, same era.
Image retrieval with structured object queries using latent ranking svm
T. Lan, W. Yang, Y. Wang, and G. Mori · 2012
Cited alongside, same era.
Object-centric spatial pooling for image classification
O. Russakovsky, Y. Lin, K. Yu, and L. Fei-Fei · 2012
Cited alongside, same era.
Devise: A deep visual-semantic embedding model
A. Frome, G. S. Corrado, J. Shlens, S. Bengio, J. Dean, T. Mikolov, et al · 2013
Cited alongside, same era.
Closest in time.
Visual semantic search: Retrieving videos via complex textual queries
D. Lin, S. Fidler, C. Kong, and R. Urtasun · 2014
Closest in time.
A multi-world approach to question answering about real-world scenes based on uncertain input
M. Malinowski and M. Fritz · 2014
Closest in time.
Explain images with multimodal recurrent neural networks
J. Mao, W. Xu, Y. Yang, J. Wang, and A. L. Yuille · 2014
Closest in time.
Grounded compositional semantics for finding and describing images with sentences
R. Socher, A. Karpathy, Q. Le, C. Manning, and A. Ng · 2014
Closest in time.