Fetching the paper…
Reading the bibliography…
In this paper we propose the construction of linguistic descriptions of images.
Cyc: A large-scale investment in knowledge infrastructure
D. B. Lenat · 1995
Earlier work this paper cites.
An information-theoretic definition of similarity
D. Lin · 1998
Earlier work this paper cites.
HAL’s Legacy: 2001’s Computer as Dream and Reality
D. G. Stork · 1998
Earlier work this paper cites.
Object recognition from local scale-invariant features
D. G. Lowe · 1999
Earlier work this paper cites.
Km-the knowledge machine 2.0: Users manual
P. Clark, B. Porter, and B. P. Works · 2004
Earlier work this paper cites.
Histograms of oriented gradients for human detection
N. Dalal and B. Triggs · 2005
Earlier work this paper cites.
On space-time interest points
I. Laptev · 2005
Earlier work this paper cites.
View-invariant modeling and recognition of human actions using grammars
A. S. Ogale, A. Karapurkar, and Y. Aloimonos · 2006
Earlier work this paper cites.
Enabling experts to build knowledge bases from science textbooks
V. K. Chaudhri, B. E. John, S. Mishra, J. Pacheco, B. Porter, and A. Spaulding · 2007
Earlier work this paper cites.
Objects in action: An approach for combining action understanding and object perception
A. Gupta and L. S. Davis · 2007
Earlier work this paper cites.
Conceptnet 3: a flexible, multilingual semantic network for common sense knowledge
C. Havasi, R. Speer, and J. Alonso · 2007
Earlier work this paper cites.
A discriminatively trained, multiscale, deformable part model
P. Felzenszwalb, D. McAllester, and D. Ramanan · 2008
Earlier work this paper cites.
Describing objects by their attributes
A. Farhadi, I. Endres, D. Hoiem, and D. Forsyth · 2009
Earlier work this paper cites.
Simplenlg: A realisation engine for practical applications
A. Gatt and E. Reiter · 2009
Earlier work this paper cites.
Learning to detect unseen object classes by between-class attribute transfer
C. H. Lampert, H. Nickisch, and S. Harmeling · 2009
Earlier work this paper cites.
Activity recognition using the velocity histories of tracked keypoints
R. Messing, C. Pal, and H. Kautz · 2009
Earlier work this paper cites.
Every picture tells a story: Generating sentences from images
A. Farhadi, M. Hejrati, M. A. Sadeghi, P. Young, C. Rashtchian, J. Hockenmaier, and D. Forsyth · 2010
Earlier work this paper cites.
Learning bayesian networks with the bnlearn R package
M. Scutari · 2010
Cited alongside, same era.
I2t: Image parsing to text description
B. Z. Yao, X. Yang, L. Lin, M. W. Lee, and S. C. Zhu · 2010
Cited alongside, same era.
Attribute-based transfer learning for object categorization with zero/one training example
X. Yu and Y. Aloimonos · 2010
Cited alongside, same era.
Baby talk: Understanding and generating image descriptions
G. Kulkarni, V. Premraj, S. Dhar, S. Li, Y. Choi, A. C. Berg, and T. L. Berg · 2011
Cited alongside, same era.
Im2text: Describing images using 1 million captioned photographs
V. Ordonez, G. Kulkarni, and T. L. Berg · 2011
Cited alongside, same era.
Action recognition by dense trajectories
H. Wang, A. Klaser, C. Schmid, and C.-L. Liu · 2011
Cited alongside, same era.
Learning everything about anything: Webly-supervised visual concept learning
S. K. Divvala, A. Farhadi, and C. Guestrin · 2014
Later among the works it cites.
Long-term recurrent convolutional networks for visual recognition and description
J. Donahue, L. A. Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell · 2014
Later among the works it cites.
Decaf: A deep convolutional activation feature for generic visual recognition
J. Donahue, Y. Jia, O. Vinyals, J. Hoffman, N. Zhang, E. Tzeng, and T. Darrell · 2014
Later among the works it cites.
Rich feature hierarchies for accurate object detection and semantic segmentation
R. Girshick, J. Donahue, T. Darrell, and J. Malik · 2014
Later among the works it cites.
Deep visual-semantic alignments for generating image descriptions
A. Karpathy and F.-F. Li · 2014
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Corpus-guided sentence generation of natural images
Y. Yang, C. L. Teo, H. Daumé, III, and Y. Aloimonos · 2011
Cited alongside, same era.
Active scene recognition with vision and language
X. Yu, C. Fermuller, C. L. Teo, Y. Yang, and Y. Aloimonos · 2011
Cited alongside, same era.
Collective generation of natural image descriptions
P. Kuznetsova, V. Ordonez, A. C. Berg, T. L. Berg, and Y. Choi · 2012
Cited alongside, same era.
Common-Sense Knowledge for a Computer Vision System for Human Action Recognition
M. Santofimia, J. Martinez-del Rincon, and J.-C. Nebel · 2012
Cited alongside, same era.
Representation learning: A review and new perspectives
Y. Bengio, A. Courville, and P. Vincent · 2013
Cited alongside, same era.
Image description using visual dependency representations
D. Elliott and F. Keller · 2013
Cited alongside, same era.
Unifying visual-semantic embeddings with multimodal neural language models
R. Kiros, R. Salakhutdinov, and R. S. Zemel · 2014
Later among the works it cites.
Explain images with multimodal recurrent neural networks
J. Mao, W. Xu, Y. Yang, J. Wang, and A. L. Yuille · 2014
Later among the works it cites.
Cnn features off-the-shelf: an astounding baseline for recognition
A. S. Razavian, H. Azizpour, J. Sullivan, and S. Carlsson · 2014
Later among the works it cites.
Grounded compositional semantics for finding and describing images with sentences
R. Socher, A. Karpathy, Q. V. Le, C. D. Manning, and A. Y. Ng · 2014
Later among the works it cites.
Show and tell: A neural image caption generator
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan · 2014
Later among the works it cites.
A cognitive system for understanding human manipulation actions
Y. Yang, C. Fermüller, Y. Aloimonos, and A. Guha · 2014
Later among the works it cites.
Learning deep features for scene recognition using places database
B. Zhou, A. Lapedriza, J. Xiao, A. Torralba, and A. Oliva · 2014
Later among the works it cites.
Commonsense reasoning and commonsense knowledge in artificial intelligence
E. Davis and G. Marcus · 2015
Closest in time.
Image retrieval using scene graphs
J. Johnson, R. Krishna, M. Stark, J. Li, M. Bernstein, and L. Fei-Fei · 2015
Closest in time.
Generating semantically precise scene graphs from textual descriptions for improved image retrieval
S. Schuster, R. Krishna, A. Chang, L. Fei-Fei, and C. D. Manning · 2015
Closest in time.
Towards addressing the winograd schema challenge - building and using a semantic parser and a knowledge hunting module
A. Sharma, N. H. Vo, S. Aditya, and C. Baral · 2015
Closest in time.
A gestaltist approach to contour-based object recognition: Combining bottom-up and top-down cues
C. L. Teo, C. Fermüller, and Y. Aloimonos · 2015
Closest in time.