Fetching the paper…
Reading the bibliography…
This paper presents a restricted visual Turing test (VTT) for story-line based deep understanding in long-term and multi-camera captured videos.
Computing machinery and intelligence
A. M. Turing · 1950
Earlier work this paper cites.
Backpropagation applied to handwritten zip code recognition
Y. LeCun, B. E. Boser, J. S. Denker, D. Henderson, R. E. Howard, W. E. Hubbard, and L. D. Jackel · 1989
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Jena: A semantic web toolkit
B. McBride · 2002
Earlier work this paper cites.
A bayesian hierarchical model for learning natural scene categories
L. Fei-Fei and P. Perona · 2005
Earlier work this paper cites.
The pyramid match kernel: Efficient learning with sets of features
K. Grauman and T. Darrell · 2005
Earlier work this paper cites.
Beyond bags of features: Spatial pyramid matching for recognizing natural scene categories
S. Lazebnik, C. Schmid, and J. Ponce · 2006
Earlier work this paper cites.
A stochastic grammar of images
S.-C. Zhu and D. Mumford · 2006
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Vizwiz: nearly real-time answers to visual questions
J. P. Bigham, C. Jayant, H. Ji, G. Little, A. Miller, R. C. Miller, R. Miller, A. Tatarowicz, B. White, S. White, and T. Yeh · 2010
Earlier work this paper cites.
Every picture tells a story: Generating sentences from images
A. Farhadi, S. M. M. Hejrati, M. A. Sadeghi, P. Young, C. Rashtchian, J. Hockenmaier, and D. A. Forsyth · 2010
Earlier work this paper cites.
Object detection with discriminatively trained part-based models
P. F. Felzenszwalb, R. B. Girshick, D. McAllester, and D. Ramanan · 2010
Earlier work this paper cites.
Hierarchical semantic indexing for large scale image retrieval
J. Deng, A. C. Berg, and F. Li · 2011
Earlier work this paper cites.
Baby talk: Understanding and generating simple image descriptions
G. Kulkarni, V. Premraj, S. Dhar, S. Li, Y. Choi, A. C. Berg, and T. L. Berg · 2011
Earlier work this paper cites.
Globally-optimal greedy algorithms for tracking a variable number of objects
H. Pirsiavash, D. Ramanan, and C. C. Fowlkes · 2011
Cited alongside, same era.
Action recognition by dense trajectories
H. Wang, A. Kläser, C. Schmid, and C.-L. Liu · 2011
Cited alongside, same era.
Analyzing 3d objects in cluttered images
M. Hejrati and D. Ramanan · 2012
Cited alongside, same era.
Undoing the damage of dataset bias
A. Khosla, T. Zhou, T. Malisiewicz, A. A. Efros, and A. Torralba · 2012
Cited alongside, same era.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Cited alongside, same era.
Translating video content to natural language descriptions
M. Rohrbach, W. Qiu, I. Titov, S. Thater, M. Pinkal, and B. Schiele · 2013
Cited alongside, same era.
Vqa: Visual question answering
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh · 2015
Closest in time.
Visual turing test for computer vision systems
D. Geman, S. Geman, N. Hallonquist, and L. Younes · 2015
Closest in time.
Fast R-CNN
R. Girshick · 2015
Closest in time.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
K. He, X. Zhang, S. Ren, and J. Sun · 2015
Closest in time.
Learning 3d object templates by quantizing geometry and appearance spaces
W. Hu and S. Zhu · 2015
Closest in time.
MOTChallenge 2015: Towards a benchmark for multi-target tracking
L. Leal-Taixé, A. Milan, I. Reid, S. Roth, and K. Schindler · 2015
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Discriminatively trained and-or tree models for object detection
X. Song, T. Wu, Y. Jia, and S.-C. Zhu · 2013
Cited alongside, same era.
Discriminatively trained and-or tree models for object detection
X. Song, T.-F. Wu, Y. Jia, and S.-C. Zhu · 2013
Cited alongside, same era.
The pascal visual object classes challenge: A retrospective
M. Everingham, S. A. Eslami, L. Van Gool, C. K. Williams, J. Winn, and A. Zisserman · 2014
Cited alongside, same era.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Cited alongside, same era.
Single-view 3d scene parsing by attributed grammar
X. Liu, Y. Zhao, and S.-C. Zhu · 2014
Cited alongside, same era.
A multi-world approach to question answering about real-world scenes based on uncertain input
M. Malinowski and M. Fritz · 2014
Cited alongside, same era.
Deep captioning with multimodal recurrent neural networks (m-rnn)
J. Mao, W. Xu, Y. Yang, J. Wang, Z. Huang, and A. Yuille · 2015
Closest in time.
Attributed grammars for joint estimation of human attributes, part and pose
S. Park and S.-C. Zhu · 2015
Closest in time.
Image question answering: A visual semantic embedding model and a new dataset
M. Ren, R. Kiros, and R. Zemel · 2015
Closest in time.
Faster R-CNN: Towards real-time object detection with region proposal networks
S. Ren, K. He, R. Girshick, and J. Sun · 2015
Closest in time.
ImageNet Large Scale Visual Recognition Challenge
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. C. Berg, and L. Fei-Fei · 2015
Closest in time.
Learning and-or models to represent context and occlusion for car detection and viewpoint estimation
T. Wu, B. Li, and S.-C. Zhu · 2015
Closest in time.
Joint action recognition and pose estimation from video
B. Xiaohan Nie, C. Xiong, and S.-C. Zhu · 2015
Closest in time.
A reconfigurable tangram model for scene representation and categorization
J. Zhu, T. Wu, S.-C. Zhu, X. Yang, and W. Zhang · 2015
Closest in time.