Non-parametric local transforms for computing visual correspondence
R. Zabih and J. Woodfill · 1994
Earlier work this paper cites.
Space-time interest points
I. Laptev and T. Lindeberg · 2003
Earlier work this paper cites.
Hello! my name is buffy–automatic naming of characters in tv video
M. Everingham, J. Sivic, and A. Zisserman · 2006
Earlier work this paper cites.
Learning realistic human actions from movies
I. Laptev, M. Marszalek, C. Schmid, and B. Rozenfeld · 2008
Earlier work this paper cites.
Annotating images by mining image search results
X.-J. Wang, L. Zhang, X. Li, and W.-Y. Ma · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Automatic annotation of human actions in video
O. Duchenne, I. Laptev, J. Sivic, F. Bach, and J. Ponce · 2009
Earlier work this paper cites.
Actions in context
M. Marszalek, I. Laptev, and C. Schmid · 2009
Earlier work this paper cites.
Exploiting weakly-labeled web images to improve object classification: a domain adaptation approach
A. Bergamo and L. Torresani · 2010
Earlier work this paper cites.
Every picture tells a story: Generating sentences from images
A. Farhadi, M. Hejrati, M. A. Sadeghi, P. Young, C. Rashtchian, J. Hockenmaier, and D. Forsyth · 2010
Earlier work this paper cites.
Learning object categories from internet image searches
R. Fergus, L. Fei-Fei, P. Perona, and A. Zisserman · 2010
Earlier work this paper cites.
HMDB: a large video database for human motion recognition
H. Kuehne, H. Jhuang, E. Garrote, T. Poggio, and T. Serre · 2011
Earlier work this paper cites.
Im2text: Describing images using 1 million captioned photographs
V. Ordonez, G. Kulkarni, and T. L. Berg · 2011
Earlier work this paper cites.
Harvesting image databases from the web
F. Schroff, A. Criminisi, and A. Zisserman · 2011
Earlier work this paper cites.
What makes paris look like paris?
C. Doersch, S. Singh, A. Gupta, J. Sivic, and A. Efros · 2012
Earlier work this paper cites.
Learning object class detectors from weakly annotated video
A. Prest, C. Leistner, J. Civera, C. Schmid, and V. Ferrari · 2012
Earlier work this paper cites.
UCF101: A dataset of 101 human actions classes from videos in the wild
K. Soomro, A. R. Zamir, and M. Shah · 2012
Earlier work this paper cites.
Finding actors and actions in movies
P. Bojanowski, F. Bach, I. Laptev, J. Ponce, C. Schmid, and J. Sivic · 2013
Earlier work this paper cites.
Neil: Extracting visual knowledge from web data
X. Chen, A. Shrivastava, and A. Gupta · 2013
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
T. Mikolov, I. Sutskever, K. Chen, G. S. Corrado, and J. Dean · 2013
Earlier work this paper cites.
Learning everything about anything: Webly-supervised visual concept learning
S. K. Divvala, A. Farhadi, and C. Guestrin · 2014
Earlier work this paper cites.
Decaf: A deep convolutional activation feature for generic visual recognition
J. Donahue, Y. Jia, O. Vinyals, J. Hoffman, N. Zhang, E. Tzeng, and T. Darrell · 2014
Earlier work this paper cites.
Rich feature hierarchies for accurate object detection and semantic segmentation
R. Girshick, J. Donahue, T. Darrell, and J. Malik · 2014
Earlier work this paper cites.
THUMOS challenge: Action recognition with a large number of classes
Y. Jiang, J. Liu, A. R. Zamir, G. Toderici, I. Laptev, M. Shah, and R. Sukthankar · 2014
Earlier work this paper cites.