Fetching the paper…
Reading the bibliography…
We present a system that produces sentential descriptions of video: who did what to whom, and where and how they did it.
Convolutional codes and their performance in communication systems
A. J. Viterbi · 1971
Earlier work this paper cites.
Logic and conversation
H. P. Grice · 1975
Earlier work this paper cites.
A threshold selection method from gray-level histograms
N. Otsu · 1979
Earlier work this paper cites.
Semantics and Cognition
Ray Jackendoff · 1983
Earlier work this paper cites.
Learnability and Cognition
Steven Pinker · 1989
Earlier work this paper cites.
Detection and tracking of point features
C. Tomasi and T. Kanade · 1991
Earlier work this paper cites.
AAAI Workshop on Integration of Natural Language and Vision Processing , 1994
P. McKevitt, editor · 1994
Earlier work this paper cites.
Good features to track
J. Shi and C. Tomasi · 1994
Earlier work this paper cites.
Towards 3-d model-based tracking and recognition of human movement
D. M. Gavrila and L. S. Davis · 1995
Earlier work this paper cites.
Integration of Natural Language and Vision Processing , volume I–IV
P. McKevitt, editor · 1996
Earlier work this paper cites.
A maximum-likelihood approach to visual event classification
J. M. Siskind and Q. Morris · 1996
Earlier work this paper cites.
Learning and recognizing human dynamics in video sequences
Christoph Bregler · 1997
Earlier work this paper cites.
Real-time American sign language recognition using desk and wearable computer based video
Thad Starner, Joshua Weaver, and Alex Pentland · 1998
Cited alongside, same era.
Motion based event recognition using HMM
Gu Xu, Yu-Fei Ma, HongJiang Zhang, and Shiqiang Yang · 2002
Cited alongside, same era.
HLT-NAACL Workshop on Learning Word Meaning from Non-Linguistic Data , 2003
R. Barzialy, E. Reiter, and J.M. Siskind, editors · 2003
Cited alongside, same era.
Recognizing human actions: A local SVM approach
C. Schuldt, I. Laptev, and B. Caputo · 2004
Cited alongside, same era.
Actions as space-time shapes
M. Blank, L. Gorelick, E. Shechtman, M. Irani, and R. Basri · 2005
Cited alongside, same era.
An HMM-based framework for video semantic analysis
Gu Xu, Yu-Fei Ma, HongJiang Zhang, and Shi-Qiang Yang · 2005
Cited alongside, same era.
Action MACH: A spatio-temporal maximum average correlation height filter for action recognition
M. D. Rodriguez, J. Ahmed, and M. Shah · 2008
Later among the works it cites.
Describing objects by their attributes
Ali Farhadi, Ian Endres, Derek Hoiem, and David Forsyth · 2009
Later among the works it cites.
Recognizing realistic actions from videos “in the wild”
J. Liu, J. Luo, and M. Shah · 2009
Later among the works it cites.
Human action recognition by semilatent topic models
Y. Wang and G. Mori · 2009
Later among the works it cites.
Object, scene and actions: Combining multiple features for human action recognition
Nazli Ikizler-Cinibis and Stan Sclaroff · 2010
Later among the works it cites.
HumanEva: Synchronized video and motion capture dataset and baseline algorithm for evaluation of articulated human motion
L. Sigal, A. Balan, and M. J. Black · 2010
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Local velocity-adapted motion events for spatio-temporal recognition
I. Laptev, B. Caputo, C. Schuldt, and T. Lindeberg · 2007
Cited alongside, same era.
A 3-dimensional SIFT descriptor and its application to action recognition
P. Scovanner, S. Ali, and M. Shah · 2007
Cited alongside, same era.
People-tracking-by-detection and people-detection-by-tracking
M. Andriluka, S. Roth, and B. Schiele · 2008
Cited alongside, same era.
Learning realistic human actions from movies
I. Laptev, M. Marszalek, C. Schmid, and B. Rozenfeld · 2008
Cited alongside, same era.
Unsupervised learning of human action categories using spatial-temporal words
J. C. Niebles, H. Wang, and L. Fei-Fei · 2008
Cited alongside, same era.
Cascade object detection with deformable part models
P. F. Felzenszwalb, R. B. Girshick, and D. McAllester
Cited in the paper.
Later among the works it cites.
I2t: Image parsing to text description
B. Z. Yao, Xiong Yang, Liang Lin, Mun Wai Lee, and Song-Chun Zhu · 2010
Later among the works it cites.
AAAI Workshop on Language-Action Tools for Cognitive Artificial Agents: Integrating Vision, Action and Language , 2011
Y. Aloimonos, L. Fadiga, G. Metta, and K. Pastra, editors · 2011
Later among the works it cites.
NIPS Workshop on Integrating Language and Vision , 2011
T. Darrell, R. Mooney, and K. Saenko, editors · 2011
Later among the works it cites.
Baby talk: Understanding and generating simple image descriptions
Girish Kulkarni, Visruth Premraj, Sagnik Dhar, Siming Li, Yejin Choi, Alexander C. Berg, and Tamara L. Berg · 2011
Later among the works it cites.
Articulated pose estimation using flexible mixtures of parts
Y. Yang and D. Ramanan · 2011
Later among the works it cites.