Fetching the paper…
Reading the bibliography…
Human communication typically has an underlying structure.
A tutorial on hidden markov models and selected applications in speech recognition
L. R. Rabiner · 1989
Earlier work this paper cites.
A factorization approach to grouping
P. Perona and W. Freeman · 1998
Earlier work this paper cites.
Automatically extracting highlights for tv baseball programs
Y. Rui, A. Gupta, and A. Acero · 2000
Earlier work this paper cites.
A mathematical theory of communication
C. E. Shannon · 2001
Earlier work this paper cites.
Recognizing action at a distance
A. A. Efros, A. C. Berg, G. Mori, and J. Malik · 2003
Earlier work this paper cites.
Infinite latent feature models and the indian buffet process
T. Griffiths and Z. Ghahramani · 2005
Earlier work this paper cites.
Clustering of time series data?a survey
T. W. Liao · 2005
Earlier work this paper cites.
Single-cluster spectral graph partitioning for robotics applications
E. Olson, M. Walter, S. J. Teller, and J. J. Leonard · 2005
Earlier work this paper cites.
Video abstraction: A systematic review and classification
B. T. Truong and S. Venkatesh · 2007
Earlier work this paper cites.
Learning realistic human actions from movies
I. Laptev, M. Marszalek, C. Schmid, and B. Rozenfeld · 2008
Earlier work this paper cites.
Automatic annotation of human actions in video
O. Duchenne, I. Laptev, J. Sivic, F. Bash, and J. Ponce · 2009
Earlier work this paper cites.
Generating photo manipulation tutorials by demonstration
F. Grabler, M. Agrawala, W. Li, M. Dontcheva, and T. Igarashi · 2009
Earlier work this paper cites.
Understanding videos, constructing plots learning a visually grounded storyline model from annotated videos
A. Gupta, P. Srinivasan, J. Shi, and L. S. Davis · 2009
Earlier work this paper cites.
Spatio-temporal relationship match: Video structure comparison for recognition of complex human activities
M. Ryoo and J. Aggarwal · 2009
Earlier work this paper cites.
Constrained parametric min-cuts for automatic object segmentation
J. Carreira and C. Sminchisescu · 2010
Earlier work this paper cites.
Every picture tells a story: Generating sentences from images
A. Farhadi, M. Hejrati, M. A. Sadeghi, P. Young, C. Rashtchian, J. Hockenmaier, and D. Forsyth · 2010
Earlier work this paper cites.
Modeling temporal structure of decomposable motion segments for activity classification
J. C. Niebles, C.-W. Chen, and L. Fei-Fei · 2010
Earlier work this paper cites.
Connecting modalities: Semi-supervised segmentation and annotation of images using unaligned text corpora
R. Socher and L. Fei-Fei · 2010
Earlier work this paper cites.
Understanding and executing instructions for everyday manipulation tasks from the world wide web
M. Tenorth, D. Nyga, and M. Beetz · 2010
Earlier work this paper cites.
Modeling mutual context of object and human pose in human-object interaction activities
B. Yao and L. Fei-Fei · 2010
Earlier work this paper cites.
Robotic roommates making pancakes
M. Beetz, U. Klank, I. Kresse, A. Maldonado, L. Mosenlechner, D. Pangercic, T. Ruhr, and M. Tenorth · 2011
Cited alongside, same era.
Bakebot: Baking cookies with the pr2
M. Bollini, J. Barry, and D. Rus · 2011
Cited alongside, same era.
Joint segmentation and classification of human actions in video
M. Hoai, Z.-Z. Lan, and F. De la Torre · 2011
Cited alongside, same era.
A large video database for human motion recognition
H. Kuehne, H. Jhuang, E. Garrote, T. Poggio, and T. Serre · 2011
Cited alongside, same era.
Key-segments for video object segmentation
Y. J. Lee, J. Kim, and K. Grauman · 2011
Cited alongside, same era.
Im2text: Describing images using 1 million captioned photographs
V. Ordonez, G. Kulkarni, and T. L. Berg · 2011
Cited alongside, same era.
Joint modeling of multiple related time series via the beta process with application to motion capture segmentation
E. Fox, M. Hughes, E. Sudderth, and M. Jordan · 2014
Later among the works it cites.
University of amsterdam at thumos challenge 2014
M. Jain, J. van Gemert, and C. G. Snoek · 2014
Later among the works it cites.
THUMOS challenge: Action recognition with a large number of classes
Y.-G. Jiang, J. Liu, A. Roshan Zamir, G. Toderici, I. Laptev, M. Shah, and R. Sukthankar · 2014
Later among the works it cites.
Efficient feature extraction, encoding and classification for action recognition
V. Kantorov and I. Laptev · 2014
Later among the works it cites.
Deep Visual-Semantic Alignments for Generating Image Descriptions
A. Karpathy and L. Fei-Fei · 2014
Later among the works it cites.
Joint summarization of large-scale collections of web images and videos for storyline reconstruction
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Barbu, A. Bridge, Z. Burchill, D. Coroian, S. Dickinson, S. Fidler, A. Michaux, S. Mussman, S. Narayanaswamy, D. Salvi, et al · 2012
Cited alongside, same era.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Cited alongside, same era.
Discovering important people and objects for egocentric video summarization
Y. J. Lee, J. Ghosh, and K. Grauman · 2012
Cited alongside, same era.
Improving video activity recognition using object recognition and text mining
T. S. Motwani and R. J. Mooney · 2012
Cited alongside, same era.
UCF101: A dataset of 101 human actions classes from videos in the wild
K. Soomro, A. Roshan Zamir, and M. Shah · 2012
Cited alongside, same era.
A thousand frames in just a few words: Lingual description of videos through latent topics and sparse object stitching
P. Das, C. Xu, R. F. Doell, and J. J. Corso · 2013
Cited alongside, same era.
G. Kim, L. Sigal, and E. P. Xing · 2014
Later among the works it cites.
Reconstructing storyline graphs for image recommendation from web community photos
G. Kim and E. P. Xing · 2014
Later among the works it cites.
Multimodal neural language models
R. Kiros, R. Salakhutdinov, and R. Zemel · 2014
Later among the works it cites.
What are you talking about? text-to-image coreference
C. Kong, D. Lin, M. Bansal, R. Urtasun, and S. Fidler · 2014
Later among the works it cites.
Learning action primitives for multi-level video event understanding
T. Lan, L. Chen, Z. Deng, G.-T. Zhou, and G. Mori · 2014
Later among the works it cites.
A hierarchical representation for future action prediction
T. Lan, T.-C. Chen, and S. Savarese · 2014
Later among the works it cites.
Cooking with semantics
J. Malmaud, E. J. Wagner, N. Chang, and K. Murphy · 2014
Later among the works it cites.
The lear submission at thumos 2014
D. Oneata, J. Verbeek, and C. Schmid · 2014
Later among the works it cites.
Parsing videos of actions with segmental grammars
H. Pirsiavash and D. Ramanan · 2014
Later among the works it cites.
Category-specific video summarization
D. Potapov, M. Douze, Z. Harchaoui, and C. Schmid · 2014
Later among the works it cites.
Robo brain: Large-scale knowledge engine for robots
A. Saxena, A. Jain, O. Sener, A. Jami, D. K. Misra, and H. S. Koppula · 2014
Later among the works it cites.
Grounded compositional semantics for finding and describing images with sentences
R. Socher, A. Karpathy, Q. V. Le, C. D. Manning, and A. Y. Ng · 2014
Later among the works it cites.
Discover: Discovering important segments for classification of video events and recounting
C. Sun and R. Nevatia · 2014
Later among the works it cites.
What’s Cookin’? Interpreting Cooking Videos using Text, Speech and Vision
J. Malmaud, J. Huang, V. Rathod, N. Johnston, A. Rabinovich, and K. Murphy · 2015
Closest in time.