Fetching the paper…
Reading the bibliography…
Humans can easily describe what they see in a coherent way and at varying level of detail.
Wordnet: a lexical database for english
G. A. Miller · 1995
Earlier work this paper cites.
Natural language processing and user modeling: Synergies and limitations
I. Zukerman and D. Litman · 2001
Earlier work this paper cites.
Accurate unlexicalized parsing
D. Klein and C. D. Manning · 2003
Earlier work this paper cites.
Feature-rich part-of-speech tagging with a cyclic dependency network
K. Toutanova, D. Klein, C. D. Manning, and Y. Singer · 2003
Earlier work this paper cites.
Moses: Open source toolkit for statistical machine translation
P. Koehn, H. Hoang, A. Birch, C. Callison-Burch, M. Federico, N. Bertoldi, B. Cowan, W. Shen, C. Moran, R. Zens, C. Dyer, O. Bojar, A. Constantin, and E. Herbst · 2007
Earlier work this paper cites.
Generalizing word lattice translation
C. Dyer, S. Muresan, and P. Resnik · 2008
Earlier work this paper cites.
Understanding videos, constructing plots learning a visually grounded storyline model from annotated videos
A. Gupta, P. Srinivasan, J. B. Shi, and L. Davis · 2009
Earlier work this paper cites.
Every picture tells a story: Generating sentences from images
A. Farhadi, M. Hejrati, M. Sadeghi, P. Young, C. Rashtchian, J. Hockenmaier, and D. Forsyth · 2010
Earlier work this paper cites.
Towards coherent natural language description of video streams
M. U. G. Khan, L. Zhang, and Y. Gotoh · 2011
Cited alongside, same era.
Baby talk: Understanding and generating simple image descriptions
G. Kulkarni, V. Premraj, S. Dhar, S. Li, Y. Choi, A. C. Berg, and T. L. Berg · 2011
Cited alongside, same era.
Towards textually describing complex video contents with audio-visual concept classifiers
C. C. Tan, Y.-G. Jiang, and C.-W. Ngo · 2011
Cited alongside, same era.
Collective generation of natural image descriptions
P. Kuznetsova, V. Ordonez, A. C. Berg, T. L. Berg, and Y. Choi · 2012
Cited alongside, same era.
Midge: Generating image descriptions from computer vision detections
M. Mitchell, J. Dodge, A. Goyal, K. Yamaguchi, K. Stratos, X. Han, A. Mensch, A. C. Berg, T. L. Berg, and H. D. III · 2012
Cited alongside, same era.
A database for fine grained activity detection of cooking activities
Youtube2text: Recognizing and describing arbitrary activities using semantic hierarchies and zero-shot recognition
S. Guadarrama, N. Krishnamoorthy, G. Malkarnenkar, R. Mooney, T. Darrell, and K. Saenko · 2013
Later among the works it cites.
Docent: A document-level decoder for phrase-based statistical machine translation
C. Hardmeier, S. Stymne, J. Tiedemann, and J. Nivre · 2013
Later among the works it cites.
Generating natural-language video descriptions using text-mined knowledge
N. Krishnamoorthy, G. Malkarnenkar, R. J. Mooney, K. Saenko, and S. Guadarrama · 2013
Later among the works it cites.
Grounding action descriptions in videos
M. Regneri, M. Rohrbach, D. Wetzel, S. Thater, B. Schiele, and M. Pinkal · 2013
Later among the works it cites.
Translating video content to natural language descriptions
M. Rohrbach, W. Qiu, I. Titov, S. Thater, M. Pinkal, and B. Schiele · 2013
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M. Rohrbach, S. Amin, M. Andriluka, and B. Schiele · 2012
Cited alongside, same era.
Script data for attribute-based recognition of composite activities
M. Rohrbach, M. Regneri, M. Andriluka, S. Amin, M. Pinkal, and B. Schiele · 2012
Cited alongside, same era.
Thousand frames in just a few words: Lingual description of videos through latent topics and sparse object stitching
P. Das, C. Xu, R. F. Doell, and J. Corso · 2013
Cited alongside, same era.
Discriminative appearance models for pictorial structures
M. Andriluka, S. Roth, and B. Schiele
Cited in the paper.
Hand detection using multiple proposals
A. Mittal, A. Zisserman, and P. Torr
Cited in the paper.
M. Schmidt · 2013
Later among the works it cites.
Dense trajectories and motion boundary descriptors for action recognition
H. Wang, A. Kläser, C. Schmid, and C. Liu · 2013
Later among the works it cites.
Grounded language learning from videos described with sentences
H. Yu and J. M. Siskind · 2013
Later among the works it cites.