Fetching the paper…
Reading the bibliography…
Descriptive video service (DVS) provides linguistic descriptions of movies and allows visually impaired people to follow a movie along with their peers.
The berkeley framenet project
C. F. Baker, C. J. Fillmore, and J. B. Lowe · 1998
Earlier work this paper cites.
WordNet: An Electronic Lexical Database
C. Fellbaum, editor · 1998
Earlier work this paper cites.
Natural language description of human activities from video images based on concept hierarchy of actions
A. Kojima, T. Tamura, and K. Fukunaga · 2002
Earlier work this paper cites.
Wordnet:: Similarity: measuring the relatedness of concepts
T. Pedersen, S. Patwardhan, and J. Michelizzi · 2004
Earlier work this paper cites.
Extending verbnet with novel verb classes
K. Kipper, A. Korhonen, N. Ryant, and M. Palmer · 2006
Earlier work this paper cites.
The semi-automatic generation of audio description from screenplays
Lakritz and Salway · 2006
Earlier work this paper cites.
Moses: Open source toolkit for statistical machine translation
P. Koehn, H. Hoang, A. Birch, C. Callison-Burch, M. Federico, N. Bertoldi, B. Cowan, W. Shen, C. Moran, R. Zens, C. Dyer, O. Bojar, A. Constantin, and E. Herbst · 2007
Earlier work this paper cites.
A corpus-based analysis of audio description
A. Salway · 2007
Earlier work this paper cites.
Associating characters with events in films
A. Salway, B. Lehane, and N. E. O’Connor · 2007
Earlier work this paper cites.
Movie/script: Alignment and parsing of video and text transcription
T. Cour, C. Jordan, E. Miltsakaki, and B. Taskar · 2008
Earlier work this paper cites.
Learning realistic human actions from movies
I. Laptev, M. Marszalek, C. Schmid, and B. Rozenfeld · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Automatic annotation of human actions in video
O. Duchenne, I. Laptev, J. Sivic, F. Bach, and J. Ponce · 2009
Earlier work this paper cites.
Understanding videos, constructing plots learning a visually grounded storyline model from annotated videos
A. Gupta, P. Srinivasan, J. Shi, and L. Davis · 2009
Earlier work this paper cites.
Actions in context
M. Marszalek, I. Laptev, and C. Schmid · 2009
Earlier work this paper cites.
Verbnet overview, extensions, mappings and applications
K. K. Schuler, A. Korhonen, and S. W. Brown · 2009
Earlier work this paper cites.
Every picture tells a story: Generating sentences from images
A. Farhadi, M. Hejrati, M. Sadeghi, P. Young, C. Rashtchian, J. Hockenmaier, and D. Forsyth · 2010
Earlier work this paper cites.
A computer-vision-assisted system for videodescription scripting
L. Gagnon, C. Chapdelaine, D. Byrns, S. Foucher, M. Heritier, and V. Gupta · 2010
Earlier work this paper cites.
Sun database: Large-scale scene recognition from abbey to zoo
J. Xiao, J. Hays, K. A. Ehinger, A. Oliva, and A. Torralba · 2010
Earlier work this paper cites.
It makes sense: A wide-coverage word sense disambiguation system for free text
Z. Zhong and H. T. Ng · 2010
Earlier work this paper cites.
Collecting highly parallel data for paraphrase evaluation
D. Chen and W. Dolan · 2011
Earlier work this paper cites.
Human focused video description
M. U. G. Khan, L. Zhang, and Y. Gotoh · 2011
Earlier work this paper cites.
Baby talk: Understanding and generating simple image descriptions
G. Kulkarni, V. Premraj, S. Dhar, S. Li, Y. Choi, A. C. Berg, and T. L. Berg · 2011
Earlier work this paper cites.
Composing simple image descriptions using web-scale N-grams
S. Li, G. Kulkarni, T. Berg, A. Berg, and Y. Choi · 2011
Cited alongside, same era.
Tvparser: An automatic tv video parsing method
C. Liang, C. Xu, J. Cheng, and H. Lu · 2011
Cited alongside, same era.
Im2text: Describing images using 1 million captioned photographs
V. Ordonez, G. Kulkarni, and T. L. Berg · 2011
Cited alongside, same era.
Towards textually describing complex video contents with audio-visual concept classifiers
C. C. Tan, Y.-G. Jiang, and C.-W. Ngo · 2011
Cited alongside, same era.
Video in sentences out
A. Barbu, A. Bridge, Z. Burchill, D. Coroian, S. Dickinson, S. Fidler, A. Michaux, S. Mussman, S. Narayanaswamy, D. Salvi, L. Schmidt, J. Shangguan, J. M. Siskind, J. Waggoner, S. Wang, J. Wei, Y. Yin, and Z. Zhang · 2012
Cited alongside, same era.
An exact dual decomposition algorithm for shallow semantic parsing with constraints
http://www.makemkv.com/ , 2014
Makemkv · 2014
Later among the works it cites.
http://www.nikse.dk/SubtitleEdit/ , 2014
Subtitle edit · 2014
Later among the works it cites.
http://www.xmedia-recode.de/ , 2014
Xmedia recode · 2014
Later among the works it cites.
Weakly supervised action labeling in videos under ordering constraints
P. Bojanowski, R. Lajugie, F. Bach, I. Laptev, J. Ponce, C. Schmid, and J. Sivic · 2014
Later among the works it cites.
Learning a recurrent visual representation for image caption generation
X. Chen and C. L. Zitnick · 2014
Later among the works it cites.
Long-term recurrent convolutional networks for visual recognition and description
J. Donahue, L. A. Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell · 2014
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
D. Das, A. F. Martins, and N. A. Smith · 2012
Cited alongside, same era.
Automated textual descriptions for a wide range of video events with 48 human actions
P. Hanckmann, K. Schutte, and G. J. Burghouts · 2012
Cited alongside, same era.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Cited alongside, same era.
Collective generation of natural image descriptions
P. Kuznetsova, V. Ordonez, A. C. Berg, T. L. Berg, and Y. Choi · 2012
Cited alongside, same era.
Midge: Generating image descriptions from computer vision detections
M. Mitchell, J. Dodge, A. Goyal, K. Yamaguchi, K. Stratos, X. Han, A. Mensch, A. C. Berg, T. L. Berg, and H. D. III · 2012
Cited alongside, same era.
Trecvid 2012 – an overview of the goals, tasks, data, evaluation mechanisms and metrics
P. Over, G. Awad, M. Michel, J. Fiscus, G. Sanders, B. Shaw, A. F. Smeaton, and G. Quéenot · 2012
Cited alongside, same era.
Semantic parsing with combinatory categorial grammars
Y. Artzi, N. FitzGerald, and L. S. Zettlemoyer · 2013
Cited alongside, same era.
Later among the works it cites.
Open question answering over curated and extracted knowledge bases
A. Fader, L. Zettlemoyer, and O. Etzioni · 2014
Later among the works it cites.
From captions to visual concepts and back
H. Fang, S. Gupta, F. N. Iandola, R. Srivastava, L. Deng, P. Dollár, J. Gao, X. He, M. Mitchell, J. C. Platt, C. L. Zitnick, and G. Zweig · 2014
Later among the works it cites.
LSDA: Large scale detection through adaptation
J. Hoffman, S. Guadarrama, E. Tzeng, J. Donahue, R. Girshick, T. Darrell, and K. Saenko · 2014
Later among the works it cites.
Deep visual-semantic alignments for generating image descriptions
A. Karpathy and L. Fei-Fei · 2014
Later among the works it cites.
Multimodal neural language models
R. Kiros, R. Salakhutdinov, and R. Zemel · 2014
Later among the works it cites.
Unifying visual-semantic embeddings with multimodal neural language models
R. Kiros, R. Salakhutdinov, and R. S. Zemel · 2014
Later among the works it cites.
Treetalk: Composition and compression of trees for image descriptions
P. Kuznetsova, V. Ordonez, T. L. Berg, U. C. Hill, and Y. Choi · 2014
Later among the works it cites.
Context-dependent semantic parsing for time expressions
K. Lee, Y. Artzi, J. Dodge, and L. Zettlemoyer · 2014
Later among the works it cites.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Later among the works it cites.
Deep captioning with multimodal recurrent neural networks (m-rnn)
J. Mao, W. Xu, Y. Yang, J. Wang, and A. L. Yuille · 2014
Later among the works it cites.
Coherent multi-sentence video description with variable level of detail
A. Rohrbach, M. Rohrbach, W. Qiu, A. Friedrich, M. Pinkal, and B. Schiele · 2014
Later among the works it cites.
ImageNet Large Scale Visual Recognition Challenge, 2014
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. C. Berg, and L. Fei-Fei · 2014
Later among the works it cites.
Grounded compositional semantics for finding and describing images with sentences
R. Socher, A. Karpathy, Q. V. Le, C. D. Manning, and A. Y. Ng · 2014
Later among the works it cites.
Integrating language and vision to generate natural language descriptions of videos in the wild
J. Thomason, S. Venugopalan, S. Guadarrama, K. Saenko, and R. J. Mooney · 2014
Later among the works it cites.
Translating videos to natural language using deep recurrent neural networks
S. Venugopalan, H. Xu, J. Donahue, M. Rohrbach, R. Mooney, and K. Saenko · 2014
Later among the works it cites.
Show and tell: A neural image caption generator
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan · 2014
Later among the works it cites.
Learning Deep Features for Scene Recognition using Places Database
B. Zhou, A. Lapedriza, J. Xiao, A. Torralba, and A. Oliva · 2014
Later among the works it cites.