Fetching the paper…
Reading the bibliography…
Audio Description (AD) provides linguistic descriptions of movies and allows visually impaired people to follow a movie along with their peers.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
The berkeley framenet project
Collin F. Baker, Charles J. Fillmore, and John B. Lowe · 1998
Earlier work this paper cites.
WordNet: An Electronic Lexical Database
Christiane Fellbaum · 1998
Earlier work this paper cites.
Natural language description of human activities from video images based on concept hierarchy of actions
Atsuhiro Kojima, Takeshi Tamura, and Kunio Fukunaga · 2002
Earlier work this paper cites.
BLEU: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei jing Zhu · 2002
Earlier work this paper cites.
Feature-rich part-of-speech tagging with a cyclic dependency network
Kristina Toutanova, Dan Klein, Christopher D. Manning, and Yoram Singer · 2003
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Chin-Yew Lin · 2004
Earlier work this paper cites.
Wordnet:: Similarity: measuring the relatedness of concepts
Ted Pedersen, Siddharth Patwardhan, and Jason Michelizzi · 2004
Earlier work this paper cites.
”hello! my name is… buffy” - automatic naming of characters in tv video
Mark Everingham, Josef Sivic, and Andrew Zisserman · 2006
Earlier work this paper cites.
Extending verbnet with novel verb classes
Karen Kipper, Anna Korhonen, Neville Ryant, and Martha Palmer · 2006
Earlier work this paper cites.
The semi-automatic generation of audio description from screenplays
Lakritz and Salway · 2006
Earlier work this paper cites.
Moses: Open source toolkit for statistical machine translation
Philipp Koehn, Hieu Hoang, Alexandra Birch, Chris Callison-Burch, Marcello Federico, Nicola Bertoldi, Brooke Cowan, Wade Shen, Christine Moran, Richard Zens, Chris Dyer, Ondrej Bojar, Alexandra Constantin, and Evan Herbst · 2007
Earlier work this paper cites.
A corpus-based analysis of audio description
Andrew Salway · 2007
Earlier work this paper cites.
Associating characters with events in films
Andrew Salway, Bart Lehane, and Noel E. O’Connor · 2007
Earlier work this paper cites.
Movie/script: Alignment and parsing of video and text transcription
Timothée Cour, Chris Jordan, Eleni Miltsakaki, and Ben Taskar · 2008
Earlier work this paper cites.
Learning realistic human actions from movies
Ivan Laptev, Marcin Marszalek, Cordelia Schmid, and Benjamin Rozenfeld · 2008
Earlier work this paper cites.
Learning from ambiguously labeled images
Timothee Cour, Benjamin Sapp, Chris Jordan, and Ben Taskar · 2009
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Automatic annotation of human actions in video
Olivier Duchenne, Ivan Laptev, Josef Sivic, Francis Bach, and Jean Ponce · 2009
Earlier work this paper cites.
Actions in context
Marcin Marszalek, Ivan Laptev, and Cordelia Schmid · 2009
Earlier work this paper cites.
Verbnet overview, extensions, mappings and applications
Karin Kipper Schuler, Anna Korhonen, and Susan Windisch Brown · 2009
Earlier work this paper cites.
”who are you?”-learning person specific classifiers from video
Josef Sivic, Mark Everingham, and Andrew Zisserman · 2009
Earlier work this paper cites.
Every picture tells a story: Generating sentences from images
Ali Farhadi, Mohsen Hejrati, M.A. Sadeghi, Peter Young, C. Rashtchian, Julia Hockenmaier, and D.A. Forsyth · 2010
Earlier work this paper cites.
A computer-vision-assisted system for videodescription scripting
Langis Gagnon, Claude Chapdelaine, David Byrns, Samuel Foucher, Maguelonne Heritier, and Vishwa Gupta · 2010
Earlier work this paper cites.
Sun database: Large-scale scene recognition from abbey to zoo
Jianxiong Xiao, James Hays, Krista A. Ehinger, Aude Oliva, and Antonio Torralba · 2010
Earlier work this paper cites.
It makes sense: A wide-coverage word sense disambiguation system for free text
Zhi Zhong and Hwee Tou Ng · 2010
Earlier work this paper cites.
Collecting highly parallel data for paraphrase evaluation
David Chen and William Dolan · 2011
Earlier work this paper cites.
Baby talk: Understanding and generating simple image descriptions
Girish Kulkarni, Visruth Premraj, Sagnik Dhar, Siming Li, Yejin Choi, Alexander C. Berg, and Tamara L. Berg · 2011
Earlier work this paper cites.
Composing simple image descriptions using web-scale N-grams
Siming Li, Girish Kulkarni, Tamara L Berg, Alexander C Berg, and Yejin Choi · 2011
Earlier work this paper cites.
Tvparser: An automatic tv video parsing method
Chao Liang, Changsheng Xu, Jian Cheng, and Hanqing Lu · 2011
Cited alongside, same era.
Im2text: Describing images using 1 million captioned photographs
Vicente Ordonez, Girish Kulkarni, and Tamara L. Berg · 2011
Cited alongside, same era.
Video in sentences out
Andrei Barbu, Alexander Bridge, Zachary Burchill, Dan Coroian, Sven Dickinson, Sanja Fidler, Aaron Michaux, Sam Mussman, Siddharth Narayanaswamy, Dhaval Salvi, Lara Schmidt, Jiangnan Shangguan, Jeffrey Mark Siskind, Jarrell Waggoner, Song Wang, Jinlian Wei, Yifan Yin, and Zhiqi Zhang · 2012
Cited alongside, same era.
An exact dual decomposition algorithm for shallow semantic parsing with constraints
Dipanjan Das, André F.T. Martins, and Noah A. Smith · 2012
Cited alongside, same era.
Improving neural networks by preventing co-adaptation of feature detectors
Geoffrey E Hinton, Nitish Srivastava, Alex Krizhevsky, Ilya Sutskever, and Ruslan R Salakhutdinov · 2012
Cited alongside, same era.
Grounded compositional semantics for finding and describing images with sentences
Richard Socher, Andrej Karpathy, Quoc V. Le, Christopher D. Manning, and Andrew Y. Ng · 2014
Later among the works it cites.
Integrating language and vision to generate natural language descriptions of videos in the wild
Jesse Thomason, Subhashini Venugopalan, Sergio Guadarrama, Kate Saenko, and Raymond J. Mooney · 2014
Later among the works it cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Peter Young, Alice Lai, Micah Hodosh, and Julia Hockenmaier · 2014
Later among the works it cites.
Learning Deep Features for Scene Recognition using Places Database
B. Zhou, A. Lapedriza, J. Xiao, A. Torralba, and A. Oliva · 2014
Later among the works it cites.
Mind’s eye: A recurrent visual representation for image caption generation
Xinlei Chen and C. Lawrence Zitnick · 2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton · 2012
Cited alongside, same era.
Collective generation of natural image descriptions
Polina Kuznetsova, Vicente Ordonez, Alexander C Berg, Tamara L Berg, and Yejin Choi · 2012
Cited alongside, same era.
Midge: Generating image descriptions from computer vision detections
Margaret Mitchell, Jesse Dodge, Amit Goyal, Kota Yamaguchi, Karl Stratos, Xufeng Han, Alyssa Mensch, Alexander C. Berg, Tamara L. Berg, and Hal Daumé III · 2012
Cited alongside, same era.
Trecvid 2012 – an overview of the goals, tasks, data, evaluation mechanisms and metrics
Paul Over, George Awad, Martial Michel, Jonathan Fiscus, Greg Sanders, B Shaw, Alan F. Smeaton, and Georges Quéenot · 2012
Cited alongside, same era.
”knock! knock! who is it?” probabilistic person identification in tv-series
Makarand Tapaswi, Martin Baeuml, and Rainer Stiefelhagen · 2012
Cited alongside, same era.
Finding actors and actions in movies
Piotr Bojanowski, Francis Bach, Ivan Laptev, Jean Ponce, Cordelia Schmid, and Josef Sivic · 2013
Cited alongside, same era.
Thousand frames in just a few words: Lingual description of videos through latent topics and sparse object stitching
Pradipto Das, Chenliang Xu, Richard Doell, and Jason Corso · 2013
Cited alongside, same era.
Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Dollár, and C Lawrence Zitnick · 2015
Later among the works it cites.
Language models for image captioning: The quirks and what works
Jacob Devlin, Hao Cheng, Hao Fang, Saurabh Gupta, Li Deng, Xiaodong He, Geoffrey Zweig, and Margaret Mitchell · 2015
Later among the works it cites.
Long-term recurrent convolutional networks for visual recognition and description
Jeff Donahue, Lisa Anne Hendricks, Sergio Guadarrama, Marcus Rohrbach, Subhashini Venugopalan, Kate Saenko, and Trevor Darrell · 2015
Later among the works it cites.
From captions to visual concepts and back
Hao Fang, Saurabh Gupta, Forrest N. Iandola, Rupesh Srivastava, Li Deng, Piotr Dollár, Jianfeng Gao, Xiaodong He, Margaret Mitchell, John C. Platt, C. Lawrence Zitnick, and Geoffrey Zweig · 2015
Later among the works it cites.
Deep visual-semantic alignments for generating image descriptions
Andrej Karpathy and Li Fei-Fei · 2015
Later among the works it cites.
Unifying visual-semantic embeddings with multimodal neural language models
Ryan Kiros, Ruslan Salakhutdinov, and Richard S. Zemel · 2015
Later among the works it cites.
Summarization-based video caption via deep neural networks
Guang Li, Shubo Ma, and Yahong Han · 2015
Later among the works it cites.
Deep captioning with multimodal recurrent neural networks (m-rnn)
Junhua Mao, Wei Xu, Yi Yang, Jiang Wang, Zhiheng Huang, and Alan Yuille · 2015
Later among the works it cites.
Rakshith Shetty and Jorma Laaksonen · 2015
Later among the works it cites.
Knowlywood: Mining activity knowledge from hollywood narratives
Niket Tandon, Gerard de Melo, Abir De, and Gerhard Weikum · 2015
Later among the works it cites.
Using descriptive video services to create a large data source for video annotation research
Atousa Torabi, Christopher Pal, Hugo Larochelle, and Aaron Courville · 2015
Later among the works it cites.
Cider: Consensus-based image description evaluation
Ramakrishna Vedantam, C Lawrence Zitnick, and Devi Parikh · 2015
Later among the works it cites.
Show and tell: A neural image caption generator
Oriol Vinyals, Alexander Toshev, Samy Bengio, and Dumitru Erhan · 2015
Later among the works it cites.
Describing videos by exploiting temporal structure
Li Yao, Atousa Torabi, Kyunghyun Cho, Nicolas Ballas, Christopher Pal, Hugo Larochelle, and Aaron Courville · 2015
Later among the works it cites.
Delving deeper into convolutional networks for learning video representations
Nicolas Ballas, Li Yao, Chris Pal, and Aaron Courville · 2016
Closest in time.
Seeing is believing: The quest for multimodal knowledge
Gerard de Melo and Niket Tandon · 2016
Closest in time.
Deep compositional captioning: Describing novel object categories without paired training data
Lisa Anne Hendricks, Subhashini Venugopalan, Marcus Rohrbach, Raymond Mooney, Kate Saenko, and Trevor Darrell · 2016
Closest in time.
Tgif: A new dataset and benchmark on animated gif description
Yuncheng Li, Yale Song, Liangliang Cao, Joel Tetreault, Larry Goldberg, Alejandro Jaimes, and Jiebo Luo · 2016
Closest in time.
Movieqa: Understanding stories in movies through question-answering
Makarand Tapaswi, Yukun Zhu, Rainer Stiefelhagen, Antonio Torralba, Raquel Urtasun, and Sanja Fidler · 2016
Closest in time.
Improving lstm-based video description with linguistic knowledge mined from text
Subhashini Venugopalan, Lisa Anne Hendricks, Raymond Mooney, and Kate Saenko · 2016
Closest in time.
Msr-vtt: A large video description dataset for bridging video and language
Jun Xu, Tao Mei, Ting Yao, and Yong Rui · 2016
Closest in time.
Empirical performance upper bounds for image and video captioning
Li Yao, Nicolas Ballas, Kyunghyun Cho, John R Smith, and Yoshua Bengio · 2016
Closest in time.
Video paragraph captioning using hierarchical recurrent neural networks
Haonan Yu, Jiang Wang, Zhiheng Huang, Yi Yang, and Wei Xu · 2016
Closest in time.