Fetching the paper…
Reading the bibliography…
We consider retrieving a specific temporal segment, or moment, from a video given a natural language text description.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Video shot boundary detection based on color histogram
J. Mas and G. Fernandez · 2003
Earlier work this paper cites.
Detecting irregularities in images and in video
O. Boiman and M. Irani · 2007
Earlier work this paper cites.
Natural language processing with Python
S. Bird, E. Klein, and E. Loper · 2009
Earlier work this paper cites.
Towards surveillance video search by natural language query
S. Tellex and D. Roy · 2009
Earlier work this paper cites.
Collecting highly parallel data for paraphrase evaluation
D. L. Chen and W. B. Dolan · 2011
Earlier work this paper cites.
A thousand frames in just a few words: Lingual description of videos through latent topics and sparse object stitching
P. Das, C. Xu, R. F. Doell, and J. J. Corso · 2013
Earlier work this paper cites.
Devise: A deep visual-semantic embedding model
A. Frome, G. S. Corrado, J. Shlens, S. Bengio, J. Dean, T. Mikolov, et al · 2013
Earlier work this paper cites.
Grounding action descriptions in videos
M. Regneri, M. Rohrbach, D. Wetzel, S. Thater, B. Schiele, and M. Pinkal · 2013
Earlier work this paper cites.
Translating video content to natural language descriptions
M. Rohrbach, W. Qiu, I. Titov, S. Thater, M. Pinkal, and B. Schiele · 2013
Earlier work this paper cites.
Grounded language learning from video described with sentences
H. Yu and J. M. Siskind · 2013
Earlier work this paper cites.
Caffe: Convolutional architecture for fast feature embedding
Y. Jia, E. Shelhamer, J. Donahue, S. Karayev, J. Long, R. Girshick, S. Guadarrama, and T. Darrell · 2014
Earlier work this paper cites.
Deep fragment embeddings for bidirectional image sentence mapping
A. Karpathy, A. Joulin, and F. F. F. Li · 2014
Earlier work this paper cites.
Referitgame: Referring to objects in photographs of natural scenes
S. Kazemzadeh, V. Ordonez, M. Matten, and T. L. Berg · 2014
Earlier work this paper cites.
Visual semantic search: Retrieving videos via complex textual queries
D. Lin, S. Fidler, C. Kong, and R. Urtasun · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
J. Pennington, R. Socher, and C. D. Manning · 2014
Earlier work this paper cites.
Coherent multi-sentence video description with variable level of detail
A. Rohrbach, M. Rohrbach, W. Qiu, A. Friedrich, M. Pinkal, and B. Schiele · 2014
Earlier work this paper cites.
Grounded compositional semantics for finding and describing images with sentences
R. Socher, A. Karpathy, Q. V. Le, C. D. Manning, and A. Y. Ng · 2014
Earlier work this paper cites.
Videoset: Video summary evaluation through text
S. Yeung, A. Fathi, and L. Fei-Fei · 2014
Earlier work this paper cites.
Weakly-supervised alignment of video with text
P. Bojanowski, R. Lajugie, E. Grave, F. Bach, I. Laptev, J. Ponce, and C. Schmid · 2015
Cited alongside, same era.
Microsoft COCO captions: Data collection and evaluation server
X. Chen, T.-Y. L. Hao Fang, R. Vedantam, S. Gupta, P. Dollár, and C. L. Zitnick · 2015
Cited alongside, same era.
Video summarization by learning submodular mixtures of objectives
M. Gygli, H. Grabner, and L. Van Gool · 2015
Cited alongside, same era.
Associating neural word embeddings with deep image representations using fisher vectors
B. Klein, G. Lev, G. Sadeh, and L. Wolf · 2015
Cited alongside, same era.
Multi-task deep visual-semantic embedding for video thumbnail selection
W. Liu, T. Mei, Y. Zhang, C. Che, and J. Luo · 2015
Cited alongside, same era.
Unsupervised learning from narrated instruction videos
J.-B. Alayrac, P. Bojanowski, N. Agrawal, J. Sivic, I. Laptev, and S. Lacoste-Julien · 2016
Later among the works it cites.
Video2gif: Automatic generation of animated gifs from video
M. Gygli, Y. Song, and L. Cao · 2016
Later among the works it cites.
Natural language object retrieval
R. Hu, H. Xu, M. Rohrbach, J. Feng, K. Saenko, and T. Darrell · 2016
Later among the works it cites.
Tgif: A new dataset and benchmark on animated gif description
Y. Li, Y. Song, L. Cao, J. Tetreault, L. Goldberg, A. Jaimes, and J. Luo · 2016
Later among the works it cites.
Generation and comprehension of unambiguous object descriptions
J. Mao, J. Huang, A. Toshev, O. Camburu, A. Yuille, and K. Murphy · 2016
Later among the works it cites.
Learning joint representations of videos and sentences with web image search
M. Otani, Y. Nakashima, E. Rahtu, J. Heikkilä, and N. Yokoya · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
B. A. Plummer, L. Wang, C. M. Cervantes, J. C. Caicedo, J. Hockenmaier, and S. Lazebnik · 2015
Cited alongside, same era.
The long-short story of movie description
A. Rohrbach, M. Rohrbach, and B. Schiele · 2015
Cited alongside, same era.
A dataset for movie description
A. Rohrbach, M. Rohrbach, N. Tandon, and B. Schiele · 2015
Cited alongside, same era.
Imagenet large scale visual recognition challenge
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, et al · 2015
Cited alongside, same era.
Unsupervised semantic parsing of video collections
O. Sener, A. R. Zamir, S. Savarese, and A. Saxena · 2015
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2015
Cited alongside, same era.
Tvsum: Summarizing web videos using titles
Y. Song, J. Vallmitjana, A. Stent, and A. Jaimes · 2015
Cited alongside, same era.
Later among the works it cites.
Jointly modeling embedding and translation to bridge video and language
Y. Pan, T. Mei, T. Yao, H. Li, and Y. Rui · 2016
Later among the works it cites.
Grounding of textual phrases in images by reconstruction
A. Rohrbach, M. Rohrbach, R. Hu, T. Darrell, and B. Schiele · 2016
Later among the works it cites.
Query-focused extractive video summarization
A. Sharghi, B. Gong, and M. Shah · 2016
Later among the works it cites.
Hollywood in homes: Crowdsourcing data collection for activity understanding
G. A. Sigurdsson, G. Varol, X. Wang, A. Farhadi, I. Laptev, and A. Gupta · 2016
Later among the works it cites.
Movieqa: Understanding stories in movies through question-answering
M. Tapaswi, Y. Zhu, R. Stiefelhagen, A. Torralba, R. Urtasun, and S. Fidler · 2016
Later among the works it cites.
Learning language-visual embedding for movie understanding with natural-language
A. Torabi, N. Tandon, and L. Sigal · 2016
Later among the works it cites.
Temporal segment networks: towards good practices for deep action recognition
L. Wang, Y. Xiong, Z. Wang, Y. Qiao, D. Lin, X. Tang, and L. Van Gool · 2016
Later among the works it cites.
Msr-vtt: A large video description dataset for bridging video and language
J. Xu, T. Mei, T. Yao, and Y. Rui · 2016
Later among the works it cites.
Highlight detection with pairwise deep ranking for first-person video summarization
T. Yao, T. Mei, and Y. Rui · 2016
Later among the works it cites.
Localizing moments in video with natural language
L. A. Hendricks, O. Wang, E. Shechtman, J. Sivic, T. Darrell, and B. Russell · 2017
Closest in time.
Modeling relationships in referential expressions with compositional modular networks
R. Hu, M. Rohrbach, J. Andreas, T. Darrell, and K. Saenko · 2017
Closest in time.
Movie description
A. Rohrbach, A. Torabi, M. Rohrbach, N. Tandon, C. Pal, H. Larochelle, A. Courville, and B. Schiele · 2017
Closest in time.