Fetching the paper…
Reading the bibliography…
Learning how to generate descriptions of images or videos received major interest both in the Computer Vision and Natural Language Processing communities.
An iterative image registration technique with an application to stereo vision
B. D. Lucas and T. Kanade · 1981
Earlier work this paper cites.
A cutting plane algorithm for a clustering problem
M. Grötschel and Y. Wakabayashi · 1989
Earlier work this paper cites.
Detection and tracking of feature points
C. Tomasi and T. Kanade · 1991
Earlier work this paper cites.
The partition problem
S. Chopra and M. Rao · 1993
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Incorporating non-local information into information extraction systems by gibbs sampling
J. R. Finkel, T. Grenager, and C. Manning · 2005
Earlier work this paper cites.
Bootstrapping path-based pronoun resolution
S. Bergsma and D. Lin · 2006
Earlier work this paper cites.
”hello! my name is… buffy” - automatic naming of characters in tv video
M. Everingham, J. Sivic, and A. Zisserman · 2006
Earlier work this paper cites.
Learning from ambiguously labeled images
T. Cour, B. Sapp, C. Jordan, and B. Taskar · 2009
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
”who are you?”-learning person specific classifiers from video
J. Sivic, M. Everingham, and A. Zisserman · 2009
Earlier work this paper cites.
”knock! knock! who is it?” probabilistic person identification in tv-series
M. Tapaswi, M. Baeuml, and R. Stiefelhagen · 2012
Earlier work this paper cites.
Finding actors and actions in movies
P. Bojanowski, F. Bach, I. Laptev, J. Ponce, C. Schmid, and J. Sivic · 2013
Earlier work this paper cites.
Youtube2text: Recognizing and describing arbitrary activities using semantic hierarchies and zero-shoot recognition
S. Guadarrama, N. Krishnamoorthy, G. Malkarnenkar, S. Venugopalan, R. Mooney, T. Darrell, and K. Saenko · 2013
Earlier work this paper cites.
Translating video content to natural language descriptions
M. Rohrbach, W. Qiu, I. Titov, S. Thater, M. Pinkal, and B. Schiele · 2013
Earlier work this paper cites.
Action recognition with improved trajectories
H. Wang and C. Schmid · 2013
Earlier work this paper cites.
Grounded language learning from videos described with sentences
H. Yu and J. M. Siskind · 2013
Earlier work this paper cites.
LSDA: Large scale detection through adaptation
J. Hoffman, S. Guadarrama, E. Tzeng, J. Donahue, R. Girshick, T. Darrell, and K. Saenko · 2014
Earlier work this paper cites.
What are you talking about? text-to-image coreference
C. Kong, D. Lin, M. Bansal, R. Urtasun, and S. Fidler · 2014
Earlier work this paper cites.
Deepreid: Deep filter pairing neural network for person re-identification
W. Li, R. Zhao, T. Xiao, and X. Wang · 2014
Earlier work this paper cites.
Visual semantic search: Retrieving videos via complex textual queries
D. Lin, S. Fidler, C. Kong, and R. Urtasun · 2014
Earlier work this paper cites.
Fine-grained activity recognition with holistic and pose based features
L. Pishchulin, M. Andriluka, and B. Schiele · 2014
Earlier work this paper cites.
Linking people in videos with ”their” names using coreference resolution
V. Ramanathan, A. Joulin, P. Liang, and L. Fei-Fei · 2014
Cited alongside, same era.
Coherent multi-sentence video description with variable level of detail
A. Rohrbach, M. Rohrbach, W. Qiu, A. Friedrich, M. Pinkal, and B. Schiele · 2014
Cited alongside, same era.
Deepface: Closing the gap to human-level performance in face verification
Y. Taigman, M. Yang, M. Ranzato, and L. Wolf · 2014
Cited alongside, same era.
Integrating language and vision to generate natural language descriptions of videos in the wild
J. Thomason, S. Venugopalan, S. Guadarrama, K. Saenko, and R. J. Mooney · 2014
Cited alongside, same era.
W. Zaremba and I. Sutskever · 2014
Cited alongside, same era.
Subgraph decomposition for multi-target tracking
S. Tang, B. Andres, M. Andriluka, and B. Schiele · 2015
Later among the works it cites.
Using descriptive video services to create a large data source for video annotation research
A. Torabi, C. Pal, H. Larochelle, and A. Courville · 2015
Later among the works it cites.
Sequence to sequence – video to text
S. Venugopalan, M. Rohrbach, J. Donahue, R. Mooney, T. Darrell, and K. Saenko · 2015
Later among the works it cites.
Show and tell: A neural image caption generator
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan · 2015
Later among the works it cites.
Show, attend and tell: Neural image caption generation with visual attention
K. Xu, J. Ba, R. Kiros, A. Courville, R. Salakhutdinov, R. Zemel, and Y. Bengio · 2015
Later among the works it cites.
Describing videos by exploiting temporal structure
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
B. Zhou, A. Lapedriza, J. Xiao, A. Torralba, and A. Oliva · 2014
Cited alongside, same era.
Neural machine translation by jointly learning to align and translate
D. Bahdanau, K. Cho, and Y. Bengio · 2015
Cited alongside, same era.
Long-term recurrent convolutional networks for visual recognition and description
J. Donahue, L. A. Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell · 2015
Cited alongside, same era.
Fast r-cnn
R. Girshick · 2015
Cited alongside, same era.
Contextual action recognition with r* cnn
G. Gkioxari, R. Girshick, and J. Malik · 2015
Cited alongside, same era.
Deep visual-semantic alignments for generating image descriptions
A. Karpathy and L. Fei-Fei · 2015
Cited alongside, same era.
Generating multi-sentence natural language descriptions of indoor scenes
D. Lin, S. Fidler, C. Kong, and R. Urtasun · 2015
Cited alongside, same era.
L. Yao, A. Torabi, K. Cho, N. Ballas, C. Pal, H. Larochelle, and A. Courville · 2015
Later among the works it cites.
Naive-deep face recognition: Touching the limit of lfw benchmark or not?
E. Zhou, Z. Cao, and Q. Yin · 2015
Later among the works it cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Later among the works it cites.
Natural language object retrieval
R. Hu, H. Xu, M. Rohrbach, J. Feng, K. Saenko, and T. Darrell · 2016
Later among the works it cites.
Visual storytelling
T.-H. Huang, F. Ferraro, N. Mostafazadeh, I. Misra, A. Agrawal, J. Devlin, R. Girshick, X. He, P. Kohli, D. Batra, C. L. Zitnick, D. Parikh, L. Vanderwende, M. Galley, and M. Mitchell · 2016
Later among the works it cites.
Densecap: Fully convolutional localization networks for dense captioning
J. Johnson, A. Karpathy, and L. Fei-Fei · 2016
Later among the works it cites.
Generation and comprehension of unambiguous object descriptions
J. Mao, J. Huang, A. Toshev, O. Camburu, A. Yuille, and K. Murphy · 2016
Later among the works it cites.
Grounding of textual phrases in images by reconstruction
A. Rohrbach, M. Rohrbach, R. Hu, T. Darrell, and B. Schiele · 2016
Later among the works it cites.
Beyond caption to narrative: Video captioning with multiple sentences
A. Shin, K. Ohnishi, and T. Harada · 2016
Later among the works it cites.
Learning deep structure-preserving image-text embeddings
L. Wang, Y. Li, and S. Lazebnik · 2016
Later among the works it cites.
Image captioning with semantic attention
Q. You, H. Jin, Z. Wang, C. Fang, and J. Luo · 2016
Later among the works it cites.
Video paragraph captioning using hierarchical recurrent neural networks
H. Yu, J. Wang, Z. Huang, Y. Yang, and W. Xu · 2016
Later among the works it cites.
Modeling context in referring expressions
L. Yu, P. Poirson, S. Yang, A. C. Berg, and T. L. Berg · 2016
Later among the works it cites.
Spatio-temporal attention models for grounded video captioning
M. Zanfir, E. Marinoiu, and C. Sminchisescu · 2016
Later among the works it cites.
Attention correctness in neural image captioning
C. Liu, J. Mao, F. Sha, and A. Yuille · 2017
Closest in time.
Movie description
A. Rohrbach, A. Torabi, M. Rohrbach, N. Tandon, C. Pal, H. Larochelle, A. Courville, and B. Schiele · 2017
Closest in time.