Fetching the paper…
Reading the bibliography…
Books are a rich source of both fine-grained information, how a character, an object or a scene looks like, as well as high-level semantics, what someone is thinking, feeling and how these states evolve through a story.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
BLEU: a method for automatic evaluation of machine translation
K. Papineni, S. Roukos, T. Ward, and W. J. Zhu · 2002
Earlier work this paper cites.
“Hello! My name is… Buffy” – Automatic Naming of Characters in TV Video
M. Everingham, J. Sivic, and A. Zisserman · 2006
Earlier work this paper cites.
Movie/script: Alignment and parsing of video and text transcription
T. Cour, C. Jordan, E. Miltsakaki, and B. Taskar · 2008
Earlier work this paper cites.
Beyond nouns: Exploiting prepositions and comparative adjectives for learning visual classifiers
A. Gupta and L. Davis · 2008
Earlier work this paper cites.
Subtitle-free Movie to Script Alignment
P. Sankar, C. V. Jawahar, and A. Zisserman · 2009
Earlier work this paper cites.
“Who are you?” - Learning person specific classifiers from video
J. Sivic, M. Everingham, and A. Zisserman · 2009
Earlier work this paper cites.
Every picture tells a story: Generating sentences for images
A. Farhadi, M. Hejrati, M. Sadeghi, P. Young, C. Rashtchian, J. Hockenmaier, and D. Forsyth · 2010
Earlier work this paper cites.
Baby talk: Understanding and generating simple image descriptions
G. Kulkarni, V. Premraj, S. Dhar, S. Li, Y. Choi, A. Berg, and T. Berg · 2011
Earlier work this paper cites.
Efficient Structured Prediction with Latent Variables for General Graphical Models
A. Schwing, T. Hazan, M. Pollefeys, and R. Urtasun · 2012
Earlier work this paper cites.
A sentence is worth a thousand pixels
S. Fidler, A. Sharma, and R. Urtasun · 2013
Earlier work this paper cites.
Recurrent continuous translation models
N. Kalchbrenner and P. Blunsom · 2013
Earlier work this paper cites.
Efficient estimation of word representations in vector space
T. Mikolov, K. Chen, G. Corrado, and J. Dean · 2013
Earlier work this paper cites.
Video event understanding using natural language descriptions
V. Ramanathan, P. Liang, and L. Fei-Fei · 2013
Cited alongside, same era.
Learning phrase representations using rnn encoder-decoder for statistical machine translation
K. Cho, B. van Merrienboer, C. Gulcehre, F. Bougares, H. Schwenk, and Y. Bengio · 2014
Cited alongside, same era.
Empirical evaluation of gated recurrent neural networks on sequence modeling
J. Chung, C. Gulcehre, K. Cho, and Y. Bengio · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
D. Kingma and J. Ba · 2014
Cited alongside, same era.
Unifying visual-semantic embeddings with multimodal neural language models
R. Kiros, R. Salakhutdinov, and R. S. Zemel · 2014
Cited alongside, same era.
Going deeper with convolutions
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich · 2014
Later among the works it cites.
Translating Videos to Natural Language Using Deep Recurrent Neural Networks
S. Venugopalan, H. Xu, J. Donahue, M. Rohrbach, R. J. Mooney, and K. Saenko · 2014
Later among the works it cites.
Show and tell: A neural image caption generator
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan · 2014
Later among the works it cites.
Learning Deep Features for Scene Recognition using Places Database
B. Zhou, A. Lapedriza, J. Xiao, A. Torralba, and A. Oliva · 2014
Later among the works it cites.
Neural machine translation by jointly learning to align and translate
D. Bahdanau, K. Cho, and Y. Bengio · 2015
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
What are you talking about? text-to-image coreference
C. Kong, D. Lin, M. Bansal, R. Urtasun, and S. Fidler · 2014
Cited alongside, same era.
Visual Semantic Search: Retrieving Videos via Complex Textual Queries
D. Lin, S. Fidler, C. Kong, and R. Urtasun · 2014
Cited alongside, same era.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Cited alongside, same era.
A multi-world approach to question answering about real-world scenes based on uncertain input
M. Malinowski and M. Fritz · 2014
Cited alongside, same era.
Explain images with multimodal recurrent neural networks
J. Mao, W. Xu, Y. Yang, J. Wang, and A. L. Yuille · 2014
Cited alongside, same era.
Inferring the why in images
H. Pirsiavash, C. Vondrick, and A. Torralba · 2014
Cited alongside, same era.
Linking People in Videos with “Their” Names Using Coreference Resolution
V. Ramanathan, A. Joulin, P. Liang, and L. Fei-Fei · 2014
Cited alongside, same era.
Closest in time.
Deep visual-semantic alignments for generating image descriptions
A. Karpathy and L. Fei-Fei · 2015
Closest in time.
Skip-Thought Vectors
R. Kiros, Y. Zhu, R. Salakhutdinov, R. S. Zemel, A. Torralba, R. Urtasun, and S. Fidler · 2015
Closest in time.
Don’t just listen, use your imagination: Leveraging visual common sense for non-visual tasks
X. Lin and D. Parikh · 2015
Closest in time.
A dataset for movie description
A. Rohrbach, M. Rohrbach, N. Tandon, and B. Schiele · 2015
Closest in time.
Book2Movie: Aligning Video scenes with Book chapters
M. Tapaswi, M. Bauml, and R. Stiefelhagen · 2015
Closest in time.
Aligning Plot Synopses to Videos for Story-based Retrieval
M. Tapaswi, M. Bäuml, and R. Stiefelhagen · 2015
Closest in time.
Show, attend and tell: Neural image caption generation with visual attention
K. Xu, J. Ba, R. Kiros, K. Cho, A. Courville, R. Salakhutdinov, R. Zemel, and Y. Bengio · 2015
Closest in time.