Fetching the paper…
Reading the bibliography…
We introduce the MovieQA dataset which aims to evaluate automatic story comprehension from both video and text.
Movie/Script: Alignment and Parsing of Video and Text Transcription
T. Cour, C. Jordan, E. Miltsakaki, and B. Taskar · 2008
Earlier work this paper cites.
Subtitle-free Movie to Script Alignment
P. Sankar, C. V. Jawahar, and A. Zisserman · 2009
Earlier work this paper cites.
“Who are you?” - Learning person specific classifiers from video
J. Sivic, M. Everingham, and A. Zisserman · 2009
Earlier work this paper cites.
Every Picture Tells a Story: Generating Sentences for Images
A. Farhadi, M. Hejrati, M. Sadeghi, P. Young, C. Rashtchian, J. Hockenmaier, and D. Forsyth · 2010
Earlier work this paper cites.
Collecting highly parallel data for paraphrase evaluation
D. L. Chen and W. B. Dolan · 2011
Earlier work this paper cites.
Baby Talk: Understanding and Generating Simple Image Descriptions
G. Kulkarni, V. Premraj, S. Dhar, S. Li, Y. Choi, A. Berg, and T. Berg · 2011
Earlier work this paper cites.
Im2Text: Describing Images Using 1 Million Captioned Photographs
V. Ordonez, G. Kulkarni, and T. Berg · 2011
Earlier work this paper cites.
Corpus-guided Sentence Generation of Natural Images
Y. Yang, C. L. Teo, H. Daumé, III, and Y. Aloimonos · 2011
Earlier work this paper cites.
Video-In-sentences Out
A. Barbu, A. Bridge, Z. Burchill, D. Coroian, S. Dickinson, S. Fidler, A. Michaux, S. Mussman, S. Narayanaswamy, D. Salvi, L. Schmidt, J. Shangguan, J. Siskind, J. Waggoner, S. Wang, J. Wei, Y. Yin, and Z. Zhang · 2012
Earlier work this paper cites.
Semi-supervised Learning with Constraints for Person Identification in Multimedia Data
M. Baeuml, M. Tapaswi, and R. Stiefelhagen · 2013
Earlier work this paper cites.
Finding Actors and Actions in Movies
P. Bojanowski, F. Bach, I. Laptev, J. Ponce, C. Schmid, and J. Sivic · 2013
Earlier work this paper cites.
A Thousand Frames in Just a Few Words: Lingual Description of Videos through Latent Topics and Sparse Object Stitching
P. Das, C. Xu, R. F. Doell, and J. J. Corso · 2013
Earlier work this paper cites.
A thousand frames in just a few words: Lingual description of videos through latent topics and sparse object stitching
P. Das, C. Xu, R. F. Doell, and J. J. Corso · 2013
Earlier work this paper cites.
Generating Natural-Language Video Descriptions Using Text-Mined Knowledge
N. Krishnamoorthy, G. Malkarnenkar, R. J. Mooney, K. Saenko, and S. Guadarrama · 2013
Earlier work this paper cites.
Learning dependency-based compositional semantics
P. Liang, M. Jordan, and D. Klein · 2013
Earlier work this paper cites.
Efficient estimation of word representations in vector space
T. Mikolov, K. Chen, G. Corrado, and J. Dean · 2013
Earlier work this paper cites.
Video Event Understanding using Natural Language Descriptions
V. Ramanathan, P. Liang, and L. Fei-Fei · 2013
Earlier work this paper cites.
Mctest: A challenge dataset for the open-domain machine comprehension of text
M. Richardson, C. J. Burges, and E. Renshaw · 2013
Cited alongside, same era.
Translating Video Content to Natural Language Descriptions
M. Rohrbach, W. Qiu, I. Titov, S. Thater, M. Pinkal, and B. Schiele · 2013
Cited alongside, same era.
Learning a Recurrent Visual Representation for Image Caption Generation
X. Chen and C. L. Zitnick · 2014
Cited alongside, same era.
Long-term Recurrent Convolutional Networks for Visual Recognition and Description
J. Donahue, L. A. Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
D. Kingma and J. Ba · 2014
Cited alongside, same era.
Adopting abstract images for semantic scene understanding
C. Zitnick, R. Vedantam, and D. Parikh · 2014
Later among the works it cites.
VQA: Visual Question Answering
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh · 2015
Closest in time.
Teaching Machines to Read and Comprehend
K. M. Hermann, T. Kočisky, E. Grefenstette, L. Espeholt, W. Kay, M. Suleyman, and P. Blunsom · 2015
Closest in time.
Deep Visual-Semantic Alignments for Generating Image Descriptions
A. Karpathy and L. Fei-Fei · 2015
Closest in time.
Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models
R. Kiros, R. Salakhutdinov, and R. S. Zemel · 2015
Closest in time.
Skip-Thought Vectors
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
What are you talking about? Text-to-Image Coreference
C. Kong, D. Lin, M. Bansal, R. Urtasun, and S. Fidler · 2014
Cited alongside, same era.
Visual Semantic Search: Retrieving Videos via Complex Textual Queries
D. Lin, S. Fidler, C. Kong, and R. Urtasun · 2014
Cited alongside, same era.
Microsoft COCO: Common Objects in Context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Cited alongside, same era.
A Multi-World Approach to Question Answering about Real-World Scenes based on Uncertain Input
M. Malinowski and M. Fritz · 2014
Cited alongside, same era.
Inferring the Why in Images
H. Pirsiavash, C. Vondrick, and A. Torralba · 2014
Cited alongside, same era.
Linking People in Videos with “Their” Names Using Coreference Resolution
V. Ramanathan, A. Joulin, P. Liang, and L. Fei-Fei · 2014
Cited alongside, same era.
Going deeper with convolutions
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich · 2014
Cited alongside, same era.
R. Kiros, Y. Zhu, R. Salakhutdinov, R. Zemel, A. Torralba, R. Urtasun, and S. Fidler · 2015
Closest in time.
Ask Your Neurons: A Neural-based Approach to Answering Questions about Images
M. Malinowski, M. Rohrbach, and M. Fritz · 2015
Closest in time.
Exploring Models and Data for Image Question Answering
M. Ren, R. Kiros, and R. Zemel · 2015
Closest in time.
A Dataset for Movie Description
A. Rohrbach, M. Rohrbach, N. Tandon, and B. Schiele · 2015
Closest in time.
S. Sukhbaatar, A. Szlam, J. Weston, and R. Fergus · 2015
Closest in time.
Book2Movie: Aligning Video scenes with Book chapters
M. Tapaswi, M. Bauml, and R. Stiefelhagen · 2015
Closest in time.
Aligning Plot Synopses to Videos for Story-based Retrieval
M. Tapaswi, M. Bäuml, and R. Stiefelhagen · 2015
Closest in time.
Learning Common Sense Through Visual Abstraction
R. Vedantam, X. Lin, T. Batra, C. L. Zitnick, and D. Parikh · 2015
Closest in time.
Machine Comprehension with Syntax, Frames, and Semantics
H. Wang, M. Bansal, K. Gimpel, and D. McAllester · 2015
Closest in time.
Visual Madlibs: Fill in the blank Image Generation and Question Answering
L. Yu, E. Park, A. C. Berg, and T. L. Berg · 2015
Closest in time.
Aligning Books and Movies: Towards Story-like Visual Explanations by Watching Movies and Reading Books
Y. Zhu, R. Kiros, R. Zemel, R. Salakhutdinov, R. Urtasun, A. Torralba, and S. Fidler · 2015
Closest in time.