Fetching the paper…
Reading the bibliography…
Our objective in this work is long range understanding of the narrative structure of movies.
Dynamic programming algorithm optimization for spoken word recognition
Sakoe, H., Chiba, S.: · 1978
Earlier work this paper cites.
Algorithms for clustering data
Jain, A.K., Dubes, R.C.: · 1988
Earlier work this paper cites.
“Hello! My name is… Buffy” – automatic naming of characters in TV video
Everingham, M., Sivic, J., Zisserman, A.: · 2006
Earlier work this paper cites.
Learning realistic human actions from movies
Laptev, I., Marszałek, M., Schmid, C., Rozenfeld, B.: · 2008
Earlier work this paper cites.
Learning from ambiguously labeled images
Cour, T., Sapp, B., Jordan, C., Taskar, B.: · 2009
Earlier work this paper cites.
“Who are you?” – learning person specific classifiers from video
Sivic, J., Everingham, M., Zisserman, A.: · 2009
Earlier work this paper cites.
Automatic annotation of human actions in video
Duchenne, O., Laptev, I., Sivic, J., Bach, F., Ponce, J.: · 2009
Earlier work this paper cites.
Actions in context
Marszałek, M., Laptev, I., Schmid, C.: · 2009
Earlier work this paper cites.
MTLD, vocd-D, and HD-D: A validation study of sophisticated approaches to lexical diversity assessment
McCarthy, P.M., Jarvis, S.: · 2010
Earlier work this paper cites.
A similarity measure of Jumping Dynamic Time Warping
Feng, L., Zhao, X., Liu, Y., Yao, Y., Jin, B.: · 2010
Earlier work this paper cites.
“Knock! Knock! Who is it?” probabilistic person identification in TV-series
Tapaswi, M., Bäuml, M., Stiefelhagen, R.: · 2012
Earlier work this paper cites.
StoViz: story visualization of TV series
Ercolessi, P., Bredin, H., Sénac, C.: · 2012
Earlier work this paper cites.
Finding actors and actions in movies
Bojanowski, P., Bach, F., Laptev, I., Ponce, J., Schmid, C., Sivic, J.: · 2013
Earlier work this paper cites.
Storygraphs: visualizing character interactions as a timeline
Tapaswi, M., Bauml, M., Stiefelhagen, R.: · 2014
Earlier work this paper cites.
Grounded compositional semantics for finding and describing images with sentences
Socher, R., Karpathy, A., Le, Q.V., Manning, C.D., Ng, A.Y.: · 2014
Earlier work this paper cites.
Densenet: Implementing efficient convnet descriptor pyramids
Iandola, F., Moskewicz, M., Karayev, S., Girshick, R., Darrell, T., Keutzer, K.: · 2014
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization
Kingma, D.P., Ba, J.: · 2014
Earlier work this paper cites.
Unsupervised semantic parsing of video collections
Sener, O., Zamir, A.R., Savarese, S., Saxena, A.: · 2015
Earlier work this paper cites.
Aligning plot synopses to videos for story-based retrieval
Tapaswi, M., Bäuml, M., Stiefelhagen, R.: · 2015
Earlier work this paper cites.
ActivityNet: A large-scale video benchmark for human activity understanding
Caba Heilbron, F., Escorcia, V., Ghanem, B., Carlos Niebles, J.: · 2015
Earlier work this paper cites.
Book2Movie: Aligning Video scenes with Book chapters
Tapaswi, M., Bäuml, M., Stiefelhagen, R.: · 2015
Earlier work this paper cites.
Weakly-Supervised Alignment of Video With Text
Bojanowski, P., Lajugie, R., Grave, E., Bach, F., Laptev, I., Ponce, J., Schmid, C.: · 2015
Earlier work this paper cites.
{MSR-VTT:} A large video description dataset for bridging video and language
Xu, J., Mei, T., Yao, T., Rui, Y.: · 2016
Cited alongside, same era.
MovieQA: Understanding Stories in Movies through Question-Answering
Tapaswi, M., Zhu, Y., Stiefelhagen, R., Torralba, A., Urtasun, R., Fidler, S.: · 2016
Cited alongside, same era.
Unsupervised learning from narrated instruction videos
Alayrac, J.B., Bojanowski, P., Agrawal, N., Sivic, J., Laptev, I., Lacoste-Julien, S.: · 2016
Cited alongside, same era.
Aligning movies with scripts by exploiting temporal ordering constraints
Naim, I., Al Mamun, A., Song, Y.C., Luo, J., Kautz, H., Gildea, D.: · 2016
Cited alongside, same era.
NetVLAD: CNN architecture for weakly supervised place recognition
Arandjelovic, R., Gronat, P., Torii, A., Pajdla, T., Sivic, J.: · 2016
Cited alongside, same era.
WIDER FACE: A Face Detection Benchmark
Yang, S., Luo, P., Loy, C.C., Tang, X.: · 2016
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.W., Lee, K., Toutanova, K.: · 2018
Later among the works it cites.
Squeeze-and-excitation networks
Hu, J., Shen, L., Sun, G.: · 2018
Later among the works it cites.
A Graph-Based Framework to Bridge Movies and Synopses
Xiong, Y., Huang, Q., Guo, L., Zhou, H., Zhou, B., Lin, D.: · 2019
Later among the works it cites.
HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips
Miech, A., Zhukov, D., Alayrac, J.B., Tapaswi, M., Laptev, I., Sivic, J.: · 2019
Later among the works it cites.
Coin: A large-scale dataset for comprehensive instructional video analysis
Tang, Y., Ding, D., Rao, Y., Zheng, Y., Zhang, D., Zhao, L., Lu, J., Zhou, J.: · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Movie description
Rohrbach, A., Torabi, A., Rohrbach, M., Tandon, N., Pal, C., Larochelle, H., Courville, A., Schiele, B.: · 2017
Cited alongside, same era.
Localizing moments in video with natural language
Hendricks, L.A., Wang, O., Shechtman, E., Sivic, J., Darrell, T., Russell, B.: · 2017
Cited alongside, same era.
The {Kinetics} Human Action Video Dataset
Kay, W., Carreira, J., Simonyan, K., Zhang, B., Hillier, C., Vijayanarasimhan, S., Viola, F., Green, T., Back, T., Natsev, P., Suleyman, M., Zisserman, A.: · 2017
Cited alongside, same era.
From benedict cumberbatch to sherlock holmes: Character identification in tv series without a script
Nagrani, A., Zisserman, A.: · 2017
Cited alongside, same era.
Dense-captioning events in videos
Krishna, R., Hata, K., Ren, F., Fei-Fei, L., Carlos Niebles, J.: · 2017
Cited alongside, same era.
Movie Description
Rohrbach, A., Torabi, A., Rohrbach, M., Tandon, N., Pal, C., Larochelle, H., Courville, A., Schiele, B.: · 2017
Cited alongside, same era.
Ignat, O., Burdick, L., Deng, J., Mihalcea, R.: · 2019
Later among the works it cites.
Moments in time dataset: one million videos for event understanding
Monfort, M., Andonian, A., Zhou, B., Ramakrishnan, K., Bargal, S.A., Yan, Y., Brown, L., Fan, Q., Gutfreund, D., Vondrick, C., Others: · 2019
Later among the works it cites.
Videobert: A joint model for video and language representation learning
Sun, C., Myers, A., Vondrick, C., Murphy, K., Schmid, C.: · 2019
Later among the works it cites.
Are we asking the right questions in {MovieQA?}
Jasani, B., Girdhar, R., Ramanan, D.: · 2019
Later among the works it cites.
TVQA+: Spatio-Temporal Grounding for Video Question Answering
Lei, J., Yu, L., Berg, T.L., Bansal, M.: · 2019
Later among the works it cites.
Use what you have: Video retrieval using representations from collaborative experts
Liu, Y., Albanie, S., Nagrani, A., Zisserman, A.: · 2019
Later among the works it cites.
Long-Term Feature Banks for Detailed Video Understanding
Wu, C.Y., Feichtenhofer, C., Fan, H., He, K., Krahenbuhl, P., Girshick, R.: · 2019
Later among the works it cites.
Voxceleb: Large-scale speaker verification in the wild
Nagrani, A., Chung, J.S., Xie, W., Zisserman, A.: · 2019
Later among the works it cites.
Squeeze-and-Excitation Networks
Hu, J., Shen, L., Albanie, S., Sun, G., Wu, E.: · 2019
Later among the works it cites.
DSFD: dual shot face detector
Li, J., Wang, Y., Wang, C., Tai, Y., Qian, J., Yang, J., Wang, C., Li, J., Huang, F.: · 2019
Later among the works it cites.
MovieNet: A Holistic Dataset for Movie Understanding
Huang, Q., Xiong, Y., Rao, A., Wang, J., Lin, D.: · 2020
Closest in time.
Caption-Supervised Face Recognition: Training a State-of-the-Art Face Model without Manual Annotation
Huang, Q., Yang, L., Huang, H., Wu, T., Lin, D.: · 2020
Closest in time.
Speech2Action: Cross-Modal Supervision for Action Recognition
Nagrani, A., Sun, C., Ross, D., Sukthankar, R., Schmid, C., Zisserman, A.: · 2020
Closest in time.
A Local-to-Global Approach to Multi-modal Movie Scene Segmentation
Rao, A., Xu, L., Xiong, Y., Xu, G., Huang, Q., Zhou, B., Lin, D.: · 2020
Closest in time.
Learning Interactions and Relationships between Movie Characters
Kukleva, A., Tapaswi, M., Laptev, I.: · 2020
Closest in time.
End-to-End Learning of Visual Representations From Uncurated Instructional Videos
Miech, A., Alayrac, J.B., Smaira, L., Laptev, I., Sivic, J., Zisserman, A.: · 2020
Closest in time.