Fetching the paper…
Reading the bibliography…
The ability to distinguish between different movie scenes is critical for understanding the storyline of a movie.
The film encyclopedia
Ephraim Katz · 1979
Earlier work this paper cites.
Robert sklar. film, an international history of the medium, 1993;; kristin thompson, david bordwell. film history, an introduction, 1994
Paul-Louis Thirard and Lorenzo Codelli · 1994
Earlier work this paper cites.
Constructing table-of-content for videos
Yong Rui, Thomas S Huang, and Sharad Mehrotra · 1999
Earlier work this paper cites.
Constructing table-of-content for videos
Xiang Sean Zhou, Yong Rui, and Thomas S Huang · 2003
Earlier work this paper cites.
Detection and representation of scenes in videos
Zeeshan Rasheed and Mubarak Shah · 2005
Earlier work this paper cites.
Video shot detection and condensed representation. a review
Costas Cotsaces, Nikos Nikolaidis, and Ioannis Pitas · 2006
Earlier work this paper cites.
Scene detection in videos using shot clustering and sequence alignment
Vasileios T Chasanis, Aristidis C Likas, and Nikolaos P Galatsanos · 2008
Earlier work this paper cites.
Video scene segmentation using a novel boundary evaluation criterion and dynamic programming
Bo Han and Weiguo Wu · 2011
Earlier work this paper cites.
Temporal video segmentation to scenes using high-level audiovisual features
Panagiotis Sidiropoulos, Vasileios Mezaris, Ioannis Kompatsiaris, Hugo Meinedo, Miguel Bugalho, and Isabel Trancoso · 2011
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, et al · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
The language of actions: Recovering the syntax and semantics of goal-directed human activities
Hilde Kuehne, Ali Arslan, and Thomas Serre · 2014
Earlier work this paper cites.
Storygraphs: visualizing character interactions as a timeline
Makarand Tapaswi, Martin Bauml, and Rainer Stiefelhagen · 2014
Earlier work this paper cites.
Analysis and re-use of videos in educational digital libraries with automatic scene detection
Lorenzo Baraldi, Costantino Grana, and Rita Cucchiara · 2015
Earlier work this paper cites.
A deep siamese network for scene detection in broadcast videos
Lorenzo Baraldi, Costantino Grana, and Rita Cucchiara · 2015
Earlier work this paper cites.
Shot and scene detection via hierarchical clustering for re-using broadcast video
Lorenzo Baraldi, Costantino Grana, and Rita Cucchiara · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Bridging nonlinearities and stochastic regularizers with gaussian error linear units
Dan Hendrycks and Kevin Gimpel · 2016
Cited alongside, same era.
Robust and efficient video scene detection using optimal sequential grouping
Daniel Rotman, Dror Porat, and Gal Ashour · 2016
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Pyscenedetect: Intelligent scene cut detection and video splitting tool, 2018
Brandon Castellano · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Cited alongside, same era.
A local-to-global approach to multi-modal movie scene segmentation
Anyi Rao, Linning Xu, Yu Xiong, Guodong Xu, Qingqiu Huang, Bolei Zhou, and Dahua Lin · 2020
Later among the works it cites.
Comprehensive instructional video analysis: The coin dataset and performance evaluation
Yansong Tang, Jiwen Lu, and Jie Zhou · 2020
Later among the works it cites.
Big bird: Transformers for longer sequences
Manzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie, Chris Alberti, Santiago Ontanon, Philip Pham, Anirudh Ravula, Qifan Wang, Li Yang, et al · 2020
Later among the works it cites.
Is space-time attention all you need for video understanding?
Gedas Bertasius, Heng Wang, and Lorenzo Torresani · 2021
Later among the works it cites.
Shot contrastive self-supervised learning for scene boundary detection
Shixing Chen, Xiaohan Nie, David Fan, Dongqing Zhang, Vimal Bhat, and Raffay Hamid · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc V Le, and Ruslan Salakhutdinov · 2019
Cited alongside, same era.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu · 2019
Cited alongside, same era.
Coin: A large-scale dataset for comprehensive instructional video analysis
Yansong Tang, Dajun Ding, Yongming Rao, Yu Zheng, Danyang Zhang, Lili Zhao, Jiwen Lu, and Jie Zhou · 2019
Cited alongside, same era.
Xlnet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R Salakhutdinov, and Quoc V Le · 2019
Cited alongside, same era.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Cited alongside, same era.
End-to-end object detection with transformers
Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko · 2020
Cited alongside, same era.
Albert Gu, Karan Goel, and Christopher Ré · 2021
Later among the works it cites.
Combining recurrent, convolutional, and continuous-time models with linear state space layers
Albert Gu, Isys Johnson, Karan Goel, Khaled Kamal Saab, Tri Dao, Atri Rudra, and Christopher Ré · 2021
Later among the works it cites.
Keeping your eye on the ball: Trajectory attention in video transformers
Mandela Patrick, Dylan Campbell, Yuki Asano, Ishan Misra, Florian Metze, Christoph Feichtenhofer, Andrea Vedaldi, and João F Henriques · 2021
Later among the works it cites.
Towards long-form video understanding
Chao-Yuan Wu and Philipp Krahenbuhl · 2021
Later among the works it cites.
Graph-based high-order relation modeling for long-term action recognition
Jiaming Zhou, Kun-Yu Lin, Haoxin Li, and Wei-Shi Zheng · 2021
Later among the works it cites.
It’s raw! audio generation with state-space models
Karan Goel, Albert Gu, Chris Donahue, and Christopher Ré · 2022
Closest in time.
Diagonal state spaces are as effective as structured state spaces
Ankit Gupta · 2022
Closest in time.
Transformer quality in linear time
Weizhe Hua, Zihang Dai, Hanxiao Liu, and Quoc Le · 2022
Closest in time.
Learning to recognize procedural activities with distant supervision
Xudong Lin, Fabio Petroni, Gedas Bertasius, Marcus Rohrbach, Shih-Fu Chang, and Lorenzo Torresani · 2022
Closest in time.
Long range language modeling via gated state spaces
Harsh Mehta, Ankit Gupta, Ashok Cutkosky, and Behnam Neyshabur · 2022
Closest in time.
Long movie clip classification with state-space video models
Md Mohaiminul Islam and Gedas Bertasius · 2022
Closest in time.
Boundary-aware self-supervised learning for video scene segmentation
Jonghwan Mun, Minchul Shin, Gunsoo Han, Sangho Lee, Seongsu Ha, Joonseok Lee, and Eun-Sol Kim · 2022
Closest in time.