Fetching the paper…
Reading the bibliography…
Self-supervised learning has drawn attention through its effectiveness in learning in-domain representations with no ground-truth annotations; in particular, it is shown that properly designed pretext tasks (e.g., contrastive prediction task) bring significant performance gains for downstream tasks (e.g., classification task).
Learning Video Representations using Contrastive Bidirectional Transformer
Chen Sun, Fabien Baradel, Kevin Murphy, and Cordelia Schmid · 1906
Earlier work this paper cites.
Using Dynamic Time Warping to Find Patterns in Time Series
Donald J Berndt and James Clifford · 1994
Earlier work this paper cites.
Exploring Video Structure beyond The Shots
Yong Rui, Thomas S Huang, and Sharad Mehrotra · 1998
Earlier work this paper cites.
Scene Detection in Hollywood Movies and TV Shows
Zeeshan Rasheed and Mubarak Shah · 2003
Earlier work this paper cites.
Detection and Representation of Scenes in Videos
Zeeshan Rasheed and Mubarak Shah · 2005
Earlier work this paper cites.
Scene Detection in Videos using Shot Clustering and Sequence Alignment
Vasileios T Chasanis, Aristidis C Likas, and Nikolaos P Galatsanos · 2008
Earlier work this paper cites.
ImageNet: A Large-scale Hierarchical Image Database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
A Novel Role-Based Movie Scene Segmentation Method
Chao Liang, Yifan Zhang, Jian Cheng, Changsheng Xu, and Hanqing Lu · 2009
Earlier work this paper cites.
Video Scene Segmentation using A Novel Boundary Evaluation Criterion and Dynamic Programming
Bo Han and Weiguo Wu · 2011
Earlier work this paper cites.
Temporal Video Segmentation to Scenes using High-level Audiovisual Features
Panagiotis Sidiropoulos, Vasileios Mezaris, Ioannis Kompatsiaris, Hugo Meinedo, Miguel Bugalho, and Isabel Trancoso · 2011
Earlier work this paper cites.
Event perception
Barbara Tversky and Jeffrey M Zacks · 2013
Earlier work this paper cites.
Dropout: A Simple Way to Prevent Neural Networks from Overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Earlier work this paper cites.
StoryGraphs: Visualizing Character Interactions as a Timeline
Makarand Tapaswi, Martin Bauml, and Rainer Stiefelhagen · 2014
Earlier work this paper cites.
A Deep Siamese Network for Scene Detection in Broadcast Videos
Lorenzo Baraldi, Costantino Grana, and Rita Cucchiara · 2015
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization
Diederik P Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Unsupervised Learning of Video Representations using LSTMs
Nitish Srivastava, Elman Mansimov, and Ruslan Salakhudinov · 2015
Earlier work this paper cites.
Deep Residual Learning for Image Recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Gaussian Error Linear Units (GELUs)
Dan Hendrycks and Kevin Gimpel · 2016
Earlier work this paper cites.
Segmental Spatiotemporal CNNs for Fine-Grained Action Segmentation
Colin Lea, Austin Reiter, René Vidal, and Gregory D Hager · 2016
Earlier work this paper cites.
Shuffle and Learn: Unsupervised Learning using Temporal Order Verification
Ishan Misra, C Lawrence Zitnick, and Martial Hebert · 2016
Earlier work this paper cites.
Robust and efficient video scene detection using optimal sequential grouping
Daniel Rotman, Dror Porat, and Gal Ashour · 2016
Cited alongside, same era.
Generating Videos with Scene Dynamics
Carl Vondrick, Hamed Pirsiavash, and Antonio Torralba · 2016
Cited alongside, same era.
Unsupervised Representation Learning by Sorting Sequences
Hsin-Ying Lee, Jia-Bin Huang, Maneesh Singh, and Ming-Hsuan Yang · 2017
Cited alongside, same era.
Optimal Sequential Grouping for Robust Video Scene Detection using Multiple Modalities
Daniel Rotman, Dror Porat, and Gal Ashour · 2017
Cited alongside, same era.
Attention Is All You Need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Learning to Segment Actions from Observation and Narration
Daniel Fried, Jean-Baptiste Alayrac, Phil Blunsom, Chris Dyer, Stephen Clark, and Aida Nematzadeh · 2020
Later among the works it cites.
Momentum Contrast for Unsupervised Visual Representation Learning
Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick · 2020
Later among the works it cites.
MovieNet: A Holistic Dataset for Movie Understanding
Qingqiu Huang, Yu Xiong, Anyi Rao, Jiaze Wang, and Dahua Lin · 2020
Later among the works it cites.
Set-Constrained Viterbi for Set-Supervised Action Segmentation
Jun Li and Sinisa Todorovic · 2020
Later among the works it cites.
A Local-to-Global Approach to Multi-modal Movie Scene Segmentation
Anyi Rao, Linning Xu, Yu Xiong, Guodong Xu, Qingqiu Huang, Bolei Zhou, and Dahua Lin · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yang You, Igor Gitman, and Boris Ginsburg · 2017
Cited alongside, same era.
Places: A 10 million Image Database for Scene Recognition
Bolei Zhou, Agata Lapedriza, Aditya Khosla, Aude Oliva, and Antonio Torralba · 2017
Cited alongside, same era.
DiscrimNet: Semi-Supervised Action Recognition from Videos using Generative Adversarial Networks
Unaiza Ahsan, Chen Sun, and Irfan Essa · 2018
Cited alongside, same era.
Self-Supervised Spatiotemporal Feature Learning by Video Geometric Transformations
Longlong Jing and Yingli Tian · 2018
Cited alongside, same era.
A hybrid RNN-HMM Approach for Weakly Supervised Temporal Action Segmentation
Hilde Kuehne, Alexander Richard, and Juergen Gall · 2018
Cited alongside, same era.
Representation Learning with Contrastive Predictive Coding
Aaron van den Oord, Yazhe Li, and Oriol Vinyals · 2018
Cited alongside, same era.
Tracking Emerges by Colorizing Videos
Carl Vondrick, Abhinav Shrivastava, Alireza Fathi, Sergio Guadarrama, and Kevin Murphy · 2018
Cited alongside, same era.
Mengmeng Xu, Juan-Manuel Pérez-Rúa, Victor Escorcia, Brais Martinez, Xiatian Zhu, Li Zhang, Bernard Ghanem, and Tao Xiang · 2020
Later among the works it cites.
Shot Contrastive Self-Supervised Learning for Scene Boundary Detection
Shixing Chen, Xiaohan Nie, David Fan, Dongqing Zhang, Vimal Bhat, and Raffay Hamid · 2021
Later among the works it cites.
TCLR: Temporal Contrastive Learning for Video Representation
Ishan Dave, Rohit Gupta, Mamshad Nayeem Rizve, and Mubarak Shah · 2021
Later among the works it cites.
A Large-Scale Study on Unsupervised Spatiotemporal Representation Learning
Christoph Feichtenhofer, Haoqi Fan, Bo Xiong, Ross Girshick, and Kaiming He · 2021
Later among the works it cites.
Unsupervised Activity Segmentation by Joint Representation Learning and Online Clustering
Sateesh Kumar, Sanjay Haresh, Awais Ahmed, Andrey Konin, M Zeeshan Zia, and Quoc-Huy Tran · 2021
Later among the works it cites.
Action Shuffle Alternating Learning for Unsupervised Action Segmentation
Jun Li and Sinisa Todorovic · 2021
Later among the works it cites.
VALUE: A Multi-Task Benchmark for Video-and-Language Understanding Evaluation
Linjie Li, Jie Lei, Zhe Gan, Licheng Yu, Yen-Chun Chen, Rohit Pillai, Yu Cheng, Luowei Zhou, Xin Eric Wang, William Yang Wang, et al · 2021
Later among the works it cites.
Spatiotemporal Contrastive Video Representation Learning
Rui Qian, Tianjian Meng, Boqing Gong, Ming-Hsuan Yang, Huisheng Wang, Serge Belongie, and Yin Cui · 2021
Later among the works it cites.
Spatially Consistent Representation Learning
Byungseok Roh, Wuhyun Shin, Ildoo Kim, and Sungwoong Kim · 2021
Later among the works it cites.
Learning To Segment Actions From Visual and Language Instructions via Differentiable Weak Sequence Alignment
Yuhan Shen, Lu Wang, and Ehsan Elhamifar · 2021
Later among the works it cites.
Generic event boundary detection: A benchmark for event segmentation
Mike Zheng Shou, Stan W Lei, Weiyao Wang, Deepti Ghadiyaram, and Matt Feiszli · 2021
Later among the works it cites.
Fast Weakly Supervised Action Segmentation using Mutual Consistency
Yaser Souri, Mohsen Fayyaz, Luca Minciullo, Gianpiero Francesca, and Juergen Gall · 2021
Later among the works it cites.
Joint Visual-Temporal Embedding for Unsupervised Learning of Actions in Untrimmed Sequences
Rosaura G VidalMata, Walter J Scheirer, Anna Kukleva, David Cox, and Hilde Kuehne · 2021
Later among the works it cites.
Unsupervised Action Segmentation with Self-supervised Feature Learning and Co-occurrence Parsing
Zhe Wang, Hao Chen, Xinyu Li, Chunhui Liu, Yuanjun Xiong, Joseph Tighe, and Charless Fowlkes · 2021
Later among the works it cites.