Fetching the paper…
Reading the bibliography…
We present a novel approach for unsupervised activity segmentation which uses video frame clustering as a pretext task and simultaneously performs representation learning and online clustering.
Autoencoders, minimum description length and helmholtz free energy
Geoffrey E Hinton and Richard S Zemel · 1994
Earlier work this paper cites.
Dimensionality reduction by learning an invariant mapping
Raia Hadsell, Sumit Chopra, and Yann LeCun · 2006
Earlier work this paper cites.
Extracting and composing robust features with denoising autoencoders
Pascal Vincent, Hugo Larochelle, Yoshua Bengio, and Pierre-Antoine Manzagol · 2008
Earlier work this paper cites.
Slow, decorrelated features for pretraining complex cell-like networks
Yoshua Bengio and James S Bergstra · 2009
Earlier work this paper cites.
Deep learning from temporal coherence in video
Hossein Mobahi, Ronan Collobert, and Jason Weston · 2009
Earlier work this paper cites.
Unsupervised learning of visual invariance with temporal coherence
Will Y Zou, Andrew Y Ng, and Kai Yu · 2011
Earlier work this paper cites.
Deep learning of invariant features via simulated fixations in video
Will Zou, Shenghuo Zhu, Kai Yu, and Andrew Y Ng · 2012
Earlier work this paper cites.
Sinkhorn distances: Lightspeed computation of optimal transport
Marco Cuturi · 2013
Earlier work this paper cites.
Combining embedded accelerometers with computer vision for recognizing food preparation activities
Sebastian Stein and Stephen J McKenna · 2013
Earlier work this paper cites.
Action recognition with improved trajectories
Heng Wang and Cordelia Schmid · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
The language of actions: Recovering the syntax and semantics of goal-directed human activities
Hilde Kuehne, Ali Arslan, and Thomas Serre · 2014
Earlier work this paper cites.
Unsupervised learning of spatiotemporally coherent metrics
Ross Goroshin, Joan Bruna, Jonathan Tompson, David Eigen, and Yann LeCun · 2015
Earlier work this paper cites.
What’s cookin’? interpreting cooking videos using text, speech and vision
Jonathan Malmaud, Jonathan Huang, Vivek Rathod, Nicholas Johnston, Andrew Rabinovich, and Kevin Murphy · 2015
Earlier work this paper cites.
Unsupervised semantic parsing of video collections
Ozan Sener, Amir R Zamir, Silvio Savarese, and Ashutosh Saxena · 2015
Earlier work this paper cites.
Unsupervised learning of video representations using lstms
Nitish Srivastava, Elman Mansimov, and Ruslan Salakhudinov · 2015
Earlier work this paper cites.
Learning spatiotemporal features with 3d convolutional networks
Du Tran, Lubomir Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri · 2015
Earlier work this paper cites.
Unsupervised learning from narrated instruction videos
Jean-Baptiste Alayrac, Piotr Bojanowski, Nishant Agrawal, Josef Sivic, Ivan Laptev, and Simon Lacoste-Julien · 2016
Earlier work this paper cites.
Cliquecnn: Deep unsupervised exemplar learning
Miguel Ángel Bautista, Artsiom Sanakoyeu, Ekaterina Tikhoncheva, and Björn Ommer · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Connectionist temporal modeling for weakly supervised action labeling
De-An Huang, Li Fei-Fei, and Juan Carlos Niebles · 2016
Earlier work this paper cites.
An end-to-end generative framework for video segmentation and recognition
Hilde Kuehne, Juergen Gall, and Thomas Serre · 2016
Earlier work this paper cites.
Learning representations for automatic colorization
Gustav Larsson, Michael Maire, and Gregory Shakhnarovich · 2016
Earlier work this paper cites.
Segmental spatiotemporal cnns for fine-grained action segmentation
Colin Lea, Austin Reiter, René Vidal, and Gregory D Hager · 2016
Earlier work this paper cites.
Shuffle and learn: unsupervised learning using temporal order verification
Ishan Misra, C Lawrence Zitnick, and Martial Hebert · 2016
Earlier work this paper cites.
Temporal action detection using a statistical language model
Alexander Richard and Juergen Gall · 2016
Earlier work this paper cites.
Temporal action localization in untrimmed videos via multi-stage cnns
Zheng Shou, Dongang Wang, and Shih-Fu Chang · 2016
Earlier work this paper cites.
Improved deep metric learning with multi-class n-pair loss objective
Kihyuk Sohn · 2016
Earlier work this paper cites.
Generating videos with scene dynamics
Carl Vondrick, Hamed Pirsiavash, and Antonio Torralba · 2016
Earlier work this paper cites.
Unsupervised deep embedding for clustering analysis
Junyuan Xie, Ross Girshick, and Ali Farhadi · 2016
Earlier work this paper cites.
Joint unsupervised learning of deep representations and image clusters
Jianwei Yang, Devi Parikh, and Dhruv Batra · 2016
Earlier work this paper cites.
Quo vadis, action recognition? a new model and the kinetics dataset
Joao Carreira and Andrew Zisserman · 2017
Cited alongside, same era.
Self-supervised video representation learning with odd-one-out networks
Basura Fernando, Hakan Bilen, Efstratios Gavves, and Stephen Gould · 2017
Cited alongside, same era.
Weakly supervised learning of actions from transcripts
Hilde Kuehne, Alexander Richard, and Juergen Gall · 2017
Cited alongside, same era.
Colorization as a proxy task for visual understanding
Gustav Larsson, Michael Maire, and Gregory Shakhnarovich · 2017
Cited alongside, same era.
Unsupervised representation learning by sorting sequences
Hsin-Ying Lee, Jia-Bin Huang, Maneesh Singh, and Ming-Hsuan Yang · 2017
Cited alongside, same era.
Deep supervision with shape concepts for occlusion-aware 3d object parsing
Chi Li, M Zeeshan Zia, Quoc-Huy Tran, Xiang Yu, Gregory D Hager, and Manmohan Chandraker · 2017
D3tw: Discriminative differentiable dynamic time warping for weakly supervised action alignment and segmentation
Chien-Yi Chang, De-An Huang, Yanan Sui, Li Fei-Fei, and Juan Carlos Niebles · 2019
Later among the works it cites.
Dynamonet: Dynamic action and motion network
Ali Diba, Vivek Sharma, Luc Van Gool, and Rainer Stiefelhagen · 2019
Later among the works it cites.
Temporal cycle-consistency learning
Debidatta Dwibedi, Yusuf Aytar, Jonathan Tompson, Pierre Sermanet, and Andrew Zisserman · 2019
Later among the works it cites.
Self-supervised representation learning by rotation feature decoupling
Zeyu Feng, Chang Xu, and Dacheng Tao · 2019
Later among the works it cites.
Predicting the future: A jointly learnt model for action anticipation
Harshala Gammulle, Simon Denman, Sridha Sridharan, and Clinton Fookes · 2019
Later among the works it cites.
Memorizing normality to detect anomaly: Memory-augmented deep autoencoder for unsupervised anomaly detection
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Representation learning by learning to count
Mehdi Noroozi, Hamed Pirsiavash, and Paolo Favaro · 2017
Cited alongside, same era.
Weakly supervised action learning with rnn based fine-to-coarse modeling
Alexander Richard, Hilde Kuehne, and Juergen Gall · 2017
Cited alongside, same era.
Cdc: Convolutional-de-convolutional networks for precise temporal action localization in untrimmed videos
Zheng Shou, Jonathan Chan, Alireza Zareian, Kazuyuki Miyazawa, and Shih-Fu Chang · 2017
Cited alongside, same era.
Order-preserving wasserstein distance for sequence matching
Bing Su and Gang Hua · 2017
Cited alongside, same era.
Discrimnet: Semi-supervised action recognition from videos using generative adversarial networks
Unaiza Ahsan, Chen Sun, and Irfan Essa · 2018
Cited alongside, same era.
Deep clustering for unsupervised learning of visual features
Mathilde Caron, Piotr Bojanowski, Armand Joulin, and Matthijs Douze · 2018
Cited alongside, same era.
Dong Gong, Lingqiao Liu, Vuong Le, Budhaditya Saha, Moussa Reda Mansour, Svetha Venkatesh, and Anton van den Hengel · 2019
Later among the works it cites.
Video representation learning by dense predictive coding
Tengda Han, Weidi Xie, and Andrew Zisserman · 2019
Later among the works it cites.
Unsupervised deep learning by neighbourhood discovery
Jiabo Huang, Qi Dong, Shaogang Gong, and Xiatian Zhu · 2019
Later among the works it cites.
Self-supervised video representation learning with space-time cubic puzzles
Dahun Kim, Donghyeon Cho, and In So Kweon · 2019
Later among the works it cites.
Unsupervised learning of action classes with continuous temporal embedding
Anna Kukleva, Hilde Kuehne, Fadime Sener, and Jurgen Gall · 2019
Later among the works it cites.
Weakly supervised energy-based learning for action segmentation
Jun Li, Peng Lei, and Sinisa Todorovic · 2019
Later among the works it cites.
Self-supervised spatiotemporal learning via video clip order prediction
Dejing Xu, Jun Xiao, Zhou Zhao, Jian Shao, Di Xie, and Yueting Zhuang · 2019
Later among the works it cites.
Graph convolutional networks for temporal action localization
Runhao Zeng, Wenbing Huang, Mingkui Tan, Yu Rong, Peilin Zhao, Junzhou Huang, and Chuang Gan · 2019
Later among the works it cites.
Degeneracy in self-calibration revisited and a deep learning solution for uncalibrated slam
Bingbing Zhuang, Quoc-Huy Tran, Gim Hee Lee, Loong Fah Cheong, and Manmohan Chandraker · 2019
Later among the works it cites.
Local aggregation for unsupervised learning of visual embeddings
Chengxu Zhuang, Alex Lin Zhai, and Daniel Yamins · 2019
Later among the works it cites.
Unsupervised learning of visual features by contrasting cluster assignments
Mathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal, Piotr Bojanowski, and Armand Joulin · 2020
Later among the works it cites.
Action segmentation with joint self-supervised temporal domain adaptation
Min-Hung Chen, Baopu Li, Yingze Bao, Ghassan AlRegib, and Zsolt Kira · 2020
Later among the works it cites.
Shuffle and attend: Video domain adaptation
Jinwoo Choi, Gaurav Sharma, Samuel Schulter, and Jia-Bin Huang · 2020
Later among the works it cites.
Sct: Set constrained temporal transformer for set supervised action segmentation
Mohsen Fayyaz and Jurgen Gall · 2020
Later among the works it cites.
Learning representations by predicting bags of visual words
Spyros Gidaris, Andrei Bursuc, Nikos Komodakis, Patrick Pérez, and Matthieu Cord · 2020
Later among the works it cites.
Towards anomaly detection in dashcam videos
Sanjay Haresh, Sateesh Kumar, M Zeeshan Zia, and Quoc-Huy Tran · 2020
Later among the works it cites.
Set-constrained viterbi for set-supervised action segmentation
Jun Li and Sinisa Todorovic · 2020
Later among the works it cites.
Ms-tcn++: Multi-stage temporal convolutional network for action segmentation
Shi-Jie Li, Yazan AbuFarha, Yun Liu, Ming-Ming Cheng, and Juergen Gall · 2020
Later among the works it cites.
Clusterfit: Improving generalization of visual representations
Xueting Yan, Ishan Misra, Abhinav Gupta, Deepti Ghadiyaram, and Dhruv Mahajan · 2020
Later among the works it cites.
Learning by aligning videos in time
Sanjay Haresh, Sateesh Kumar, Huseyin Coskun, Shahram Najam Syed, Andrey Konin, Muhammad Zeeshan Zia, and Quoc-Huy Tran · 2021
Closest in time.
Action shuffle alternating learning for unsupervised action segmentation
Jun Li and Sinisa Todorovic · 2021
Closest in time.
Temporal action segmentation from timestamp supervision
Zhe Li, Yazan Abu Farha, and Jurgen Gall · 2021
Closest in time.
Unsupervised discriminative embedding for sub-action learning in complex activities
Sirnam Swetha, Hilde Kuehne, Yogesh S Rawat, and Mubarak Shah · 2021
Closest in time.
Joint visual-temporal embedding for unsupervised learning of actions in untrimmed sequences
Rosaura G VidalMata, Walter J Scheirer, Anna Kukleva, David Cox, and Hilde Kuehne · 2021
Closest in time.
Learning to align sequential actions in the wild
Weizhe Liu, Bugra Tekin, Huseyin Coskun, Vibhav Vineet, Pascal Fua, and Marc Pollefeys · 2022
Closest in time.