Fetching the paper…
Reading the bibliography…
We introduce a self-supervised representation learning method based on the task of temporal alignment between videos.
An introduction to multivariate statistical analysis
Theodore Wilbur Anderson · 1958
Earlier work this paper cites.
Neighbourhood components analysis
Jacob Goldberger, Geoffrey E Hinton, Sam T Roweis, and Ruslan R Salakhutdinov · 2005
Earlier work this paper cites.
Matching local self-similarities across images and videos
Eli Shechtman and Michal Irani · 2007
Earlier work this paper cites.
Disambiguating visual relations using loop constraints
Christopher Zach, Manfred Klopschitz, and Marc Pollefeys · 2010
Earlier work this paper cites.
Ucf101: A dataset of 101 human actions classes from videos in the wild
Khurram Soomro, Amir Roshan Zamir, and Mubarak Shah · 2012
Earlier work this paper cites.
Deep canonical correlation analysis
Galen Andrew, Raman Arora, Jeff Bilmes, and Karen Livescu · 2013
Earlier work this paper cites.
Event retrieval in large video collections with circulant temporal encoding
Jérôme Revaud, Matthijs Douze, Cordelia Schmid, and Hervé Jégou · 2013
Earlier work this paper cites.
Image co-segmentation via consistent functional maps
Fan Wang, Qixing Huang, and Leonidas J Guibas · 2013
Earlier work this paper cites.
Action recognition with improved trajectories
Heng Wang and Cordelia Schmid · 2013
Earlier work this paper cites.
Network principles for sfm: Disambiguating repeated structures with local context
Kyle Wilson and Noah Snavely · 2013
Earlier work this paper cites.
From actemes to action: A strongly-supervised representation for detailed action understanding
Weiyu Zhang, Menglong Zhu, and Konstantinos G Derpanis · 2013
Earlier work this paper cites.
Weakly supervised action labeling in videos under ordering constraints
Piotr Bojanowski, Rémi Lajugie, Francis Bach, Ivan Laptev, Jean Ponce, Cordelia Schmid, and Josef Sivic · 2014
Earlier work this paper cites.
Return of the devil in the details: Delving deep into convolutional nets
Ken Chatfield, Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman · 2014
Earlier work this paper cites.
The language of actions: Recovering the syntax and semantics of goal-directed human activities
Hilde Kuehne, Ali Arslan, and Thomas Serre · 2014
Earlier work this paper cites.
Parsing videos of actions with segmental grammars
Hamed Pirsiavash and Deva Ramanan · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Unsupervised multi-class joint image segmentation
Fan Wang, Qixing Huang, Maks Ovsjanikov, and Leonidas J Guibas · 2014
Earlier work this paper cites.
Articulated motion discovery using pairs of trajectories
Luca Del Pero, Susanna Ricco, Rahul Sukthankar, and Vittorio Ferrari · 2015
Earlier work this paper cites.
Unsupervised visual representation learning by context prediction
Carl Doersch, Abhinav Gupta, and Alexei A Efros · 2015
Cited alongside, same era.
Action recognition by hierarchical mid-level action elements
Tian Lan, Yuke Zhu, Amir Roshan Zamir, and Silvio Savarese · 2015
Cited alongside, same era.
Unsupervised semantic parsing of video collections
Ozan Sener, Amir R Zamir, Silvio Savarese, and Ashutosh Saxena · 2015
Cited alongside, same era.
Flowweb: Joint image set alignment by weaving consistent, pixel-wise correspondences
Tinghui Zhou, Yong Jae Lee, Stella X Yu, and Alyosha A Efros · 2015
Cited alongside, same era.
Multi-image matching via fast alternating minimization
Xiaowei Zhou, Menglong Zhu, and Kostas Daniilidis · 2015
Cited alongside, same era.
Unsupervised learning from narrated instruction videos
Jean-Baptiste Alayrac, Piotr Bojanowski, Nishant Agrawal, Ivan Laptev, Josef Sivic, and Simon Lacoste-Julien · 2016
Self-supervised video representation learning with odd-one-out networks
Basura Fernando, Hakan Bilen, Efstratios Gavves, and Stephen Gould · 2017
Later among the works it cites.
Actionvlad: Learning spatio-temporal aggregation for action classification
Rohit Girdhar, Deva Ramanan, Abhinav Gupta, Josef Sivic, and Bryan Russell · 2017
Later among the works it cites.
Cycada: Cycle-consistent adversarial domain adaptation
Judy Hoffman, Eric Tzeng, Taesung Park, Jun-Yan Zhu, Phillip Isola, Kate Saenko, Alexei A Efros, and Trevor Darrell · 2017
Later among the works it cites.
No fuss distance metric learning using proxies
Yair Movshovitz-Attias, Alexander Toshev, Thomas K Leung, Sergey Ioffe, and Saurabh Singh · 2017
Later among the works it cites.
Asynchronous temporal fields for action recognition
Gunnar A Sigurdsson, Santosh Kumar Divvala, Ali Farhadi, and Abhinav Gupta · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Unsupervised feature extraction by time-contrastive learning and nonlinear ica
Aapo Hyvarinen and Hiroshi Morioka · 2016
Cited alongside, same era.
Segmental spatiotemporal cnns for fine-grained action segmentation
Colin Lea, Austin Reiter, René Vidal, and Gregory D Hager · 2016
Cited alongside, same era.
Learning activity progression in lstms for activity detection and early detection
Shugao Ma, Leonid Sigal, and Stan Sclaroff · 2016
Cited alongside, same era.
Shuffle and learn: unsupervised learning using temporal order verification
Ishan Misra, C Lawrence Zitnick, and Martial Hebert · 2016
Cited alongside, same era.
Improved deep metric learning with multi-class n-pair loss objective
Kihyuk Sohn · 2016
Cited alongside, same era.
Prototypical networks for few-shot learning
Jake Snell, Kevin Swersky, and Richard Zemel · 2017
Later among the works it cites.
Temporal action detection with structured segment networks
Yue Zhao, Yuanjun Xiong, Limin Wang, Zhirong Wu, Xiaoou Tang, and Dahua Lin · 2017
Later among the works it cites.
Unpaired image-to-image translation using cycle-consistent adversarial networks
Jun-Yan Zhu, Taesung Park, Phillip Isola, and Alexei A Efros · 2017
Later among the works it cites.
Playing hard exploration games by watching youtube
Yusuf Aytar, Tobias Pfaff, David Budden, Tom Le Paine, Ziyu Wang, and Nando de Freitas · 2018
Later among the works it cites.
Scaling egocentric vision: The epic-kitchens dataset
Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Sanja Fidler, Antonino Furnari, Evangelos Kazakos, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, et al · 2018
Later among the works it cites.
Action completion: A temporal model for moment detection
Farnoosh Heidarivincheh, Majid Mirmehdi, and Dima Damen · 2018
Later among the works it cites.
Neighbourhood consensus networks
Ignacio Rocco, Mircea Cimpoi, Relja Arandjelović, Akihiko Torii, Tomas Pajdla, and Josef Sivic · 2018
Later among the works it cites.
Deep unsupervised learning of visual similarities
Artsiom Sanakoyeu, Miguel A Bautista, and Björn Ommer · 2018
Later among the works it cites.
Unsupervised learning and segmentation of complex activities from video
Fadime Sener and Angela Yao · 2018
Later among the works it cites.
Actor and observer: Joint modeling of first and third-person videos
Gunnar Sigurdsson, Abhinav Gupta, Cordelia Schmid, Ali Farhadi, and Karteek Alahari · 2018
Later among the works it cites.
Coefficient of determination — Wikipedia, the free encyclopedia, 2018
Wikipedia contributors · 2018
Later among the works it cites.
Kendall rank correlation coefficient — Wikipedia, the free encyclopedia, 2018
Wikipedia contributors · 2018
Later among the works it cites.
Every moment counts: Dense detailed labeling of actions in complex videos
Serena Yeung, Olga Russakovsky, Ning Jin, Mykhaylo Andriluka, Greg Mori, and Li Fei-Fei · 2018
Later among the works it cites.