Fetching the paper…
Reading the bibliography…
A large part of the current success of deep learning lies in the effectiveness of data -- more precisely: labelled data.
Contrastive bidirectional transformer for temporal representation learning
Chen Sun, Fabien Baradel, Kevin Murphy, and Cordelia Schmid · 1906
Earlier work this paper cites.
The hungarian method for the assignment problem
Harold W Kuhn · 1955
Earlier work this paper cites.
Direct clustering of a data matrix
John A Hartigan · 1972
Earlier work this paper cites.
Worst-case analysis of a new heuristic for the travelling salesman problem
Nicos Christofides · 1976
Earlier work this paper cites.
Learning classification with unlabeled data
Virginia R. de Sa · 1994
Earlier work this paper cites.
Greedy randomized adaptive search procedures
Thomas Feo and Mauricio Resende · 1995
Earlier work this paper cites.
A simple framework for contrastive learning of visual representations
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey E. Hinton · 2002
Earlier work this paper cites.
An annotated bibliography of grasp
Paola Festa and Mauricio GC Resende · 2004
Earlier work this paper cites.
Blind audiovisual source separation based on sparse redundant representations
Anna Llagostera Casanovas, Gianluca Monaci, Pierre Vandergheynst, and Rémi Gribonval · 2010
Earlier work this paper cites.
HMDB: a large video database for human motion recognition
H. Kuehne, H. Jhuang, E. Garrote, T. Poggio, and T. Serre · 2011
Earlier work this paper cites.
UCF101: A dataset of 101 human action classes from videos in the wild
Khurram Soomro, Amir Roshan Zamir, and Mubarak Shah · 2012
Earlier work this paper cites.
Sinkhorn distances: Lightspeed computation of optimal transport
Marco Cuturi · 2013
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean · 2015
Earlier work this paper cites.
YouTube-8M: A large-scale video classification benchmark
Sami Abu-El-Haija, Nisarg Kothari, Joonseok Lee, Paul Natsev, George Toderici, Balakrishnan Varadarajan, and Sudheendra Vijayanarasimhan · 2016
Earlier work this paper cites.
Soundnet: Learning sound representations from unlabeled video
Yusuf Aytar, Carl Vondrick, and Antonio Torralba · 2016
Earlier work this paper cites.
Cliquecnn: Deep unsupervised exemplar learning
Miguel A Bautista, Artsiom Sanakoyeu, Ekaterina Tikhoncheva, and Bjorn Ommer · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Shuffle and learn: unsupervised learning using temporal order verification
Ishan Misra, C Lawrence Zitnick, and Martial Hebert · 2016
Earlier work this paper cites.
Unsupervised learning of visual representations by solving jigsaw puzzles
Mehdi Noroozi and Paolo Favaro · 2016
Earlier work this paper cites.
Ambient sound provides supervision for visual learning
Andrew Owens, Jiajun Wu, Josh H McDermott, William T Freeman, and Antonio Torralba · 2016
Earlier work this paper cites.
Unsupervised deep embedding for clustering analysis
Junyuan Xie, Ross Girshick, and Ali Farhadi · 2016
Earlier work this paper cites.
Joint unsupervised learning of deep representations and image clusters
Jianwei Yang, Devi Parikh, and Dhruv Batra · 2016
Earlier work this paper cites.
Colorful image colorization
Richard Zhang, Phillip Isola, and Alexei A Efros · 2016
Earlier work this paper cites.
Look, listen and learn
Relja Arandjelovic and Andrew Zisserman · 2017
Earlier work this paper cites.
Deep unsupervised similarity learning using partially ordered sets
Miguel A Bautista, Artsiom Sanakoyeu, and Bjorn Ommer · 2017
Earlier work this paper cites.
Deep adaptive image clustering
Jianlong Chang, Lingfeng Wang, Gaofeng Meng, Shiming Xiang, and Chunhong Pan · 2017
Earlier work this paper cites.
Accurate, large minibatch SGD: training imagenet in 1 hour
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2017
Cited alongside, same era.
Learning discrete representations via information maximizing self-augmented training
Weihua Hu, Takeru Miyato, Seiya Tokui, Eiichi Matsumoto, and Masashi Sugiyama · 2017
Cited alongside, same era.
Billion-scale similarity search with gpus
Jeff Johnson, Matthijs Douze, and Hervé Jégou · 2017
Cited alongside, same era.
The kinetics human action video dataset
Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, Mustafa Suleyman, and Andrew Zisserman · 2017
Cited alongside, same era.
Unsupervised representation learning by sorting sequences
Video representation learning by dense predictive coding
Tengda Han, Weidi Xie, and Andrew Zisserman · 2019
Later among the works it cites.
Deep multimodal clustering for unsupervised audiovisual learning
Di Hu, Feiping Nie, and Xuelong Li · 2019
Later among the works it cites.
A novel distributed multitask fuzzy clustering algorithm for automatic mr brain image segmentation
Yizhang Jiang, Kaifa Zhao, Kaijian Xia, Jing Xue, Leyuan Zhou, Yang Ding, and Pengjiang Qian · 2019
Later among the works it cites.
Self-supervised video representation learning with space-time cubic puzzles
Dahun Kim, Donghyeon Cho, and In So Kweon · 2019
Later among the works it cites.
Juho Lee, Yoonho Lee, and Yee Whye Teh · 2019
Later among the works it cites.
Clu-cnns: Object detection for medical images
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hsin-Ying Lee, Jia-Bin Huang, Maneesh Singh, and Ming-Hsuan Yang · 2017
Cited alongside, same era.
Representation learning by learning to count
Mehdi Noroozi, Hamed Pirsiavash, and Paolo Favaro · 2017
Cited alongside, same era.
Dmitry Ulyanov, Andrea Vedaldi, and Victor Lempitsky · 2017
Cited alongside, same era.
Objects that sound
Relja Arandjelović and Andrew Zisserman · 2018
Cited alongside, same era.
Improving spatiotemporal self-supervision by deep reinforcement learning
Uta Buchler, Biagio Brattoli, and Bjorn Ommer · 2018
Cited alongside, same era.
Deep clustering for unsupervised learning of visual features
Mathilde Caron, Piotr Bojanowski, Armand Joulin, and Matthijs Douze · 2018
Cited alongside, same era.
Learning to separate object sounds by watching unlabeled video
Ruohan Gao, Rogerio Feris, and Kristen Grauman · 2018
Cited alongside, same era.
Unsupervised representation learning by predicting image rotations
Spyros Gidaris, Praveer Singh, and Nikos Komodakis · 2018
Cited alongside, same era.
Zhuoling Li, Minghui Dong, Shiping Wen, Xiang Hu, Pan Zhou, and Zhigang Zeng · 2019
Later among the works it cites.
End-to-end learning of visual representations from uncurated instructional videos, 2019
Antoine Miech, Jean-Baptiste Alayrac, Lucas Smaira, Ivan Laptev, Josef Sivic, and Andrew Zisserman · 2019
Later among the works it cites.
Greedy randomized adaptive search procedures: Advances and extensions
Mauricio GC Resende and Celso C Ribeiro · 2019
Later among the works it cites.
Self-supervised audio-visual co-segmentation
Andrew Rouditchenko, Hang Zhao, Chuang Gan, Josh McDermott, and Antonio Torralba · 2019
Later among the works it cites.
Yonglong Tian, Dilip Krishnan, and Phillip Isola · 2019
Later among the works it cites.
Self-supervised spatio-temporal representation learning for videos by predicting motion and appearance statistics
Jiangliu Wang, Jianbo Jiao, Linchao Bao, Shengfeng He, Yunhui Liu, and Wei Liu · 2019
Later among the works it cites.
Self-supervised spatiotemporal learning via video clip order prediction
Dejing Xu, Jun Xiao, Zhou Zhao, Jian Shao, Di Xie, and Yueting Zhuang · 2019
Later among the works it cites.
The sound of motions
Hang Zhao, Chuang Gan, Wei-Chiu Ma, and Antonio Torralba · 2019
Later among the works it cites.
Speednet: Learning the speediness in videos, 2020
Sagie Benaim, Ariel Ephrat, Oran Lang, Inbar Mosseri, William T. Freeman, Michael Rubinstein, Michal Irani, and Tali Dekel · 2020
Closest in time.
Self-supervised spatio-temporal representation learning using variable playback speed prediction
Hyeon Cho, Taehoon Kim, Hyung Jin Chang, and Wonjun Hwang · 2020
Closest in time.
Learning representations by predicting bags of visual words, 2020
Spyros Gidaris, Andrei Bursuc, Nikos Komodakis, Patrick Pérez, and Matthieu Cord · 2020
Closest in time.
Prototypical contrastive learning of unsupervised representations
Junnan Li, Pan Zhou, Caiming Xiong, Richard Socher, and Steven CH Hoi · 2020
Closest in time.
Learning spatiotemporal features via video and text pair discrimination, 2020
Tianhao Li and Limin Wang · 2020
Closest in time.
Video cloze procedure for self-supervised spatio-temporal learning
Dezhao Luo, Chang Liu, Yu Zhou, Dongbao Yang, Can Ma, Qixiang Ye, and Weiping Wang · 2020
Closest in time.
Audio-visual instance discrimination with cross-modal agreement, 2020
Pedro Morgado, Nuno Vasconcelos, and Ishan Misra · 2020
Closest in time.
Speech2action: Cross-modal supervision for action recognition
Arsha Nagrani, Chen Sun, David Ross, Rahul Sukthankar, Cordelia Schmid, and Andrew Zisserman · 2020
Closest in time.
Multi-modal self-supervision from generalized data transformations, 2020
Mandela Patrick, Yuki M. Asano, Ruth Fong, João F. Henriques, Geoffrey Zweig, and Andrea Vedaldi · 2020
Closest in time.
Evolving losses for unsupervised video representation learning, 2020
AJ Piergiovanni, Anelia Angelova, and Michael S. Ryoo · 2020
Closest in time.
Scan: Learning to classify images without labels
Wouter Van Gansbeke, Simon Vandenhende, Stamatios Georgoulis, Marc Proesmans, and Luc Van Gool · 2020
Closest in time.
Audiovisual slowfast networks for video recognition
Fanyi Xiao, Yong Jae Lee, Kristen Grauman, Jitendra Malik, and Christoph Feichtenhofer · 2020
Closest in time.
ClusterFit: Improving Generalization of Visual Representations
Xueting Yan, Ishan Misra, Abhinav Gupta, Deepti Ghadiyaram, and Dhruv Mahajan · 2020
Closest in time.