Fetching the paper…
Reading the bibliography…
We present a self-supervised learning approach to learn audio-visual representations from video and audio.
Learning classification with unlabeled data
Virginia R de Sa · 1994
Earlier work this paper cites.
Emergence of simple-cell receptive field properties by learning a sparse code for natural images
Bruno A Olshausen and David J Field · 1996
Earlier work this paper cites.
Combining labeled and unlabeled data with co-training
Avrim Blum and Tom Mitchell · 1998
Earlier work this paper cites.
Sparse coding of time-varying natural images
Bruno A Olshausen · 2000
Earlier work this paper cites.
Multiple kernel learning, conic duality, and the smo algorithm
Francis R Bach, Gert RG Lanckriet, and Michael I Jordan · 2004
Earlier work this paper cites.
Multi-view clustering
Steffen Bickel and Tobias Scheffer · 2004
Earlier work this paper cites.
Learning the kernel matrix with semidefinite programming
Gert RG Lanckriet, Nello Cristianini, Peter Bartlett, Laurent El Ghaoui, and Michael I Jordan · 2004
Earlier work this paper cites.
Pixels that sound
Einat Kidron, Yoav Y Schechner, and Michael Elad · 2005
Earlier work this paper cites.
Semi-supervised self-training of object detection models
Chuck Rosenberg, Martial Hebert, and Henry Schneiderman · 2005
Earlier work this paper cites.
Dimensionality reduction by learning an invariant mapping
Raia Hadsell, Sumit Chopra, and Yann LeCun · 2006
Earlier work this paper cites.
Efficient sparse coding algorithms
Honglak Lee, Alexis Battle, Rajat Raina, and Andrew Y Ng · 2007
Earlier work this paper cites.
Unsupervised learning of invariant feature hierarchies with applications to object recognition
Marc’aurelio Ranzato, Fu Jie Huang, Y-Lan Boureau, and Yann LeCun · 2007
Earlier work this paper cites.
Analyzing co-training style algorithms
Wei Wang and Zhi-Hua Zhou · 2007
Earlier work this paper cites.
Multiview fisher discriminant analysis
Tom Diethe, David R Hardoon, and John Shawe-Taylor · 2008
Earlier work this paper cites.
Extracting and composing robust features with denoising autoencoders
Pascal Vincent, Hugo Larochelle, Yoshua Bengio, and Pierre-Antoine Manzagol · 2008
Earlier work this paper cites.
Deep learning from temporal coherence in video
Hossein Mobahi, Ronan Collobert, and Jason Weston · 2009
Earlier work this paper cites.
Deep boltzmann machines
Ruslan Salakhutdinov and Geoffrey Hinton · 2009
Earlier work this paper cites.
Noise-contrastive estimation: A new estimation principle for unnormalized statistical models
Michael Gutmann and Aapo Hyvärinen · 2010
Earlier work this paper cites.
The local rademacher complexity of lp-norm multiple kernel learning
Marius Kloft and Gilles Blanchard · 2011
Earlier work this paper cites.
HMDB: a large video database for human motion recognition
Hildegard Kuehne, Hueihan Jhuang, Estíbaliz Garrote, Tomaso Poggio, and Thomas Serre · 2011
Earlier work this paper cites.
Co-regularized multi-view spectral clustering
Abhishek Kumar, Piyush Rai, and Hal Daume · 2011
Earlier work this paper cites.
Ensemble of exemplar-svms for object detection and beyond
Tomasz Malisiewicz, Abhinav Gupta, and Alexei A Efros · 2011
Earlier work this paper cites.
Stacked convolutional auto-encoders for hierarchical feature extraction
Jonathan Masci, Ueli Meier, Dan Cireşan, and Jürgen Schmidhuber · 2011
Earlier work this paper cites.
Learning multi-view neighborhood preserving projections
Novi Quadrianto and Christoph Lampert · 2011
Earlier work this paper cites.
UCF101: A dataset of 101 human actions classes from videos in the wild
Khurram Soomro, Amir Roshan Zamir, and Mubarak Shah · 2012
Earlier work this paper cites.
Building high-level features using large scale unsupervised learning
Quoc V Le, Marc’Aurelio Ranzato, Rajat Monga, Matthieu Devin, Kai Chen, Greg S Corrado, Jeff Dean, and Andrew Y Ng · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Learning large-scale automatic image colorization
Aditya Deshpande, Jason Rock, and David Forsyth · 2015
Earlier work this paper cites.
Unsupervised visual representation learning by context prediction
Carl Doersch, Abhinav Gupta, and Alexei A Efros · 2015
Earlier work this paper cites.
Environmental sound classification with convolutional neural networks
Karol J Piczak · 2015
Earlier work this paper cites.
Esc: Dataset for environmental sound classification
Karol J Piczak · 2015
Earlier work this paper cites.
Facenet: A unified embedding for face recognition and clustering
Florian Schroff, Dmitry Kalenichenko, and James Philbin · 2015
Earlier work this paper cites.
Detection and classification of acoustic scenes and events
Dan Stowell, Dimitrios Giannoulis, Emmanouil Benetos, Mathieu Lagrange, and Mark D Plumbley · 2015
Cited alongside, same era.
Going deeper with convolutions
Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich · 2015
Cited alongside, same era.
Unsupervised learning of visual representations using videos
Xiaolong Wang and Abhinav Gupta · 2015
Cited alongside, same era.
Soundnet: Learning sound representations from unlabeled video
Yusuf Aytar, Carl Vondrick, and Antonio Torralba · 2016
Cited alongside, same era.
Out of time: automated lip sync in the wild
Joon Son Chung and Andrew Zisserman · 2016
Cited alongside, same era.
Discriminative unsupervised feature learning with exemplar convolutional neural networks
Alexey Dosovitskiy, Philipp Fischer, Jost Tobias Springenberg, Martin Riedmiller, and Thomas Brox · 2016
Self-supervised spatiotemporal feature learning by video geometric transformations
Longlong Jing and Yingli Tian · 2018
Later among the works it cites.
Cooperative learning of audio and video models from self-supervised synchronization
Bruno Korbar, Du Tran, and Lorenzo Torresani · 2018
Later among the works it cites.
Self-supervised generation of spatial audio for 360 video
Pedro Morgado, Nuno Nvasconcelos, Timothy Langlois, and Oliver Wang · 2018
Later among the works it cites.
Representation learning with contrastive predictive coding
Aaron van den Oord, Yazhe Li, and Oriol Vinyals · 2018
Later among the works it cites.
Audio-visual scene analysis with self-supervised multisensory features
Andrew Owens and Alexei A Efros · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Learning representations for automatic colorization
Gustav Larsson, Michael Maire, and Gregory Shakhnarovich · 2016
Cited alongside, same era.
Shuffle and learn: unsupervised learning using temporal order verification
Ishan Misra, C Lawrence Zitnick, and Martial Hebert · 2016
Cited alongside, same era.
Unsupervised learning of visual representations by solving jigsaw puzzles
Mehdi Noroozi and Paolo Favaro · 2016
Cited alongside, same era.
Ambient sound provides supervision for visual learning
Andrew Owens, Jiajun Wu, Josh H McDermott, William T Freeman, and Antonio Torralba · 2016
Cited alongside, same era.
Colorful image colorization
Richard Zhang, Phillip Isola, and Alexei A Efros · 2016
Cited alongside, same era.
Look, listen and learn
Relja Arandjelovic and Andrew Zisserman · 2017
Cited alongside, same era.
Deep co-training for semi-supervised image recognition
Siyuan Qiao, Wei Shen, Zhishuai Zhang, Bo Wang, and Alan Yuille · 2018
Later among the works it cites.
Learning to localize sound source in visual scenes
Arda Senocak, Tae-Hyun Oh, Junsik Kim, Ming-Hsuan Yang, and In So Kweon · 2018
Later among the works it cites.
Time-contrastive networks: Self-supervised learning from video
Pierre Sermanet, Corey Lynch, Yevgen Chebotar, Jasmine Hsu, Eric Jang, Stefan Schaal, Sergey Levine, and Google Brain · 2018
Later among the works it cites.
A closer look at spatiotemporal convolutions for action recognition
Du Tran, Heng Wang, Lorenzo Torresani, Jamie Ray, Yann LeCun, and Manohar Paluri · 2018
Later among the works it cites.
Unsupervised feature learning via non-parametric instance discrimination
Zhirong Wu, Yuanjun Xiong, Stella X Yu, and Dahua Lin · 2018
Later among the works it cites.
The sound of pixels
Hang Zhao, Chuang Gan, Andrew Rouditchenko, Carl Vondrick, Josh McDermott, and Antonio Torralba · 2018
Later among the works it cites.
Unsupervised pre-training of image features on non-curated data
Mathilde Caron, Piotr Bojanowski, Julien Mairal, and Armand Joulin · 2019
Later among the works it cites.
Temporal cycle-consistency learning
Debidatta Dwibedi, Yusuf Aytar, Jonathan Tompson, Pierre Sermanet, and Andrew Zisserman · 2019
Later among the works it cites.
Self-supervised representation learning by rotation feature decoupling
Zeyu Feng, Chang Xu, and Dacheng Tao · 2019
Later among the works it cites.
Co-separating sounds of visual objects
Ruohan Gao and Kristen Grauman · 2019
Later among the works it cites.
Scaling and benchmarking self-supervised visual representation learning
Priya Goyal, Dhruv Mahajan, Abhinav Gupta, and Ishan Misra · 2019
Later among the works it cites.
Video representation learning by dense predictive coding
Tengda Han, Weidi Xie, and Andrew Zisserman · 2019
Later among the works it cites.
Contrastive bidirectional transformer for temporal representation learning
Chen Sun, Fabien Baradel, Kevin Murphy, and Cordelia Schmid · 2019
Later among the works it cites.
Contrastive multiview coding
Yonglong Tian, Dilip Krishnan, and Phillip Isola · 2019
Later among the works it cites.
Self-supervised spatiotemporal learning via video clip order prediction
Dejing Xu, Jun Xiao, Zhou Zhao, Jian Shao, Di Xie, and Yueting Zhuang · 2019
Later among the works it cites.
Unsupervised embedding learning via invariant and spreading instance feature
Mang Ye, Xu Zhang, Pong C Yuen, and Shih-Fu Chang · 2019
Later among the works it cites.
Aet vs. aed: Unsupervised representation learning by auto-encoding transformations rather than data
Liheng Zhang, Guo-Jun Qi, Liqiang Wang, and Jiebo Luo · 2019
Later among the works it cites.
The sound of motions
Hang Zhao, Chuang Gan, Wei-Chiu Ma, and Antonio Torralba · 2019
Later among the works it cites.
Local aggregation for unsupervised learning of visual embeddings
Chengxu Zhuang, Alex Lin Zhai, and Daniel Yamins · 2019
Later among the works it cites.
Self-supervised learning by cross-modal audio-video clustering
Humam Alwassel, Dhruv Mahajan, Bruno Korbar, Lorenzo Torresani, Bernard Ghanem, and Du Tran · 2020
Closest in time.
Music gesture for visual sound separation
Chuang Gan, Deng Huang, Hang Zhao, Joshua B Tenenbaum, and Antonio Torralba · 2020
Closest in time.
Self-supervised co-training for video representation learning
Tengda Han, Weidi Xie, and Andrew Zisserman · 2020
Closest in time.
Momentum contrast for unsupervised visual representation learning
Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick · 2020
Closest in time.
Data-efficient image recognition with contrastive predictive coding
Olivier Henaff · 2020
Closest in time.
Contrastive learning with adversarial examples
Chih-Hui Ho and Nuno Vasconcelos · 2020
Closest in time.
Self-supervised learning of pretext-invariant representations
Ishan Misra and Laurens van der Maaten · 2020
Closest in time.