Fetching the paper…
Reading the bibliography…
We consider the question: what can be learnt by looking at and listening to a large number of unlabelled videos? There is a valuable, but so far untapped, source of information contained in the video itself -- the correspondence between the visual and the audio streams, and we introduce a novel "Audio-Visual Correspondence" learning task that makes use of this.
Combining labeled and unlabeled data with co-training
A. Blum and T. Mitchell · 1998
Earlier work this paper cites.
Pixels that sound
E. Kidron, Y. Y. Schechner, and M. Elad · 2005
Earlier work this paper cites.
Visualizing data using t-SNE
L. Van der Maaten and G. Hinton · 2008
Earlier work this paper cites.
Neural correlates of interspecies perspective taking in the post-mortem Atlantic salmon: An argument for multiple comparisons correction
C. M. Bennett, M. B. Miller, and G. L. Wolford · 2009
Earlier work this paper cites.
ImageNet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
Devise: A deep visual-semantic embedding model
A. Frome, G. S. Corrado, J. Shlens, S. Bengio, J. Dean, M. A. Ranzato, and T. Mikolov · 2013
Earlier work this paper cites.
Auditory scene classification using machine learning techniques
D. Li, J. Tam, and D. Toub · 2013
Earlier work this paper cites.
Recurrence quantification analysis features for environmental sound recognition
G. Roma, W. Nogueira, and P. Herrera · 2013
Earlier work this paper cites.
Zero-shot learning through cross-modal transfer
R. Socher, M. Ganjoo, C. D. Manning, and A. Ng · 2013
Earlier work this paper cites.
Discriminative unsupervised feature learning with convolutional neural networks
A. Dosovitskiy, J. T. Springenberg, M. Riedmiller, and T. Brox · 2014
Earlier work this paper cites.
Learning to see by moving
P. Agrawal, J. Carreira, and J. Malik · 2015
Earlier work this paper cites.
Unsupervised visual representation learning by context prediction
C. Doersch, A. Gupta, and A. A. Efros · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
S. Ioffe and C. Szegedy · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2015
Cited alongside, same era.
Data-dependent initializations of convolutional neural networks
P. Krähenbühl, C. Doersch, J. Donahue, and T. Darrell · 2015
Cited alongside, same era.
Predicting deep zero-shot convolutional neural networks using textual descriptions
J. Lei Ba, K. Swersky, S. Fidler, and R. Salakhutdinov · 2015
Cited alongside, same era.
Environmental sound classification with convolutional neural networks
K. J. Piczak · 2015
Cited alongside, same era.
ESC: Dataset for environmental sound classification
K. J. Piczak · 2015
Cited alongside, same era.
Out of time: Automated lip sync in the wild
J. S. Chung and A. Zisserman · 2016
Later among the works it cites.
Unsupervised learning of spoken language with visual context
D. Harwath, A. Torralba, and J. R. Glass · 2016
Later among the works it cites.
Shuffle and learn: Unsupervised learning using temporal order verification
I. Misra, C. L. Zitnick, and M. Herbert · 2016
Later among the works it cites.
Unsupervised learning of visual representations by solving jigsaw puzzles
M. Noroozi and P. Favaro · 2016
Later among the works it cites.
Visually indicated sounds
A. Owens, P. Isola, J. McDermott, A. Torralba, E. Adelson, and W. Freeman · 2016
Later among the works it cites.
Ambient sound provides supervision for visual learning
A. Owens, W. Jiajun, J. McDermott, W. Freeman, and A. Torralba · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Histogram of gradients of time-frequency representations for audio scene classification
A. Rakotomamonjy and G. Gasso · 2015
Cited alongside, same era.
Imagenet large scale visual recognition challenge
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, S. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. Berg, and F. Li · 2015
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2015
Cited alongside, same era.
Detection and classification of acoustic scenes and events
D. Stowell, D. Giannoulis, E. Benetos, M. Lagrange, and M. D. Plumbley · 2015
Cited alongside, same era.
Unsupervised learning of visual representations using videos
X. Wang and A. Gupta · 2015
Cited alongside, same era.
SoundNet: Learning sound representations from unlabeled video
Y. Aytar, C. Vondrick, and A. Torralba · 2016
Cited alongside, same era.
Context encoders: Feature learning by inpainting
D. Pathak, P. Krähenbühl, J. Donahue, T. Darrell, and A. A. Efros · 2016
Later among the works it cites.
Rethinking the Inception architecture for computer vision
C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna · 2016
Later among the works it cites.
Colorful image colorization
R. Zhang, P. Isola, and A. A. Efros · 2016
Later among the works it cites.
Adversarial feature learning
J. Donahue, P. Krähenbühl, and T. Darrell · 2017
Closest in time.
Audio Set: An ontology and human-labeled dataset for audio events
J. F. Gemmeke, D. P. W. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter · 2017
Closest in time.
The Kinetics human action video dataset
W. Kay, J. Carreira, K. Simonyan, B. Zhang, C. Hillier, S. Vijayanarasimhan, F. Viola, T. Green, T. Back, P. Natsev, M. Suleyman, and A. Zisserman · 2017
Closest in time.