Fetching the paper…
Reading the bibliography…
In this paper, we investigate how to learn rich and robust feature representations for audio classification from visual data and acoustic images, a novel audio data modality.
Integration of acoustic and visual speech signals using neural networks
B. P. Yuhas, M. H. Goldstein, and T. J. Sejnowski · 1989
Earlier work this paper cites.
Learning classification with unlabeled data
V. R. DeSa · 1993
Earlier work this paper cites.
What in the world do we hear?: An ecological approach to auditory event perception
W. W. Gaver · 1993
Earlier work this paper cites.
Converging influences from visual, auditory, and somatosensory cortices onto output neurons of the superior colliculus
M. T. Wallace, M. A. Meredith, and B. E. Stein · 1993
Earlier work this paper cites.
Object Recognition with Gradient-Based Learning
Y. LeCun, P. Haffner, L. Bottou, and Y. Bengio · 1999
Earlier work this paper cites.
Detection, Estimation, and Modulation Theory, Optimum Array Processing
H. Van Trees · 2002
Earlier work this paper cites.
Knn model-based approach in classification
G. Guo, H. Wang, D. A. Bell, Y. Bi, and K. Greer · 2003
Earlier work this paper cites.
Pixels that sound
E. Kidron, Y. Y. Schechner, and M. Elad · 2005
Earlier work this paper cites.
A statistical model of timbre perception
H. Terasawa, M. Slaney, and J. Berger · 2006
Earlier work this paper cites.
Meaningful auditory information enhances perception of visual biological motion
R. Arrighi, F. Marini, and D. Burr · 2009
Earlier work this paper cites.
Multimodal deep learning
J. Ngiam, A. Khosla, M. Kim, J. Nam, H. Lee, and A. Y. Ng · 2011
Earlier work this paper cites.
Distilling the knowledge in a neural network
G. E. Hinton, O. Vinyals, and J. Dean · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
S. Ioffe and C. Szegedy · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2015
Earlier work this paper cites.
Esc: Dataset for environmental sound classification
K. J. Piczak · 2015
Earlier work this paper cites.
Seeing the sound: A new multimodal imaging device for computer vision
A. Zunino, M. Crocco, S. Martelli, A. Trucco, A. D. Bue, and V. Murino · 2015
Earlier work this paper cites.
Soundnet: Learning sound representations from unlabeled video
Y. Aytar, C. Vondrick, and A. Torralba · 2016
Cited alongside, same era.
Learning aligned cross-modal representations from weakly aligned data
L. Castrejón, Y. Aytar, C. Vondrick, H. Pirsiavash, and A. Torralba · 2016
Cited alongside, same era.
Unsupervised learning of spoken language with visual context
D. Harwath, A. Torralba, and J. Glass · 2016
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
Learning with side information through modality hallucination
J. Hoffman, S. Gupta, and T. Darrell · 2016
Cited alongside, same era.
Unifying distillation and privileged information
D. Lopez-Paz, L. Bottou, B. Schölkopf, and V. Vapnik · 2016
Visual speech enhancement
A. Gabbay, A. Shamir, and S. Peleg · 2018
Later among the works it cites.
Modality distillation with multiple stream networks for action recognition
N. C. Garcia, P. Morerio, and V. Murino · 2018
Later among the works it cites.
Unsupervised adversarial domain adaptation for acoustic scene classification
S. Gharib, K. Drossos, E. Çakir, D. Serdyuk, and T. Virtanen · 2018
Later among the works it cites.
Acoustic scene classification using convolutional neural networks and different channels representations and its fusion
A. Golubkov and A. Lavrentyev · 2018
Later among the works it cites.
Acoustic scene classification using multi-scale features
Y. Liping, C. Xinxing, and T. Lianjie · 2018
Later among the works it cites.
A multi-device dataset for urban acoustic scene classification
A. Mesaros, T. Heittola, and T. Virtanen · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Visually indicated sounds
A. Owens, P. Isola, J. McDermott, A. Torralba, E. H. Adelson, and W. T. Freeman · 2016
Cited alongside, same era.
Ambient sound provides supervision for visual learning
A. Owens, J. Wu, J. H. McDermott, W. T. Freeman, and A. Torralba · 2016
Cited alongside, same era.
Look, listen and learn
R. Arandjelovic and A. Zisserman · 2017
Cited alongside, same era.
See, hear, and read: Deep aligned representations
Y. Aytar, C. Vondrick, and A. Torralba · 2017
Cited alongside, same era.
Audio set: An ontology and human-labeled dataset for audio events
J. F. Gemmeke, D. P. W. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter · 2017
Cited alongside, same era.
The kinetics human action video dataset
W. Kay, J. Carreira, K. Simonyan, B. Zhang, C. Hillier, S. Vijayanarasimhan, F. Viola, T. Green, T. Back, A. Natsev, M. Suleyman, and A. Zisserman · 2017
Cited alongside, same era.
Later among the works it cites.
Audio-visual scene analysis with self-supervised multisensory features
A. Owens and A. A. Efros · 2018
Later among the works it cites.
Learning sight from sound: Ambient sound provides supervision for visual learning
A. Owens, J. Wu, J. H. McDermott, W. T. Freeman, and A. Torralba · 2018
Later among the works it cites.
Weakly supervised representation learning for unsynchronized audio-visual events
S. Parekh, S. Essid, A. Ozerov, N. Q. K. Duong, P. Perez, and G. Richard · 2018
Later among the works it cites.
Speaker recognition from raw waveform with sincnet
M. Ravanelli and Y. Bengio · 2018
Later among the works it cites.
Learning to localize sound source in visual scenes
A. Senocak, T.-H. Oh, J. Kim, M.-H. Yang, and I. So Kweon · 2018
Later among the works it cites.
On learning association of sound source and visual scenes
A. Senocak, T.-H. Oh, J. Kim, M.-H. Yang, and I. So Kweon · 2018
Later among the works it cites.
Audio-visual event localization in unconstrained videos
Y. Tian, J. Shi, B. Li, Z. Duan, and C. Xu · 2018
Later among the works it cites.
A closer look at spatiotemporal convolutions for action recognition
D. Tran, H. Wang, L. Torresani, J. Ray, Y. LeCun, and M. Paluri · 2018
Later among the works it cites.
The sound of pixels
H. Zhao, C. Gan, A. Rouditchenko, C. Vondrick, J. McDermott, and A. Torralba · 2018
Later among the works it cites.
Multi-stream network with temporal attention for environmental sound classification
X. Li, V. Chebiyyam, and K. Kirchhoff · 2019
Closest in time.