Fetching the paper…
Reading the bibliography…
We introduce a new approach for audio-visual speech separation.
Signal estimation from modified short-time fourier transform
D. Griffin and J. Lim · 1984
Earlier work this paper cites.
Understanding face recognition
V. Bruce and A. Young · 1986
Earlier work this paper cites.
Auditory scene analysis: The perceptual organization of sound
A. S. Bregman · 1994
Earlier work this paper cites.
Audio vision: Using audio-visual synchrony to locate sounds
J. R. Hershey and J. R. Movellan · 2000
Earlier work this paper cites.
Learning joint statistical models for audio-visual fusion and segregation
J. W. Fisher III, T. Darrell, W. T. Freeman, and P. A. Viola · 2001
Earlier work this paper cites.
Perceptual evaluation of speech quality (pesq)-a new method for speech quality assessment of telephone networks and codecs
A. W. Rix, J. G. Beerends, M. P. Hollier, and A. P. Hekstra · 2001
Earlier work this paper cites.
One microphone source separation
S. T. Roweis · 2001
Earlier work this paper cites.
Real-time speaker localization and speech separation by audio-visual integration
K. Nakadai, K.-i. Hidai, H. G. Okuno, and H. Kitano · 2002
Earlier work this paper cites.
Cuave: A new audio-visual database for multimodal human-computer interface research
E. K. Patterson, S. Gurbuz, Z. Tufekci, and J. N. Gowdy · 2002
Earlier work this paper cites.
Audio/visual independent components
P. Smaragdis and M. Casey · 2003
Earlier work this paper cites.
Blind separation of speech mixtures via time-frequency masking
O. Yilmaz and S. Rickard · 2004
Earlier work this paper cites.
Visual cues can modulate integration and segregation of objects in auditory scene analysis
T. Rahne, M. Böckmann, H. von Specht, and E. S. Sussman · 2007
Earlier work this paper cites.
Supervised and semi-supervised separation of sounds from single-channel mixtures
P. Smaragdis, B. Raj, and M. Shashanka · 2007
Earlier work this paper cites.
Visualizing data using t-sne
L. v. d. Maaten and G. Hinton · 2008
Earlier work this paper cites.
Source-filter based clustering for monaural blind source separation
M. Spiertz and G. Volker · 2009
Earlier work this paper cites.
Blind audiovisual source separation based on sparse redundant representations
A. L. Casanovas, G. Monaci, P. Vandergheynst, and R. Gribonval · 2010
Earlier work this paper cites.
Under-determined reverberant audio source separation using a full-rank spatial covariance model
N. Q. Duong, E. Vincent, and R. Gribonval · 2010
Earlier work this paper cites.
An algorithm for intelligibility prediction of time–frequency weighted noisy speech
C. H. Taal, R. C. Hendriks, R. Heusdens, and J. Jensen · 2011
Earlier work this paper cites.
Do body mass index and fat volume influence vocal quality, phonatory range, and aerodynamics in females?
B. Barsties, R. Verfaillie, N. Roy, and Y. Maryn · 2013
Earlier work this paper cites.
Visual input enhances selective speech envelope tracking in auditory cortex at a “cocktail party”
E. Z. Golumbic, G. B. Cogan, C. E. Schroeder, and D. Poeppel · 2013
Earlier work this paper cites.
Deep learning for monaural speech separation
P.-S. Huang, M. Kim, M. Hasegawa-Johnson, and P. Smaragdis · 2014
Earlier work this paper cites.
mir_eval: A transparent implementation of common mir metrics
C. Raffel, B. McFee, E. J. Humphrey, J. Salamon, O. Nieto, D. Liang, D. P. Ellis, and C. C. Raffel · 2014
Earlier work this paper cites.
On training targets for supervised speech separation
Y. Wang, A. Narayanan, and D. Wang · 2014
Earlier work this paper cites.
Tcd-timit: An audio-visual corpus of continuous speech
N. Harte and E. Gillen · 2015
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
O. Ronneberger, P. Fischer, and T. Brox · 2015
Earlier work this paper cites.
Speech enhancement with lstm recurrent neural networks and its application to noise-robust asr
F. Weninger, H. Erdogan, S. Watanabe, E. Vincent, J. Le Roux, J. R. Hershey, and B. Schuller · 2015
Earlier work this paper cites.
Complex ratio masking for monaural speech separation
D. S. Williamson, Y. Wang, and D. Wang · 2015
Earlier work this paper cites.
Soundnet: Learning sound representations from unlabeled video
Y. Aytar, C. Vondrick, and A. Torralba · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
Deep clustering: Discriminative embeddings for segmentation and separation
J. R. Hershey, Z. Chen, J. Le Roux, and S. Watanabe · 2016
Cited alongside, same era.
Speech enhancement in multiple-noise conditions using deep neural networks
A. Kumar and D. Florencio · 2016
Cited alongside, same era.
Visually indicated sounds
A. Owens, P. Isola, J. McDermott, A. Torralba, E. H. Adelson, and W. T. Freeman · 2016
Cited alongside, same era.
Ambient sound provides supervision for visual learning
A. Owens, J. Wu, J. H. McDermott, W. T. Freeman, and A. Torralba · 2016
Cited alongside, same era.
Look, listen and learn
R. Arandjelovic and A. Zisserman · 2017
Cited alongside, same era.
Learning to localize sound source in visual scenes
A. Senocak, T.-H. Oh, J. Kim, M.-H. Yang, and I. So Kweon · 2018
Later among the works it cites.
Audio-visual event localization in unconstrained videos
Y. Tian, J. Shi, B. Li, Z. Duan, and C. Xu · 2018
Later among the works it cites.
The sound of pixels
H. Zhao, C. Gan, A. Rouditchenko, C. Vondrick, J. McDermott, and A. Torralba · 2018
Later among the works it cites.
Visual to sound: Generating natural sound for videos in the wild
Y. Zhou, Z. Wang, C. Fang, T. Bui, and T. L. Berg · 2018
Later among the works it cites.
My lips are concealed: Audio-visual speech enhancement through obstructions
T. Afouras, J. S. Chung, and A. Zisserman · 2019
Later among the works it cites.
Perfect match: Improved cross-modal embeddings for audio-visual synchronisation
S.-W. Chung, J. S. Chung, and H.-G. Kang · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Audio set: An ontology and human-labeled dataset for audio events
J. F. Gemmeke, D. P. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter · 2017
Cited alongside, same era.
Multitalker speech separation with utterance-level permutation invariant training of deep recurrent neural networks
M. Kolbæk, D. Yu, Z.-H. Tan, and J. Jensen · 2017
Cited alongside, same era.
Voxceleb: a large-scale speaker identification dataset
A. Nagrani, J. S. Chung, and A. Zisserman · 2017
Cited alongside, same era.
Motion informed audio source separation
S. Parekh, S. Essid, A. Ozerov, N. Q. Duong, P. Pérez, and G. Richard · 2017
Cited alongside, same era.
Audio-visual object localization and separation using low-rank and sparsity
J. Pu, Y. Panagakis, S. Petridis, and M. Pantic · 2017
Cited alongside, same era.
Permutation invariant training of deep models for speaker-independent multi-talker speech separation
D. Yu, M. Kolbæk, Z.-H. Tan, and J. Jensen · 2017
Cited alongside, same era.
2.5d visual sound
R. Gao and K. Grauman · 2019
Later among the works it cites.
Co-separating sounds of visual objects
R. Gao and K. Grauman · 2019
Later among the works it cites.
Epic-fusion: Audio-visual temporal binding for egocentric action recognition
E. Kazakos, A. Nagrani, A. Zisserman, and D. Damen · 2019
Later among the works it cites.
Speech2face: Learning the face behind a voice
T.-H. Oh, T. Dekel, C. Kim, I. Mosseri, W. T. Freeman, M. Rubinstein, and W. Matusik · 2019
Later among the works it cites.
Disjoint mapping network for cross-modal matching of voices and faces
Y. Wen, M. A. Ismail, W. Liu, B. Raj, and R. Singh · 2019
Later among the works it cites.
Recursive visual sound separation using minus-plus net
X. Xu, B. Dai, and D. Lin · 2019
Later among the works it cites.
The sound of motions
H. Zhao, C. Gan, W.-C. Ma, and A. Torralba · 2019
Later among the works it cites.
Talking face generation by adversarially disentangled audio-visual representation
H. Zhou, Y. Liu, Z. Liu, P. Luo, and X. Wang · 2019
Later among the works it cites.
Self-supervised learning of audio-visual objects from video
T. Afouras, A. Owens, J.-S. Chung, and A. Zisserman · 2020
Later among the works it cites.
Spot the conversation: speaker diarisation in the wild
J. S. Chung, J. Huh, A. Nagrani, T. Afouras, and A. Zisserman · 2020
Later among the works it cites.
Facefilter: Audio-visual speech separation using still images
S.-W. Chung, S. Choe, J. S. Chung, and H.-G. Kang · 2020
Later among the works it cites.
Seeing voices and hearing voices: learning discriminative embeddings using cross-modal self-supervision
S.-W. Chung, H. G. Kang, and J. S. Chung · 2020
Later among the works it cites.
Music gesture for visual sound separation
C. Gan, D. Huang, H. Zhao, J. B. Tenenbaum, and A. Torralba · 2020
Later among the works it cites.
Visualechoes: Spatial image representation learning through echolocation
R. Gao, C. Chen, Z. Al-Halab, C. Schissler, and K. Grauman · 2020
Later among the works it cites.
Listen to look: Action recognition by previewing audio
R. Gao, T.-H. Oh, K. Grauman, and L. Torresani · 2020
Later among the works it cites.
Discriminative sounding objects localization via self-supervised audiovisual matching
D. Hu, R. Qian, M. Jiang, X. Tan, S. Wen, E. Ding, W. Lin, and D. Dou · 2020
Later among the works it cites.
Towards practical lipreading with distilled and efficient models
P. Ma, B. Martinez, S. Petridis, and M. Pantic · 2020
Later among the works it cites.
Lipreading using temporal convolutional networks
B. Martinez, P. Ma, S. Petridis, and M. Pantic · 2020
Later among the works it cites.
Voice separation with an unknown number of multiple speakers
E. Nachmani, Y. Adi, and L. Wolf · 2020
Later among the works it cites.
Disentangled speech embeddings using cross-modal self-supervision
A. Nagrani, J. S. Chung, S. Albanie, and A. Zisserman · 2020
Later among the works it cites.
Sep-stereo: Visually guided stereophonic audio generation by associating source separation
H. Zhou, X. Xu, D. Lin, X. Wang, and Z. Liu · 2020
Later among the works it cites.
Into the wild with audioscope: Unsupervised audio-visual separation of on-screen sounds
E. Tzinis, S. Wisdom, A. Jansen, S. Hershey, T. Remez, D. P. Ellis, and J. R. Hershey · 2021
Closest in time.