Fetching the paper…
Reading the bibliography…
Speech enhancement and speech separation are two related tasks, whose purpose is to extract either one or more target speech signals, respectively, from a mixture of sounds generated by several sources.
E. Lombard, “Le signe de l’elevation de la voix,” Annales des Maladies de L’Oreille et du Larynx , vol. 37, no. 2, pp. 101–119, 1911
1911
Earlier work this paper cites.
H. Fletcher and J. Steinberg, “Articulation testing methods,” The Bell System Technical Journal , vol. 8, no. 4, pp. 806–854, 1929
1929
Earlier work this paper cites.
S. S. Stevens, J. Volkmann, and E. B. Newman, “A scale for the measurement of the psychological magnitude pitch,” The Journal of the Acoustical Society of America , vol. 8, no. 3, pp. 185–190, 1937
1937
Earlier work this paper cites.
J. P. Egan, “Articulation testing methods,” The Laryngoscope , vol. 58, no. 9, pp. 955–991, 1948
1948
Earlier work this paper cites.
H. Robbins and S. Monro, “A stochastic approximation method,” The Annals of Mathematical Statistics , pp. 400–407, 1951
1951
Earlier work this paper cites.
J. Kiefer, J. Wolfowitz et al. , “Stochastic estimation of the maximum of a regression function,” The Annals of Mathematical Statistics , vol. 23, no. 3, pp. 462–466, 1952
1952
Earlier work this paper cites.
E. C. Cherry, “Some experiments on the recognition of speech, with one and with two ears,” The Journal of the Acoustical Society of America , vol. 25, no. 5, pp. 975–979, 1953
1953
Earlier work this paper cites.
W. H. Sumby and I. Pollack, “Visual contribution to speech intelligibility in noise,” The Journal of the Acoustical Society of America , vol. 26, no. 2, 1954
1954
Earlier work this paper cites.
G. A. Miller and P. E. Nicely, “An analysis of perceptual confusions among some English consonants,” The Journal of the Acoustical Society of America , vol. 27, no. 2, pp. 338–352, 1955
1955
Earlier work this paper cites.
G. Fairbanks, “Test of phonemic differentiation: The rhyme test,” The Journal of the Acoustical Society of America , vol. 30, no. 7, pp. 596–600, 1958
1958
Earlier work this paper cites.
A. S. House, C. E. Williams, M. H. Hecker, and K. D. Kryter, “Articulation-testing methods: consonantal differentiation with a closed-response set,” The Journal of the Acoustical Society of America , vol. 37, no. 1, pp. 158–166, 1965
1965
Earlier work this paper cites.
J. Capon, “High-resolution frequency-wavenumber spectrum analysis,” Proceedings of the IEEE , vol. 57, no. 8, pp. 1408–1418, 1969
1969
Earlier work this paper cites.
H. McGurk and J. MacDonald, “Hearing lips and seeing voices,” Nature , vol. 264, no. 5588, pp. 746–748, 1976
1976
Earlier work this paper cites.
D. N. Kalikow, K. N. Stevens, and L. L. Elliott, “Development of a test of speech intelligibility in noise using sentence materials with controlled word predictability,” The Journal of the Acoustical Society of America , vol. 61, no. 5, pp. 1337–1351, 1977
1977
Earlier work this paper cites.
B. D. Lucas and T. Kanade, “An iterative image registration technique with an application to stereo vision,” in Proc. of IJCAI , 1981
1981
Earlier work this paper cites.
B. Hagerman, “Sentences for testing speech intelligibility in noise,” Scandinavian Audiology , vol. 11, no. 2, pp. 79–87, 1982
1982
Earlier work this paper cites.
W. D. Voiers, “Evaluating processed speech using the diagnostic rhyme test,” Speech Technology , pp. 30–39, 1983
1983
Earlier work this paper cites.
Y. Ephraim and D. Malah, “Speech enhancement using a minimum-mean square error short-time spectral amplitude estimator,” IEEE Transactions on Acoustics, Speech, and Signal Processing , vol. 32, no. 6, pp. 1109–1121, 1984
1984
Earlier work this paper cites.
D. Griffin and J. Lim, “Signal estimation from modified short-time fourier transform,” IEEE Transactions on Acoustics, Speech, and Signal Processing , vol. 32, no. 2, pp. 236–243, 1984
1984
Earlier work this paper cites.
D. E. Rumelhart, G. E. Hinton, and R. J. Williams, “Learning representations by back-propagating errors,” Nature , vol. 323, no. 6088, pp. 533–536, 1986
1986
Earlier work this paper cites.
G. Cybenko, “Approximation by superpositions of a sigmoidal function,” Mathematics of Control, Signals and Systems , vol. 2, no. 4, pp. 303–314, 1989
1989
Earlier work this paper cites.
K. Hornik, M. Stinchcombe, H. White et al. , “Multilayer feedforward networks are universal approximators.” Neural Networks , vol. 2, no. 5, pp. 359–366, 1989
1989
Earlier work this paper cites.
Y. LeCun, “Generalization and network design strategies,” Connectionism in Perspective , vol. 19, pp. 143–155, 1989
1989
Earlier work this paper cites.
ITU-R, “Recommendation BS.562: Subjective assessment of sound quality,” 1990
1990
Earlier work this paper cites.
P. J. Werbos, “Backpropagation through time: what it does and how to do it,” Proceedings of the IEEE , vol. 78, no. 10, pp. 1550–1560, 1990
1990
Earlier work this paper cites.
K. Hornik, “Approximation capabilities of multilayer feedforward networks,” Neural Networks , vol. 4, no. 2, pp. 251–257, 1991
1991
Earlier work this paper cites.
C. Tomasi and T. Kanade, “Detection and tracking of point features,” Technical Report CMU-CS-91-132 , 1991
1991
Earlier work this paper cites.
J. F. Feuerstein, “Monaural versus binaural hearing: Ease of listening, word recognition, and attentional effort,” Ear and Hearing , vol. 13, no. 2, pp. 80–86, 1992
1992
Earlier work this paper cites.
Q. Summerfield, “Lipreading and audio-visual speech perception,” Philosophical Transactions of the Royal Society of London. Series B: Biological Sciences , vol. 335, no. 1273, pp. 71–78, 1992
1992
Earlier work this paper cites.
Y. Bengio, P. Simard, and P. Frasconi, “Learning long-term dependencies with gradient descent is difficult,” IEEE Transactions on Neural Networks , vol. 5, no. 2, pp. 157–166, 1994
1994
Earlier work this paper cites.
M. Nilsson, S. D. Soli, and J. A. Sullivan, “Development of the hearing in noise test for the measurement of speech reception thresholds in quiet and in noise,” The Journal of the Acoustical Society of America , vol. 95, no. 2, pp. 1085–1099, 1994
1994
Earlier work this paper cites.
ITU-T, “P.830 : Subjective performance assessment of telephone-band and wideband digital codecs,” 1996
1996
Earlier work this paper cites.
R. Caruana, “Multitask learning,” Machine learning , vol. 28, no. 1, pp. 41–75, 1997
1997
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural Computation , vol. 9, no. 8, pp. 1735–1780, 1997
1997
Earlier work this paper cites.
A. N. S. Institute, American National Standard: Methods for Calculation of the Speech Intelligibility Index . Acoustical Society of America, 1997
1997
Earlier work this paper cites.
M. Schuster and K. K. Paliwal, “Bidirectional recurrent neural networks,” IEEE Transactions on Signal Processing , vol. 45, no. 11, pp. 2673–2681, 1997
1997
Earlier work this paper cites.
——, “Recommendation BT.1359-1: Relative timing of sound and vision for broadcasting,” 1998
1998
Earlier work this paper cites.
H. Yehia, P. Rubin, and E. Vatikiotis-Bateson, “Quantitative association of vocal-tract and facial behavior,” Speech Communication , vol. 26, no. 1, 1998
1998
Earlier work this paper cites.
J. P. Barker and F. Berthommier, “Evidence of correlation between acoustic and visual features of speech,” in Proc. of ICPhS , 1999
1999
Earlier work this paper cites.
H. Kawahara, I. Masuda-Katsuse, and A. De Cheveigne, “Restructuring speech representations using a pitch-adaptive time–frequency smoothing and an instantaneous-frequency-based F0 extraction: Possible role of a repetitive structure in sounds,” Speech Communication , vol. 27, no. 3-4, pp. 187–207, 1999
1999
Earlier work this paper cites.
S. Partan and P. Marler, “Communication goes multimodal,” Science , vol. 283, no. 5406, pp. 1272–1273, 1999
1999
Earlier work this paper cites.
K. Wagener, T. Brand, and B. Kollmeier, “Entwicklung und evaluation eines satztests in deutscher sprache - Teil II: Optimierung des Oldenburger satztests,” Zeitschrift für Audiologie , no. 38, pp. 44–56, 1999
1999
Earlier work this paper cites.
——, “Entwicklung und evaluation eines satztests in deutscher sprache - Teil III: Evaluierung des Oldenburger satztests,” Zeitschrift für Audiologie , no. 38, pp. 86–95, 1999
1999
Earlier work this paper cites.
K. Wagener, V. Kühnel, and B. Kollmeier, “Entwicklung und evaluation eines satztests in deutscher sprache - Teil I: Design des Oldenburger satztests,” Zeitschrift für Audiologie , no. 38, pp. 4–15, 1999
1999
Earlier work this paper cites.
A. W. Bronkhorst, “The cocktail party phenomenon: A review of research on speech intelligibility in multiple-talker conditions,” Acta Acustica united with Acustica , vol. 86, no. 1, pp. 117–128, 2000
2000
Earlier work this paper cites.
T. Darrell, J. W. Fisher, and P. Viola, “Audio-visual segmentation and “the cocktail party effect”,” in Proc. of ICMI , 2000
2000
Earlier work this paper cites.
J. R. Deller, J. H. L. Hansen, and J. G. Proakis, Discrete-Time Processing of Speech Signals . Wiley-IEEE Press, 2000
2000
Earlier work this paper cites.
F. A. Gers, J. Schmidhuber, and F. Cummins, “Learning to forget: Continual prediction with LSTM,” Neural Computation , vol. 12, no. 10, pp. 2451–2471, 2000
2000
Earlier work this paper cites.
T. F. Cootes, G. J. Edwards, and C. J. Taylor, “Active appearance models,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 23, no. 6, pp. 681–685, 2001
2001
Earlier work this paper cites.
L. Girin, J.-L. Schwartz, and G. Feng, “Audio-visual enhancement of speech in noise,” The Journal of the Acoustical Society of America , vol. 109, no. 6, pp. 3007–3020, 2001
2001
Earlier work this paper cites.
——, “Recommendation P.862: Perceptual evaluation of speech quality (PESQ): an objective method for end-to-end speech quality assessment of narrow-band telephone networks and speech codecs,” 2001
2001
Earlier work this paper cites.
A. W. Rix, J. G. Beerends, M. P. Hollier, and A. P. Hekstra, “Perceptual evaluation of speech quality (PESQ) - A new method for speech quality assessment of telephone networks and codecs,” in Proc. of ICASSP , 2001
2001
Earlier work this paper cites.
E. K. Patterson, S. Gurbuz, Z. Tufekci, and J. N. Gowdy, “CUAVE: A new audio-visual database for multimodal human-computer interface research,” in Proc. of ICASSP , 2002
2002
Earlier work this paper cites.
J.-L. Schwartz, F. Berthommier, and C. Savariaux, “Audio-visual scene analysis: Evidence for a “very-early” integration process in audio-visual speech perception,” in Proc. of ICSLP - Interspeech , 2002
2002
Earlier work this paper cites.
D. Sodoyer, J.-L. Schwartz, L. Girin, J. Klinkisch, and C. Jutten, “Separation of audio-visual speech sources: A new approach exploiting the audio-visual coherence of speech stimuli,” EURASIP Journal on Advances in Signal Processing , no. 11, pp. 1165–1173, 2002
2002
Earlier work this paper cites.
——, “Recommendation BS.1534-1: Method for the subjective assessment of intermediate quality levels of coding systems,” 2003
2003
Earlier work this paper cites.
——, “Recommendation P.835: Subjective test methodology for evaluating speech communication systems that include noise suppression algorithm,” 2003
2003
Earlier work this paper cites.
——, “Recommendation P.862.1: Mapping function for transforming P.862 raw result scores to MOS-LQO,” 2003
2003
Earlier work this paper cites.
C. Sanderson and K. K. Paliwal, “Noise compensation in a person verification system using face and multiple speech features,” Pattern Recognition , vol. 36, no. 2, pp. 293–302, 2003
2003
Earlier work this paper cites.
M. L. Hawley, R. Y. Litovsky, and J. F. Culling, “The benefit of binaural hearing in a cocktail party: Effect of location and type of interferer,” The Journal of the Acoustical Society of America , vol. 115, no. 2, pp. 833–843, 2004
2004
Earlier work this paper cites.
J. Kates and K. Arehart, “Coherence and the speech intelligibility index,” The Journal of the Acoustical Society of America , vol. 115, no. 5, pp. 2604–2604, 2004
2004
Earlier work this paper cites.
C. T. Kello and D. C. Plaut, “A neural network model of the articulatory-acoustic forward mapping trained on recordings of articulatory parameters,” The Journal of the Acoustical Society of America , vol. 116, no. 4, pp. 2354–2364, 2004
2004
Earlier work this paper cites.
D. Sodoyer, L. Girin, C. Jutten, and J.-L. Schwartz, “Developing an audio-visual speech source separation algorithm,” Speech Communication , vol. 44, no. 1-4, pp. 113–125, 2004
2004
Earlier work this paper cites.
P. Viola and M. J. Jones, “Robust real-time face detection,” International Journal of Computer Vision , vol. 57, no. 2, pp. 137–154, 2004
2004
Earlier work this paper cites.
——, “Recommendation P.862.2: Wideband extension to recommendation P.862 for the assessment of wideband telephone networks and speech codecs,” 2005
2005
Earlier work this paper cites.
K. S. Rhebergen and N. J. Versfeld, “A speech intelligibility index-based approach to predict the speech reception threshold for sentences in fluctuating noise for normal-hearing listeners,” The Journal of the Acoustical Society of America , vol. 117, no. 4, pp. 2181–2192, 2005
2005
Earlier work this paper cites.
V. Vaillancourt, C. Laroche, C. Mayer, C. Basque, M. Nali, A. Eriks-Brophy, S. D. Soli, and C. Giguère, “Adaptation of the HINT (hearing in noise test) for adult Canadian Francophone populations,” International Journal of Audiology , vol. 44, no. 6, pp. 358–361, 2005
2005
Earlier work this paper cites.
L. L. Wong and S. D. Soli, “Development of the Cantonese hearing in noise test (CHINT),” Ear and Hearing , vol. 26, no. 3, pp. 276–289, 2005
2005
Earlier work this paper cites.
I. Almajai, B. Milner, and J. Darch, “Analysis of correlation between audio and visual speech features for clean audio feature prediction in noise,” in Proc. of Interspeech , 2006
2006
Earlier work this paper cites.
J. Chen, J. Benesty, Y. Huang, and S. Doclo, “New insights into the noise reduction Wiener filter,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 14, no. 4, pp. 1218–1234, 2006
2006
Earlier work this paper cites.
M. Cooke, J. Barker, S. Cunningham, and X. Shao, “An audio-visual corpus for speech perception and automatic speech recognition,” The Journal of the Acoustical Society of America , vol. 120, no. 5, pp. 2421–2424, 2006
2006
Earlier work this paper cites.
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, “Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,” in Proc. of ICML , 2006
2006
Earlier work this paper cites.
M. Hällgren, B. Larsby, and S. Arlinger, “A Swedish version of the hearing in noise test (HINT) for measurement of speech recognition,” International Journal of Audiology , vol. 45, no. 4, pp. 227–237, 2006
2006
Earlier work this paper cites.
U. Jekosch, Voice and Speech Quality Perception: Assessment and Evaluation . Springer Science & Business Media, 2006
2006
Earlier work this paper cites.
B. Rivet, L. Girin, and C. Jutten, “Mixing audiovisual speech processing and blind source separation for the extraction of speech signals from convolutive mixtures,” IEEE transactions on audio, speech, and language processing , vol. 15, no. 1, pp. 96–108, 2006
2006
Earlier work this paper cites.
E. Vincent, R. Gribonval, and C. Févotte, “Performance measurement in blind audio source separation,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 14, no. 4, pp. 1462–1469, 2006
2006
Earlier work this paper cites.
D. L. Wang and G. J. Brown, Computational Auditory Scene Analysis: Principles, Algorithms, and Applications . Wiley-IEEE Press, 2006
2006
Earlier work this paper cites.
Z. Barzelay and Y. Y. Schechner, “Harmony in motion,” in Proc. of CVPR , 2007
2007
Earlier work this paper cites.
Y. Hu and P. C. Loizou, “Evaluation of objective quality measures for speech enhancement,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 16, no. 1, pp. 229–238, 2007
2007
Earlier work this paper cites.
K. Kumar, T. Chen, and R. M. Stern, “Profile view lip reading,” in Proc. of ICASSP , 2007
2007
Earlier work this paper cites.
H. K. Maganti, D. Gatica-Perez, and I. McCowan, “Speech enhancement and recognition in meetings with an audio–visual sensor array,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 15, no. 8, pp. 2257–2269, 2007
2007
Earlier work this paper cites.
——, “Visual voice activity detection as a help for speech source separation from convolutive mixtures,” Speech Communication , vol. 49, no. 7-8, pp. 667–677, 2007
2007
Earlier work this paper cites.
P. Lichtsteiner, C. Posch, and T. Delbruck, “A 128 × \times 128 120 dB 15 μ \mu s latency asynchronous temporal contrast vision sensor,” IEEE Journal of Solid-State Circuits , vol. 43, no. 2, pp. 566–576, 2008
2008
Earlier work this paper cites.
B. G. Shinn-Cunningham and V. Best, “Selective attention in normal and impaired hearing,” Trends in Amplification , vol. 12, no. 4, pp. 283–299, 2008
2008
Earlier work this paper cites.
D. E. King, “Dlib-ml: A machine learning toolkit,” Journal of Machine Learning Research , vol. 10, pp. 1755–1758, 2009
2009
Earlier work this paper cites.
U. Kjems, J. B. Boldt, M. S. Pedersen, T. Lunner, and D. Wang, “Role of mask pattern in intelligibility of ideal binary-masked noisy speech,” The Journal of the Acoustical Society of America , vol. 126, no. 3, pp. 1415–1426, 2009
2009
Earlier work this paper cites.
J. B. Nielsen and T. Dau, “Development of a Danish speech intelligibility test,” International Journal of Audiology , vol. 48, no. 10, pp. 729–741, 2009
2009
Earlier work this paper cites.
C. Richie, S. Warburton, and M. Carter, Audiovisual database of spoken American English . Linguistic Data Consortium, 2009
2009
Earlier work this paper cites.
G. Zhao, M. Barnard, and M. Pietikainen, “Lipreading with local spatiotemporal descriptors,” IEEE Transactions on Multimedia , vol. 11, no. 7, pp. 1254–1265, 2009
2009
Earlier work this paper cites.
I. Almajai and B. Milner, “Visually derived Wiener filters for speech enhancement,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 19, no. 6, pp. 1642–1651, 2010
2010
Earlier work this paper cites.
A. L. Casanovas, G. Monaci, P. Vandergheynst, and R. Gribonval, “Blind audiovisual source separation based on sparse redundant representations,” IEEE Transactions on Multimedia , vol. 12, no. 5, pp. 358–371, 2010
2010
Earlier work this paper cites.
N. Dehak, P. J. Kenny, R. Dehak, P. Dumouchel, and P. Ouellet, “Front-end factor analysis for speaker verification,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 19, no. 4, pp. 788–798, 2010
2010
Earlier work this paper cites.
B. Denby, T. Schultz, K. Honda, T. Hueber, J. M. Gilbert, and J. S. Brumberg, “Silent speech interfaces,” Speech Communication , vol. 52, no. 4, pp. 270–287, 2010
2010
Earlier work this paper cites.
X. Glorot and Y. Bengio, “Understanding the difficulty of training deep feedforward neural networks,” in Proc. of AISTATS , 2010
2010
Earlier work this paper cites.
C. Jorgensen and S. Dusan, “Speech interfaces based upon surface electromyography,” Speech Communication , vol. 52, no. 4, pp. 354–366, 2010
2010
Earlier work this paper cites.
J. M. Kates and K. H. Arehart, “The hearing-aid speech quality index (HASQI),” Journal of the Audio Engineering Society , vol. 58, no. 5, pp. 363–381, 2010
2010
Earlier work this paper cites.
S. M. Naqvi, M. Yu, and J. A. Chambers, “A multimodal approach to blind source separation of moving sources,” IEEE Journal of Selected Topics in Signal Processing , vol. 4, no. 5, pp. 895–910, 2010
2010
Earlier work this paper cites.
E. Ozimek, A. Warzybok, and D. Kutzner, “Polish sentence matrix test for speech intelligibility measurement in noise,” International Journal of Audiology , vol. 49, no. 6, pp. 444–454, 2010
2010
Earlier work this paper cites.
A. A. Zekveld, S. E. Kramer, and J. M. Festen, “Pupil response as an indication of effortful listening: The influence of sentence intelligibility,” Ear and Hearing , vol. 31, no. 4, pp. 480–490, 2010
2010
Earlier work this paper cites.
H. Brumm and S. A. Zollinger, “The evolution of the Lombard effect: 100 years of psychoacoustic research,” Behaviour , vol. 148, no. 11-13, pp. 1173–1198, 2011
2011
Cited alongside, same era.
——, “Recommendation P.863: Perceptual objective listening quality assessment,” 2011
2011
Cited alongside, same era.
J. Ngiam, A. Khosla, M. Kim, J. Nam, H. Lee, and A. Y. Ng, “Multimodal deep learning,” in Proc. of ICML , 2011
2011
Cited alongside, same era.
——, “The Danish hearing in noise test,” International Journal of Audiology , vol. 50, no. 3, pp. 202–208, 2011
2011
Cited alongside, same era.
C. H. Taal, R. C. Hendriks, R. Heusdens, and J. Jensen, “An algorithm for intelligibility prediction of time–frequency weighted noisy speech,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 19, no. 7, 2011
2011
Cited alongside, same era.
A. Gabbay, A. Ephrat, T. Halperin, and S. Peleg, “Seeing through noise: Visually driven speaker separation and enhancement,” in Proc. of ICASSP , 2018
2018
Later among the works it cites.
A. Gabbay, A. Shamir, and S. Peleg, “Visual speech enhancement,” Proc. of Interspeech , 2018
2018
Later among the works it cites.
R. Gao, R. Feris, and K. Grauman, “Learning to separate object sounds by watching unlabeled video,” in Proc. of ECCV , 2018
2018
Later among the works it cites.
M. Gogate, A. Adeel, R. Marxer, J. Barker, and A. Hussain, “DNN driven speaker independent audio-visual mask estimation for speech separation,” in Proc. of Interspeech , 2018
2018
Later among the works it cites.
C. Gu, C. Sun, D. A. Ross, C. Vondrick, C. Pantofaru, Y. Li, S. Vijayanarasimhan, G. Toderici, S. Ricco, R. Sukthankar et al. , “Ava: A video dataset of spatio-temporally localized atomic visual actions,” in Proc. of CVPR , 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Hines, J. Skoglund, A. Kokaram, and N. Harte, “ViSQOL: The virtual speech quality objective listener,” in proc. of IWAENC , 2012
2012
Cited alongside, same era.
S. Hochmuth, T. Brand, M. A. Zokoll, F. Z. Castro, N. Wardenga, and B. Kollmeier, “A spanish matrix sentence test for assessing speech reception thresholds in noise,” International Journal of Audiology , vol. 51, no. 7, pp. 536–544, 2012
2012
Cited alongside, same era.
Y. Lan, B.-J. Theobald, and R. Harvey, “View independent computer lip-reading,” in Proc. of ICME , 2012
2012
Cited alongside, same era.
Y. Liang, S. M. Naqvi, and J. A. Chambers, “Audio video based fast fixed-point independent vector analysis for multisource separation in a room environment,” EURASIP Journal on Advances in Signal Processing , vol. 2012, no. 1, p. 183, 2012
2012
Cited alongside, same era.
S. M. Naqvi, W. Wang, M. S. Khan, M. Barnard, and J. A. Chambers, “Multimodal (audio–visual) source separation exploiting multi-speaker tracking, robust beamforming and time–frequency masking,” IET Signal Processing , vol. 6, no. 5, pp. 466–477, 2012
2012
Cited alongside, same era.
T. Tieleman and G. Hinton, “Lecture 6.5 - RmsProp: Divide the gradient by a running average of its recent magnitude,” COURSERA: Neural Networks for Machine Learning, 2012
2012
Cited alongside, same era.
R. Xia, J. Li, M. Akagi, and Y. Yan, “Evaluation of objective intelligibility prediction measures for noise-reduced signals in Mandarin,” in Proc. of ICASSP , 2012
2012
Cited alongside, same era.
2018
Later among the works it cites.
J.-C. Hou, S.-S. Wang, Y.-H. Lai, Y. Tsao, H.-W. Chang, and H.-M. Wang, “Audio-visual speech enhancement using multimodal deep convolutional neural networks,” IEEE Transactions on Emerging Topics in Computational Intelligence , vol. 2, no. 2, pp. 117–128, 2018
2018
Later among the works it cites.
J. Hu, L. Shen, and G. Sun, “Squeeze-and-excitation networks,” in Proc. of CVPR , 2018
2018
Later among the works it cites.
Y. Jia, Y. Zhang, R. Weiss, Q. Wang, J. Shen, F. Ren, P. Nguyen, R. Pang, I. L. Moreno, Y. Wu et al. , “Transfer learning from speaker verification to multispeaker text-to-speech synthesis,” in Proc. of NeurIPS , 2018
2018
Later among the works it cites.
F. U. Khan, B. P. Milner, and T. Le Cornu, “Using visual speech information in masking methods for audio speaker separation,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 26, no. 10, pp. 1742–1754, 2018
2018
Later among the works it cites.
C. Kim, H. V. Shin, T.-H. Oh, A. Kaspar, M. Elgharib, and W. Matusik, “On learning associations of faces and voices,” in Proc. of ACCV , 2018
2018
Later among the works it cites.
Y. Kumar, M. Aggarwal, P. Nawal, S. Satoh, R. R. Shah, and R. Zimmermann, “Harnessing AI for speech reconstruction using multi-view silent video feed,” in Proc. of ACM-MM , 2018
2018
Later among the works it cites.
Y. Kumar, R. Jain, M. Salik, R. R. Shah, R. Zimmermann, and Y. Yin, “MyLipper: A personalized system for speech reconstruction using multi-view visual feeds,” in Proc. of ISM , 2018
2018
Later among the works it cites.
S. Leglaive, L. Girin, and R. Horaud, “A variance modeling framework based on variational autoencoders for speech enhancement,” in Proc. of MLSP , 2018
2018
Later among the works it cites.
K. Liu, Y. Li, N. Xu, and P. Natarajan, “Learn to combine modalities in multimodal deep learning,” Proc. of KDD BigMine , 2018
2018
Later among the works it cites.
S. R. Livingstone and F. A. Russo, “The Ryerson audio-visual database of emotional speech and song (RAVDESS): A dynamic, multimodal set of facial and vocal expressions in North American English,” PLOS ONE , vol. 13, no. 5, 2018
2018
Later among the works it cites.
R. Lu, Z. Duan, and C. Zhang, “Listen and look: Audio–visual matching assisted speech source separation,” IEEE Signal Processing Letters , vol. 25, no. 9, pp. 1315–1319, 2018
2018
Later among the works it cites.
Y. Luo, Z. Chen, and N. Mesgarani, “Speaker-independent speech separation with deep attractor network,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 26, no. 4, pp. 787–796, 2018
2018
Later among the works it cites.
M. Morise and Y. Watanabe, “Sound quality comparison among high-quality vocoders by using re-synthesized speech,” Acoustical Science and Technology , vol. 39, no. 3, pp. 263–265, 2018
2018
Later among the works it cites.
A. Owens and A. A. Efros, “Audio-visual scene analysis with self-supervised multisensory features,” in Proc. of ECCV , 2018
2018
Later among the works it cites.
A. Pandey and D. Wang, “On adversarial training and loss functions for speech enhancement,” in Proc. of ICASSP , 2018
2018
Later among the works it cites.
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Skerrv-Ryan et al. , “Natural TTS synthesis by conditioning WaveNet on mel spectrogram predictions,” in Proc. of ICASSP , 2018
2018
Later among the works it cites.
T. M. F. Taha and A. Hussain, “A survey on techniques for enhancing speech,” International Journal of Computer Applications , vol. 179, no. 17, pp. 1–14, 2018
2018
Later among the works it cites.
D. L. Wang and J. Chen, “Supervised speech separation based on deep learning: An overview,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , 2018
2018
Later among the works it cites.
Z.-Q. Wang, J. Le Roux, and J. R. Hershey, “Multi-channel deep clustering: Discriminative spectral and spatial embeddings for speaker-independent speech separation,” in Proc. of ICASSP , 2018
2018
Later among the works it cites.
D. Ward, H. Wierstorf, R. D. Mason, E. M. Grais, and M. D. Plumbley, “BSS Eval or PEASS? Predicting the perception of singing-voice separation,” in Proc. of ICASSP , 2018
2018
Later among the works it cites.
C. Zhang, K. Koishida, and J. H. Hansen, “Text-independent speaker verification based on triplet convolutional neural network embeddings,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 26, no. 9, pp. 1633–1644, 2018
2018
Later among the works it cites.
H. Zhao, C. Gan, A. Rouditchenko, C. Vondrick, J. McDermott, and A. Torralba, “The sound of pixels,” in Proc. of ECCV , 2018
2018
Later among the works it cites.
A. Adeel, J. Ahmad, H. Larijani, and A. Hussain, “A novel real-time, lightweight chaotic-encryption scheme for next-generation audio-visual hearing aids,” Cognitive Computation , vol. 12, no. 3, pp. 589–601, 2019
2019
Later among the works it cites.
A. Adeel, M. Gogate, A. Hussain, and W. M. Whitmer, “Lip-reading driven deep learning approach for speech enhancement,” IEEE Transactions on Emerging Topics in Computational Intelligence , 2019
2019
Later among the works it cites.
——, “My lips are concealed: Audio-visual speech enhancement through obstructions,” in Proc. of Interspeech , 2019
2019
Later among the works it cites.
2019
Later among the works it cites.
R. Gao and K. Grauman, “2.5D visual sound,” in Proc. of CVPR , 2019
2019
Later among the works it cites.
——, “Co-separating sounds of visual objects,” in Proc. of ICCV , 2019
2019
Later among the works it cites.
2019
Later among the works it cites.
R. Gu, L. Chen, S.-X. Zhang, J. Zheng, Y. Xu, M. Yu, D. Su, Y. Zou, and D. Yu, “Neural spatial filter: Target speaker speech separation assisted with directional information,” in Proc. of Interspeech , 2019
2019
Later among the works it cites.
2019
Later among the works it cites.
J. Hu, Y. Zhang, and T. Okatani, “Visualization of convolutional neural networks for monocular depth estimation,” in Proc. of ICCV , 2019
2019
Later among the works it cites.
E. Ideli, “Audio-visual speech processing using deep learning techniques.” MSc thesis, Applied Sciences: School of Engineering Science, 2019
2019
Later among the works it cites.
E. Ideli, B. Sharpe, I. V. Bajić, and R. G. Vaughan, “Visually assisted time-domain speech enhancement,” in Proc. of GlobalSIP , 2019
2019
Later among the works it cites.
B. İnan, M. Cernak, H. Grabner, H. P. Tukuljac, R. C. Pena, and B. Ricaud, “Evaluating audiovisual source separation in the context of video conferencing,” Proc. of Interspeech , 2019
2019
Later among the works it cites.
——, “Recommendation BS.1284-2: General methods for the subjective assessment of sound quality,” 2019
2019
Later among the works it cites.
Y. Kumar, R. Jain, K. M. Salik, R. R. Shah, Y. Yin, and R. Zimmermann, “Lipper: Synthesizing thy speech using multi-view lipreading,” in Proc. of AAAI , 2019
2019
Later among the works it cites.
J. Le Roux, S. Wisdom, H. Erdogan, and J. R. Hershey, “SDR–half-baked or well done?” in Proc. of ICASSP , 2019
2019
Later among the works it cites.
——, “Audio–visual deep clustering for speech separation,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 27, no. 11, pp. 1697–1712, 2019
2019
Later among the works it cites.
Y. Luo and N. Mesgarani, “Conv-TasNet: Surpassing ideal time–frequency magnitude masking for speech separation,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 27, no. 8, pp. 1256–1266, 2019
2019
Later among the works it cites.
Y. Luo, J. Wang, X. Wang, L. Wen, and L. Wang, “Audio-visual speech separation using i-Vectors,” in Proc. of ICICSP , 2019
2019
Later among the works it cites.
D. Michelsanti, Z.-H. Tan, S. Sigurdsson, and J. Jensen, “On training targets and objective functions for deep-learning-based audio-visual speech enhancement,” in Proc. of ICASSP , 2019
2019
Later among the works it cites.
D. Michelsanti, Z.-H. Tan, S. Sigurdsson, and J. Jensen, “Deep-learning-based audio-visual speech enhancement in presence of Lombard effect,” Speech Communication , vol. 115, pp. 38–50, 2019
2019
Later among the works it cites.
——, “Effects of Lombard reflex on the performance of deep-learning-based audio-visual speech enhancement systems,” in Proc. of ICASSP , 2019
2019
Later among the works it cites.
G. Morrone, S. Bergamaschi, L. Pasa, L. Fadiga, V. Tikhanoff, and L. Badino, “Face landmark-based speaker-independent audio-visual speech enhancement in multi-talker environments,” in Proc. of ICASSP , 2019
2019
Later among the works it cites.
T. Ochiai, M. Delcroix, K. Kinoshita, A. Ogawa, and T. Nakatani, “Multimodal SpeakerBeam: Single channel target speech extraction with audio-visual speaker clues,” Proc. Interspeech , 2019
2019
Later among the works it cites.
T.-H. Oh, T. Dekel, C. Kim, I. Mosseri, W. T. Freeman, M. Rubinstein, and W. Matusik, “Speech2Face: Learning the face behind a voice,” in Proc. of CVPR , 2019
2019
Later among the works it cites.
S. Parekh, A. Ozerov, S. Essid, N. Q. Duong, P. Pérez, and G. Richard, “Identify, locate and separate: Audio-visual object extraction in large video collections using weak supervision,” in Proc. of WASPAA , 2019
2019
Later among the works it cites.
J. Rincón-Trujillo and D. M. Córdova-Esparza, “Analysis of speech separation methods based on deep learning,” International Journal of Computer Applications , vol. 148, no. 9, pp. 21–29, 2019
2019
Later among the works it cites.
A. Rouditchenko, H. Zhao, C. Gan, J. McDermott, and A. Torralba, “Self-supervised audio-visual co-segmentation,” in Proc. of ICASSP , 2019
2019
Later among the works it cites.
2019
Later among the works it cites.
Y. Takashima, T. Takiguchi, and Y. Ariki, “Exemplar-based lip-to-speech synthesis using convolutional neural networks,” in Proc. of IW-FCV , 2019
2019
Later among the works it cites.
S. Uttam, Y. Kumar, D. Sahrawat, M. Aggarwal, R. R. Shah, D. Mahata, and A. Stent, “Hush-hush speak: Speech reconstruction using silent videos,” in Proc. of Interspeech , 2019
2019
Later among the works it cites.
K. Vougioukas, P. Ma, S. Petridis, and M. Pantic, “Video-driven speech reconstruction using generative adversarial networks,” in Proc. of Interspeech , 2019
2019
Later among the works it cites.
Q. Wang, H. Muckenhirn, K. Wilson, P. Sridhar, Z. Wu, J. R. Hershey, R. A. Saurous, R. J. Weiss, Y. Jia, and I. L. Moreno, “VoiceFilter: Targeted voice separation by speaker-conditioned spectrogram masking,” Proc. Interspeech , 2019
2019
Later among the works it cites.
G. Wichern, J. Antognini, M. Flynn, L. R. Zhu, E. McQuinn, D. Crow, E. Manilow, and J. Le Roux, “WHAM!: Extending speech separation to noisy environments,” Proc. of Interspeech , 2019
2019
Later among the works it cites.
J. Wu, Y. Xu, S.-X. Zhang, L.-W. Chen, M. Yu, L. Xie, and D. Yu, “Time domain audio visual speech separation,” in Proc. of ASRU , 2019
2019
Later among the works it cites.
W. Xie, A. Nagrani, J. S. Chung, and A. Zisserman, “Utterance-level aggregation for speaker recognition in the wild,” in Proc. of ICASSP , 2019
2019
Later among the works it cites.
X. Xu, B. Dai, and D. Lin, “Recursive visual sound separation using minus-plus net,” in Proc. of ICCV , 2019
2019
Later among the works it cites.
H. Zhao, C. Gan, W.-C. Ma, and A. Torralba, “The sound of motions,” in Proc. of ICCV , 2019
2019
Later among the works it cites.
——, “Contextual deep learning-based audio-visual switching for speech enhancement in real-world environments,” Information Fusion , vol. 59, pp. 163–170, 2020
2020
Closest in time.
2020
Closest in time.
S.-Y. Chuang, Y. Tsao, C.-C. Lo, and H.-M. Wang, “Lite audio-visual speech enhancement,” in Proc. of Interspeech , 2020
2020
Closest in time.
S.-W. Chung, S. Choe, J. S. Chung, and H.-G. Kang, “FaceFilter: Audio-visual speech separation using still images,” Proc. of Interspeech , 2020
2020
Closest in time.
C. Gan, D. Huang, H. Zhao, J. B. Tenenbaum, and A. Torralba, “Music gesture for visual sound separation,” in Proc. of CVPR , 2020
2020
Closest in time.
M. Gogate, K. Dashtipour, A. Adeel, and A. Hussain, “Cochleanet: A robust language-independent audio-visual model for speech enhancement,” Information Fusion , vol. 63, pp. 273–285, 2020
2020
Closest in time.
R. Gu, S.-X. Zhang, Y. Xu, L. Chen, Y. Zou, and D. Yu, “Multi-modal multi-channel target speech separation,” IEEE Journal of Selected Topics in Signal Processing , 2020
2020
Closest in time.
C. Han, Y. Luo, and N. Mesgarani, “Real-time binaural speech separation with preserved spatial cues,” in Proc. of ICASSP , 2020
2020
Closest in time.
M. L. Iuzzolino and K. Koishida, “AV(SE) 2 : Audio-visual squeeze-excite speech enhancement,” in Proc. of ICASSP , 2020
2020
Closest in time.
H. R. V. Joze, A. Shaban, M. L. Iuzzolino, and K. Koishida, “MMTM: Multimodal transfer module for CNN fusion,” Proc. of CVPR , 2020
2020
Closest in time.
C. Li and Y. Qian, “Deep audio-visual speech separation with attention mechanism,” in Proc. of ICASSP , 2020
2020
Closest in time.
Y. Li, Z. Liu, Y. Na, Z. Wang, B. Tian, and Q. Fu, “A visual-pilot deep fusion for target speech separation in multitalker noisy environment,” in Proc. of ICASSP , 2020
2020
Closest in time.
D. Michelsanti, O. Slizovskaia, G. Haro, E. Gómez, Z.-H. Tan, and J. Jensen, “Vocoder-based speech synthesis from silent videos,” in Proc. of Interspeech , 2020
2020
Closest in time.
S. Mun, S. Choe, J. Huh, and J. S. Chung, “The sound of my voice: speaker representation loss for target voice separation,” in Proc. of ICASSP , 2020
2020
Closest in time.
L. Pasa, G. Morrone, and L. Badino, “An analysis of speech enhancement and recognition losses in limited resources multi-talker single channel audio-visual ASR,” in Proc. of ICASSP , 2020
2020
Closest in time.
K. Prajwal, R. Mukhopadhyay, V. P. Namboodiri, and C. Jawahar, “Learning individual speaking styles for accurate lip to speech synthesis,” in Proc. of CVPR , 2020
2020
Closest in time.
L. Qu, C. Weber, and S. Wermter, “Multimodal target speech separation with voice and face references,” Proc. of Interspeech , 2020
2020
Closest in time.
J. Roth, S. Chaudhuri, O. Klejch, R. Marvin, A. Gallagher, L. Kaver, S. Ramaswamy, A. Stopczynski, C. Schmid, Z. Xi et al. , “Ava active speaker: An audio-visual dataset for active speaker detection,” in Proc. of ICASSP , 2020
2020
Closest in time.
——, “Robust unsupervised audio-visual speech enhancement using a mixture of variational autoencoders,” in Proc. of ICASSP , 2020
2020
Closest in time.
M. Sadeghi, S. Leglaive, X. Alameda-Pineda, L. Girin, and R. Horaud, “Audio-visual speech enhancement using conditional variational auto-encoders,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 28, pp. 1788–1800, 2020
2020
Closest in time.
2020
Closest in time.
Z. Sun, Y. Wang, and L. Cao, “An attention based speaker-independent audio-visual deep learning model for speech enhancement,” in Proc. of MMM , 2020
2020
Closest in time.
K. Tan, Y. Xu, S.-X. Zhang, M. Yu, and D. Yu, “Audio-visual speech separation and dereverberation with a two-stage multimodal network,” IEEE Journal of Selected Topics in Signal Processing , vol. 14, no. 3, pp. 542–553, 2020
2020
Closest in time.
W. Wang, C. Xing, D. Wang, X. Chen, and F. Sun, “A robust audio-visual speech enhancement model,” in Proc. of ICASSP , 2020
2020
Closest in time.
Z.-Q. Wang, “Deep learning based array processing for speech separation, localization, and recognition,” Ph.D. dissertation, The Ohio State University, 2020
2020
Closest in time.
Y. Xu, M. Yu, S.-X. Zhang, L. Chen, C. Weng, J. Liu, and D. Yu, “Neural spatio-temporal beamformer for target speech separation,” Proc. of Interspeech , 2020
2020
Closest in time.
D. Yin, C. Luo, Z. Xiong, and W. Zeng, “PHASEN: A phase-and-harmonics-aware speech enhancement network,” in Proc. of AAAI , 2020
2020
Closest in time.
2020
Closest in time.
L. Zhu and E. Rahtu, “Separating sounds from a single image,” arXiv preprint arXiv:2007.07984 , 2020
2020
Closest in time.
——, “Visually guided sound source separation using cascaded opponent filter network,” Proc. of ACCV , 2020
2020
Closest in time.
A. Arriandiaga, G. Morrone, L. Pasa, L. Badino, and C. Bartolozzi, “Audio-visual target speaker extraction on multi-talker environment using event-driven cameras,” Proc. of ISCAS (to appear) , 2021
2021
Closest in time.
G. Morrone, D. Michelsanti, Z.-H. Tan, and J. Jensen, “Audio-visual speech inpainting with deep learning,” Proc. of ICASSP (to appear) , 2021
2021
Closest in time.
E. Zwicker and H. Fastl, Psychoacoustics: Facts and Models . Springer Science & Business Media, 2013, vol. 22
2021
Closest in time.