Fetching the paper…
Reading the bibliography…
Variational auto-encoders (VAEs) are deep generative latent variable models that can be used for learning the distribution of complex data.
“Visual contribution to speech intelligibility in noise,”
William Sumby and Irwin Pollack, · 1954
Earlier work this paper cites.
“Auditory-visual perception of speech,”
Norman Erber, · 1975
Earlier work this paper cites.
“Maximum likelihood from incomplete data via the EM algorithm,”
Arthur P. Dempster, Nan M. Laird, and Donald B. Rubin, · 1977
Earlier work this paper cites.
“Suppression of acoustic noise in speech using spectral subtraction,”
Steven Boll, · 1979
Earlier work this paper cites.
“Enhancement and bandwidth compression of noisy speech,”
Jae Soo Lim and Alan V Oppenheim, · 1979
Earlier work this paper cites.
Speech enhancement
Jae Soo Lim, · 1983
Earlier work this paper cites.
“Speech enhancement using a minimum-mean square error short-time spectral amplitude estimator,”
Yariv Ephraim and David Malah, · 1984
Earlier work this paper cites.
“Speech enhancement using a minimum mean-square error log-spectral amplitude estimator,”
Yariv Ephraim and David Malah, · 1985
Earlier work this paper cites.
“Quantifying the contribution of vision to speech perception in noise,”
Alison MacLeod and Quentin Summerfield, · 1987
Earlier work this paper cites.
“A Monte Carlo implementation of the EM algorithm and the poor man’s data augmentation algorithms,”
Greg C.G. Wei and Martin A. Tanner, · 1990
Earlier work this paper cites.
“TIMIT acoustic phonetic continuous speech corpus,”
John S. Garofolo, Lori F. Lamel, William M. Fisher, Jonathan G. Fiscus, David S. Pallett, Nancy L. Dahlgren, and Victor Zue, · 1993
Earlier work this paper cites.
“Noisy speech enhancement with filters estimated from the speaker’s lips,”
Laurent Girin, Gang Feng, and Jean-Luc Schwartz, · 1995
Earlier work this paper cites.
“An introduction to variational methods for graphical models,”
Michael I. Jordan, Zoubin Ghahramani, Tommi S. Jaakkola, and Lawrence K. Saul, · 1999
Earlier work this paper cites.
“Audio-visual enhancement of speech in noise,”
Laurent Girin, Jean-Luc Schwartz, and Gang Feng, · 2001
Earlier work this paper cites.
“Learning joint statistical models for audio-visual fusion and segregation,”
John W. Fisher III, Trevor Darrell, William T. Freeman, and Paul A. Viola, · 2001
Earlier work this paper cites.
“Speech enhancement for non-stationary noise environments,”
Israel Cohen and Baruch Berdugo, · 2001
Earlier work this paper cites.
“Perceptual evaluation of speech quality (PESQ)-a new method for speech quality assessment of telephone networks and codecs,”
Antony W. Rix, John G. Beerends, Michael P. Hollier, and Andries P. Hekstra, · 2001
Earlier work this paper cites.
“Audio-visual speech enhancement with AVCDCN (audio-visual codebook dependent cepstral normalization),”
Sabine Deligne, Gerasimos Potamianos, and Chalapathy Neti, · 2002
Earlier work this paper cites.
“Audio-visual sound separation via hidden Markov models,”
John R. Hershey and Michael Casey, · 2002
Earlier work this paper cites.
“Speech enhancement based on minimum mean-square error estimation and supergaussian priors,”
Rainer Martin, · 2005
Earlier work this paper cites.
Monte Carlo Statistical Methods
Christian P. Robert and George Casella, · 2005
Earlier work this paper cites.
“FaNT– filtering and noise adding tool,”
Hans-Günter Hirsch, · 2005
Earlier work this paper cites.
Speech enhancement
Jacob Benesty, Shoji Makino, and Jingdong Chen, · 2006
Earlier work this paper cites.
“An audiovisual corpus for speech perception and automatic speech recognition,”
Martin Cooke, Jon Barker, Stuart Cunningham, and Xu Shao, · 2006
Cited alongside, same era.
“Performance measurement in blind audio source separation,”
Emmanuel Vincent, Rémi Gribonval, and Cédric Févotte, · 2006
Cited alongside, same era.
Speech enhancement: theory and practice
Philipos C. Loizou, · 2007
Cited alongside, same era.
“Minimum mean-square error estimation of discrete Fourier coefficients with generalized Gamma priors,”
Jan Erkelens, Richard Hendriks, Richard Heusdens, and Jesper Jensen, · 2007
Cited alongside, same era.
“Supervised and semi-supervised separation of sounds from single-channel mixtures,”
Paris Smaragdis, Bhiksha Raj, and Madhusudana Shashanka, · 2007
Cited alongside, same era.
“Speech denoising using nonnegative matrix factorization with priors,”
“SNR-aware convolutional neural network modeling for speech enhancement.,”
Szu-Wei Fu, Yu Tsao, and Xugang Lu, · 2016
Later among the works it cites.
“NTCD-TIMIT: A new database and baseline for noise-robust audio-visual speech recognition,”
Ahmed Hussen Abdelaziz, · 2017
Later among the works it cites.
“Variational inference: A review for statisticians,”
David M. Blei, Alp Kucukelbir, and Jon D. McAuliffe, · 2017
Later among the works it cites.
“ β \beta -vae: Learning basic visual concepts with a constrained variational framework,”
Irina Higgins, Loic Matthey, Arka Pal, Christopher Burgess, Xavier Glorot, Matthew Botvinick, Shakir Mohamed, and Alexander Lerchner, · 2017
Later among the works it cites.
“The conversation: Deep audio-visual speech enhancement,”
Triantafyllos Afouras, Joon Son Chung, and Andrew Zisserman, · 2018
Later among the works it cites.
“Visual speech enhancement,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kevin Wilson, Bhiksha Raj, Paris Smaragdis, and Ajay Divakaran, · 2008
Cited alongside, same era.
“Nonnegative matrix factorization with the Itakura-Saito divergence: With application to music analysis,”
Cédric Févotte, Nancy Bertin, and Jean-Louis Durrieu, · 2009
Cited alongside, same era.
“Visually derived wiener filters for speech enhancement,”
Ibrahim Almajai and Ben Milner, · 2010
Cited alongside, same era.
“Phoneme-dependent NMF for speech enhancement in monaural mixtures,”
Bhiksha Raj, Rita Singh, and Tuomas Virtanen, · 2011
Cited alongside, same era.
“Algorithms for nonnegative matrix factorization with the β \beta -divergence,”
Cédric Févotte and Jérôme Idier, · 2011
Cited alongside, same era.
“A non-negative approach to semi-supervised separation of speech from noise with the use of temporal dynamics,”
Gautham J. Mysore and Paris Smaragdis, · 2011
Cited alongside, same era.
“An algorithm for intelligibility prediction of time–frequency weighted noisy speech,”
Cees H. Taal, Richard C. Hendriks, Richard Heusdens, and Jesper Jensen, · 2011
Cited alongside, same era.
Aviv Gabbay, Asaph Shamir, and Shmuel Peleg, · 2018
Later among the works it cites.
“Seeing through noise: Speaker separation and enhancement using visually-derived speech,”
Aviv Gabbay, Ariel Ephart, Tavi Halperin, and Shmuel Peleg, · 2018
Later among the works it cites.
“Audio-visual speech enhancement using multimodal deep convolutional neural networks,”
Jen-Cheng Hou, Syu-Siang Wang, Ying-Hui Lai, Yu Tsao, Hsiu-Wen Chang, and Hsin-Min Wang, · 2018
Later among the works it cites.
“DNN driven speaker independent audio-visual mask estimation for speech separation,”
Mandar Gogate, Ahsan Adeel, Ricard Marxer, Jon Barker, and Amir Hussain, · 2018
Later among the works it cites.
“Statistical speech enhancement based on probabilistic integration of variational autoencoder and non-negative matrix factorization,”
Yoshiaki Bando, Masato Mimura, Katsutoshi Itoyama, Kazuyoshi Yoshii, and Tatsuya Kawahara, · 2018
Later among the works it cites.
“A variance modeling framework based on variational autoencoders for speech enhancement,”
Simon Leglaive, Laurent Girin, and Radu Horaud, · 2018
Later among the works it cites.
“Bayesian multichannel speech enhancement with a deep speech prior,”
Kouhei Sekiguchi, Yoshiaki Bando, Kazuyoshi Yoshii, and Tatsuya Kawahara, · 2018
Later among the works it cites.
“Supervised speech separation based on deep learning: An overview,”
DeLiang Wang and Jitong Chen, · 2018
Later among the works it cites.
“End-to-end audiovisual speech recognition,”
Stavros Petridis, Themos Stafylakis, Pingchuan Ma, Feipeng Cai, Georgios Tzimiropoulos, and Maja Pantic, · 2018
Later among the works it cites.
“Semi-supervised multichannel speech enhancement with variational autoencoders and non-negative matrix factorization,”
Simon Leglaive, Laurent Girin, and Radu Horaud, · 2019
Closest in time.
“Speech enhancement with variational autoencoders and alpha-stable distributions,”
Simon Leglaive, Umut Şimşekli, Antoine Liutkus, Laurent Girin, and Radu Horaud, · 2019
Closest in time.
“A statistically principled and computationally efficient approach to speech enhancement using variational autoencoders,”
Manuel Pariente, Antoine Deleforge, and Emmanuel Vincent, · 2019
Closest in time.
“Multichannel speech enhancement based on time-frequency masking using subband long short-term memory,”
Xiaofei Li and Radu Horaud, · 2019
Closest in time.
“Supervised determined source separation with multichannel variational autoencoder,”
Hirokazu Kameoka, Li Li, Shota Inoue, and Shoji Makino, · 2019
Closest in time.
“Fast MVAE: Joint separation and classification of mixed sources based on multichannel variational autoencoder with auxiliary classifier,”
Li Li, Hirokazu Kameoka, and Shoji Makino, · 2019
Closest in time.
“Joint separation and dereverberation of reverberant mixtures with multichannel variational autoencoder,”
Shota Inoue, Hirokazu Kameoka, Li Li, Shogo Seki, and Shoji Makino, · 2019
Closest in time.
“A deep generative model of speech complex spectrograms,”
Aditya Arie Nugraha, Kouhei Sekiguchi, and Kazuyoshi Yoshii, · 2019
Closest in time.
“Noisy audio feature enhancement using audio-visual speech data,”
Roland Goecke, Gerasimos Potamianos, and Chalapathy Neti, · 2028
Closest in time.