Fetching the paper…
Reading the bibliography…
Dynamical variational autoencoders (DVAEs) are a class of deep generative models with latent variables, dedicated to model time series of high-dimensional data.
S. Boll, “Suppression of acoustic noise in speech using spectral subtraction,” IEEE Trans. Acoust., Speech, Signal Process. , vol. 27, no. 2, pp. 113–120, 1979
1979
Earlier work this paper cites.
J. S. Lim and A. V. Oppenheim, “Enhancement and bandwidth compression of noisy speech,” Proc. IEEE , vol. 67, no. 12, pp. 1586–1604, 1979
1979
Earlier work this paper cites.
Y. Ephraim and D. Malah, “Speech enhancement using a minimum-mean square error short-time spectral amplitude estimator,” IEEE Trans. Acoust., Speech, Signal Process. , vol. 32, no. 6, pp. 1109–1121, 1984
1984
Earlier work this paper cites.
B. D. Van Veen and K. M. Buckley, “Beamforming: A versatile approach to spatial filtering,” IEEE Acoust., Speech, Signal Process. Magazine , vol. 5, no. 2, pp. 4–24, 1988
1988
Earlier work this paper cites.
R. J. Williams and D. Zipser, “A learning algorithm for continually running fully recurrent neural networks,” Neural Comp. , vol. 1, no. 2, pp. 270–280, 1989
1989
Earlier work this paper cites.
F. D. Neeser and J. L. Massey, “Proper complex random processes with applications to information theory,” IEEE Trans. Inform. Theory , vol. 39, no. 4, pp. 1293–1302, 1993
1993
Earlier work this paper cites.
J. Garofolo, D. Graff, D. Paul, and D. Pallett, “CSR-I (WSJ0) Sennheiser LDC93S6B. https://catalog.ldc.upenn.edu/ldc93s6b,” Philadelphia: Linguistic Data Consortium , 1993
1993
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural Comp. , vol. 9, no. 8, pp. 1735–1780, 1997
1997
Earlier work this paper cites.
R. M. Neal and G. E. Hinton, “A view of the EM algorithm that justifies incremental, sparse, and other variants,” in Learning in graphical models . Springer, 1998, pp. 355–368
1998
Earlier work this paper cites.
A. Rix, J. Beerends, M. Hollier, and A. Hekstra, “Perceptual evaluation of speech quality (PESQ): A new method for speech quality assessment of telephone networks and codecs,” in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process. (ICASSP) , Salt Lake City, USA, 2001
2001
Earlier work this paper cites.
D. R. Hunter and K. Lange, “A tutorial on MM algorithms,” Am. Stat. , vol. 58, no. 1, pp. 30–37, 2004
2004
Earlier work this paper cites.
J. Benesty, S. Makino, and J. Chen, Speech enhancement . Springer Science & Business Media, 2006
2006
Earlier work this paper cites.
C. M. Bishop, Pattern Recognition and Machine Learning . Berlin: Springer-Verlag, 2006
2006
Earlier work this paper cites.
P. Smaragdis, B. Raj, and M. Shashanka, “Supervised and semi-supervised separation of sounds from single-channel mixtures,” in International Conference on Independent Component Analysis and Signal Separation , Charleston, USA, 2007
2007
Earlier work this paper cites.
J. Benesty, J. Chen, and Y. Huang, Microphone Array Signal Processing . Springer Science & Business Media, 2008
2008
Earlier work this paper cites.
M. J. Wainwright and M. I. Jordan, “Graphical models, exponential families, and variational inference,” Found. Trends Mach. Learn. , vol. 1, no. 1–2, p. 1–305, 2008
2008
Earlier work this paper cites.
C. Févotte, N. Bertin, and J.-L. Durrieu, “Nonnegative matrix factorization with the Itakura-Saito divergence: With application to music analysis,” Neural Comp. , vol. 21, no. 3, pp. 793–830, 2009
2009
Earlier work this paper cites.
G. J. Mysore and P. Smaragdis, “A non-negative approach to semi-supervised separation of speech from noise with the use of temporal dynamics,” in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process. (ICASSP) , Prague, Czech Republic, 2011
2011
Earlier work this paper cites.
E. Vincent, M. G. Jafari, S. A. Abdallah, M. D. Plumbley, and M. E. Davies, “Probabilistic modeling paradigms for audio source separation,” in Machine Audition: Principles, Algorithms and Systems . IGI global, 2011, pp. 162–185
2011
Earlier work this paper cites.
C. Févotte and J. Idier, “Algorithms for nonnegative matrix factorization with the β \beta -divergence,” Neural Comp. , vol. 23, no. 9, pp. 2421–2456, 2011
2011
Earlier work this paper cites.
ITU-R, “Recommendation BS.1770-4: Algorithms to measure audio programme loudness and true-peak audio level,” BS Series , 2011
2011
Earlier work this paper cites.
C. H. Taal, R. C. Hendriks, R. Heusdens, and J. Jensen, “An algorithm for intelligibility prediction of time–frequency weighted noisy speech,” IEEE Trans. Audio, Speech, Lang. Process. , vol. 19, no. 7, pp. 2125–2136, 2011
2011
Earlier work this paper cites.
P. C. Loizou, Speech enhancement: Theory and practice . CRC press, 2013
2013
Earlier work this paper cites.
X. Lu, Y. Tsao, S. Matsuda, and C. Hori, “Speech enhancement based on deep denoising autoencoder.” in Proc. Interspeech Conf. , Lyon, France, 2013
2013
Earlier work this paper cites.
N. Mohammadiha, P. Smaragdis, and A. Leijon, “Supervised and unsupervised speech enhancement using nonnegative matrix factorization,” IEEE Trans. Audio, Speech, Lang. Process. , vol. 21, no. 10, pp. 2140–2151, 2013
2013
Earlier work this paper cites.
D. P. Kingma, S. Mohamed, D. Jimenez Rezende, and M. Welling, “Semi-supervised learning with deep generative models,” in Advances Neural Inform. Process. Systems (NeurIPS) , Montreal, Canada, 2014
2014
Earlier work this paper cites.
D. P. Kingma and M. Welling, “Auto-encoding variational Bayes,” in Proc. Int. Conf. Learn. Repres. (ICLR) , Banff, Canada, 2014
2014
Earlier work this paper cites.
D. J. Rezende, S. Mohamed, and D. Wierstra, “Stochastic backpropagation and approximate inference in deep generative models,” in Proc. Int. Conf. Mach. Learn. (ICML) , Beijing, China, 2014
2014
Cited alongside, same era.
O. Fabius and J. R. van Amersfoort, “Variational recurrent auto-encoders,” Int. Conf. Learn. Repres. (ICLR) workshop , 2014
2014
Cited alongside, same era.
2014
Cited alongside, same era.
2015
Cited alongside, same era.
J. Chung, K. Kastner, L. Dinh, K. Goel, A. Courville, and Y. Bengio, “A recurrent latent variable model for sequential data,” in Advances Neural Inform. Process. Systems (NeurIPS) , Montreal, Canada, 2015
P.-A. Mattei and J. Frellsen, “Refit your encoder when new data comes by,” in NeurIPS Workshop on Bayesian Deep Learning , Montreal, Canada, 2018
2018
Later among the works it cites.
2018
Later among the works it cites.
S.-W. Fu, C.-F. Liao, Y. Tsao, and S.-D. Lin, “MetricGAN: Generative adversarial networks based black-box metric scores optimization for speech enhancement,” in Proc. Int. Conf. Mach. Learn. (ICML) , Long Beach, CA, 2019
2019
Later among the works it cites.
S. Leglaive, U. Şimşekli, A. Liutkus, L. Girin, and R. Horaud, “Speech enhancement with variational autoencoders and alpha-stable distributions,” in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process. (ICASSP) , Brighton, UK, 2019
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2015
Cited alongside, same era.
D. Dean, A. Kanagasundaram, H. Ghaemmaghami, M. H. Rahman, and S. Sridharan, “The QUT-NOISE-SRE protocol for the evaluation of noisy speaker recognition,” in Proc. Interspeech Conf. , Dresden, Germany, 2015
2015
Cited alongside, same era.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in Proc. Int. Conf. Learn. Repres. (ICLR) , San Diego, USA, 2015
2015
Cited alongside, same era.
S. Bengio, O. Vinyals, N. Jaitly, and N. Shazeer, “Scheduled sampling for sequence prediction with recurrent neural networks,” in Advances Neural Inform. Process. Systems (NeurIPS) , Montreal, Canada, 2015
2015
Cited alongside, same era.
C. Lea, R. Vidal, A. Reiter, and G. D. Hager, “Temporal convolutional networks: A unified approach to action segmentation,” in Proc. Europ. Conf. Computer Vision (ECCV) , Amsterdam, The Netherlands, 2016
2016
Cited alongside, same era.
M. Fraccaro, S. K. Sønderby, U. Paquet, and O. Winther, “Sequential neural models with stochastic layers,” in Advances Neural Inform. Process. Systems (NeurIPS) , Barcelona, Spain, 2016
2016
Cited alongside, same era.
C. Valentini-Botinhao, X. Wang, S. Takaki, and J. Yamagishi, “Investigating RNN-based speech enhancement methods for noise-robust text-to-speech.” in Speech Synth. Workshop , Sunnyvale, CA, 2016
2016
Cited alongside, same era.
C. K. Sønderby, T. Raiko, L. Maaløe, S. K. Sønderby, and O. Winther, “Ladder variational autoencoders,” in Advances Neural Inform. Process. Systems (NeurIPS) , Barcelona, Spain, 2016
2016
Cited alongside, same era.
M. Pariente, A. Deleforge, and E. Vincent, “A statistically principled and computationally efficient approach to speech enhancement using variational autoencoders,” in Proc. Interspeech Conf. , Graz, Austria, 2019
2019
Later among the works it cites.
S. Leglaive, L. Girin, and R. Horaud, “Semi-supervised multichannel speech enhancement with variational autoencoders and non-negative matrix factorization,” in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process. (ICASSP) , Brighton, UK, 2019
2019
Later among the works it cites.
M. Fontaine, A. A. Nugraha, R. Badeau, K. Yoshii, and A. Liutkus, “Cauchy multichannel speech enhancement with a deep speech prior,” in Proc. Europ. Signal Process. Conf. (EUSIPCO) , A Coruna, Spain, 2019
2019
Later among the works it cites.
J. Le Roux, S. Wisdom, H. Erdogan, and J. R. Hershey, “SDR: Half-baked or well done?” in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process. (ICASSP) , Brighton, UK, 2019
2019
Later among the works it cites.
F.-R. Stöter, S. Uhlich, A. Liutkus, and Y. Mitsufuji, “Open-Unmix: A reference implementation for music source separation,” Journal of Open Source Software , vol. 4, no. 41, p. 1667, 2019
2019
Later among the works it cites.
Y. Xiang and C. Bao, “A parallel-data-free speech enhancement method using multi-objective learning cycle-consistent generative adversarial network,” IEEE/ACM Trans. Audio, Speech, Lang. Process. , vol. 28, pp. 1826–1838, 2020
2020
Later among the works it cites.
Y. Bando, K. Sekiguchi, and K. Yoshii, “Adaptive neural speech enhancement with a denoising variational autoencoder.” in Proc. Interspeech Conf. , Shanghai, China, 2020, pp. 2437–2441
2020
Later among the works it cites.
S. Leglaive, X. Alameda-Pineda, L. Girin, and R. Horaud, “A recurrent variational autoencoder for speech enhancement,” in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process. (ICASSP) , Barcelona, Spain, 2020
2020
Later among the works it cites.
J. Richter, G. Carbajal, and T. Gerkmann, “Speech enhancement with stochastic temporal convolutional networks,” in Proc. Interspeech Conf. , Shanghai, China, 2020
2020
Later among the works it cites.
S. Uhlich and Y. Mitsufuji, “Open-Unmix for speech enhancement (UMX SE),” May 2020. [Online]. Available: https://doi.org/10.5281/zenodo.3786908
2020
Later among the works it cites.
A. Vahdat and J. Kautz, “NVAE: A deep hierarchical variational autoencoder,” in Advances Neural Inform. Process. Systems (NeurIPS) , Vancouver, Canada, 2020
2020
Later among the works it cites.
S.-W. Fu, C. Yu, T.-A. Hsieh, P. Plantinga, M. Ravanelli, X. Lu, and Y. Tsao, “MetricGAN+: An improved version of MetricGAN for speech enhancement,” in Proc. Interspeech Conf. , Brno, Czech Republik, 2021
2021
Closest in time.
G. Yu, Y. Wang, C. Zheng, H. Wang, and Q. Zhang, “CycleGAN-based non-parallel speech enhancement with an adaptive attention-in-attention mechanism,” in Asia-Pacific Signal Inform. Process. Assoc. Annual Conf. (APSIPA) , 2021
2021
Closest in time.
N. Alamdari, A. Azarang, and N. Kehtarnavaz, “Improving deep speech denoising by noisy2noisy signal mapping,” Applied Acoustics , vol. 172, p. 107631, 2021
2021
Closest in time.
M. M. Kashyap, A. Tambwekar, K. Manohara, and S. Natarajan, “Speech denoising without clean training data: a noise2noise approach,” in Proc. Interspeech Conf. , Brno, Czech Republic, 2021
2021
Closest in time.
T. Fujimura, Y. Koizumi, K. Yatabe, and R. Miyazaki, “Noisy-target training: A training strategy for DNN-based speech enhancement without clean speech,” in Proc. Europ. Signal Process. Conf. (EUSIPCO) , Dublin, Ireland (virtual conference), 2021
2021
Closest in time.
C. K. Reddy, V. Gopal, and R. Cutler, “DNSMOS: A non-intrusive perceptual objective speech quality metric to evaluate noise suppressors,” in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process. (ICASSP) , Toronto, Canada, 2021
2021
Closest in time.
G. Carbajal, J. Richter, and T. Gerkmann, “Guided variational autoencoder for speech enhancement with a supervised classifier,” in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process. (ICASSP) , Toronto, Canada, 2021
2021
Closest in time.
H. Fang, G. Carbajal, S. Wermter, and T. Gerkmann, “Variational autoencoder for speech enhancement with a noise-aware encoder,” in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process. (ICASSP) , Toronto, Canada, 2021, pp. 676–680
2021
Closest in time.
L. Girin, S. Leglaive, X. Bie, J. Diard, T. Hueber, and X. Alameda-Pineda, “Dynamical variational autoencoders: A comprehensive review,” Found. Trends Mach. Learn. , vol. 15, no. 1-2, pp. 1–175, 2021
2021
Closest in time.
X. Bie, L. Girin, S. Leglaive, T. Hueber, and X. Alameda-Pineda, “A benchmark of dynamical variational autoencoders applied to speech spectrogram modeling,” in Proc. Interspeech Conf. , Brno, Czech Republic, 2021
2021
Closest in time.
S.-W. Fu, C. Yu, K.-H. Hung, M. Ravanelli, and Y. Tsao, “MetricGAN-U: Unsupervised speech enhancement/dereverberation based only on noisy/reverberated speech,” in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process. (ICASSP) , 2022
2022
Closest in time.