Fetching the paper…
Reading the bibliography…
This paper proposes an approach to the joint modeling of the short-time Fourier transform magnitude and phase spectrograms with a deep generative model.
“On information and sufficiency,”
S. Kullback and R. A. Leibler, · 1951
Earlier work this paper cites.
“Short term spectral analysis, synthesis, and modification by discrete Fourier transform,”
J. Allen, · 1977
Earlier work this paper cites.
“Signal estimation from modified short-time fourier transform,”
D. W. Griffin and J. S. Lim, · 1984
Earlier work this paper cites.
“Learning representations by back-propagating errors,”
D. E. Rumelhart, G. E. Hinton, and R. J. Williams, · 1986
Earlier work this paper cites.
“Estimating and interpreting the instantaneous frequency of a signal – Part 1: Fundamentals,”
B. Boashash, · 1992
Earlier work this paper cites.
“P.862 : Perceptual evaluation of speech quality (PESQ): An objective method for end-to-end speech quality assessment of narrow-band telephone networks and speech codecs,” 2001
ITU-T, · 2001
Earlier work this paper cites.
“P.862.1 : Mapping function for transforming P.862 raw result scores to MOS-LQO,” 2003
ITU-T, · 2003
Earlier work this paper cites.
“CSR-I (WSJ0) Complete LDC93S6A,” DVD, 2007,
J. Garofalo, D. Graff, D. Paul, and D. Pallett, · 2007
Earlier work this paper cites.
“Explicit consistency constraints for stft spectrograms and their application to phase reconstruction,”
J. le Roux, N. Ono, and S. Sagayama, · 2008
Earlier work this paper cites.
Discrete-Time Signal Processing
Alan V. Oppenheim and Ronald W. Schafer, · 2009
Earlier work this paper cites.
Speech Processing in Modern Communication: Challenges and Perspectives
I. Cohen, J. Benesty, and S. Gannot, Eds., · 2010
Earlier work this paper cites.
“An algorithm for intelligibility prediction of time-frequency weighted noisy speech,”
C. H. Taal, R. C. Hendriks, R. Heusdens, and J. Jensen, · 2011
Earlier work this paper cites.
“Early stopping – but when?,”
L. Prechelt, · 2012
Earlier work this paper cites.
“On the difficulty of training recurrent neural networks,”
R. Pascanu, T. Mikolov, and Y. Bengio, · 2013
Cited alongside, same era.
“Auto-encoding variational Bayes,”
D. P. Kingma and M. Welling, · 2014
Cited alongside, same era.
“Generative adversarial nets,”
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, · 2014
Cited alongside, same era.
“Bayesian estimation of clean speech spectral coefficients given a priori knowledge of the phase,”
T. Gerkmann, · 2014
Cited alongside, same era.
“Phase processing for single-channel speech enhancement: History and recent advances,”
T. Gerkmann, M. Krawczyk-Becker, and J. le Roux, · 2015
Cited alongside, same era.
“Variational inference with normalizing flows,”
D. Rezende and S. Mohamed, · 2015
Cited alongside, same era.
“Language modeling with gated convolutional networks,”
Y. N. Dauphin, A. Fan, M. Auli, and D. Grangier, · 2017
Later among the works it cites.
“An analysis of environment, microphone and data simulation mismatches in robust speech recognition,”
E. Vincent, S. Watanabe, A. A. Nugraha, J. Barker, and R. Marxer, · 2017
Later among the works it cites.
Audio Source Separation and Speech Enhancement
E. Vincent, T. Virtanen, and S. Gannot, Eds., · 2018
Later among the works it cites.
“Model-based STFT phase recovery for audio source separation,”
P. Magron, R. Badeau, and B. David, · 2018
Later among the works it cites.
“On modeling the STFT phase of audio signals with the von Mises distribution,”
P. Magron and T. Virtanen, · 2018
Later among the works it cites.
“Phase reconstruction from amplitude spectrograms based on von-Mises-distribution deep neural network,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“U-net: Convolutional networks for biomedical image segmentation,”
O. Ronneberger, P. Fischer, and T. Brox, · 2015
Cited alongside, same era.
“Adam: A method for stochastic optimization,”
D. P. Kingma and J. Ba, · 2015
Cited alongside, same era.
“Advances in phase-aware signal processing in speech communication,”
P. Mowlaee, R. Saeidi, and Y. Stylianou, · 2016
Cited alongside, same era.
“WaveNet: A generative model for raw audio,”
A. van den Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. Senior, and K. Kavukcuoglu, · 2016
Cited alongside, same era.
“Weight normalization: A simple reparameterization to accelerate training of deep neural networks,”
T. Salimans and D. P. Kingma, · 2016
Cited alongside, same era.
“The one hundred layers tiramisu: Fully convolutional densenets for semantic segmentation,”
S. Jegou, M. Drozdzal, D. Vazquez, A. Romero, and Y. Bengio, · 2017
Cited alongside, same era.
S. Takamichi, Y. Saito, N. Takamune, D. Kitamura, and H. Saruwatari, · 2018
Later among the works it cites.
“PhaseNet: Discretized phase modeling with deep neural networks for audio source separation,”
N. Takahashi, P. Agrawal, N. Goswami, and Y. Mitsufuji, · 2018
Later among the works it cites.
“Semi-blind source separation with multichannel variational autoencoder,”
H. Kameoka, L. Li, S. Inoue, and S. Makino, · 2018
Later among the works it cites.
“Generalized multichannel variational autoencoder for underdetermined source separation,”
S. Seki, H. Kameoka, L. Li, T. Toda, and K. Takeda, · 2018
Later among the works it cites.
“Statistical speech enhancement based on probabilistic integration of variational autoencoder and non-negative matrix factorization,”
Y. Bando, M. Mimura, K. Itoyama, K. Yoshii, and T. Kawahara, · 2018
Later among the works it cites.
“A variance modeling framework based on variational autoencoders for speech enhancement,”
S. Leglaive, L. Girin, and R. Horaud, · 2018
Later among the works it cites.
“Bayesian multichannel speech enhancement with a deep speech prior,”
K. Sekiguchi, Y. Bando, K. Yoshii, and T. Kawahara, · 2018
Later among the works it cites.
“An empirical evaluation of generic convolutional and recurrent networks for sequence modeling,”
S. Bai, J. Z. Kolter, and V. Koltun, · 2018
Later among the works it cites.