Fetching the paper…
Reading the bibliography…
Denoising diffusion probabilistic models (DDPMs) and generative adversarial networks (GANs) are popular generative models for neural vocoders.
I. Yamada, M. Yukawa, and M. Yamagishi, Minimizing the Moreau envelope of nonsmooth convex functions over the fixed point set of certain quasi-nonexpansive mappings . Springer, 2011, pp. 345–390
2011
Earlier work this paper cites.
S. Theodoridis, K. Slavakis, and I. Yamada, “Adaptive learning in a world of projections,” IEEE Signal Process. Mag. , 2011
2011
Earlier work this paper cites.
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial nets,” in Proc. NeurIPS , 2014
2014
Earlier work this paper cites.
N. Parikh and S. Boyd, “Proximal algorithms,” Found. Trends Optim. , 2014
2014
Earlier work this paper cites.
D. J. Rezende and S. Mohamed, “Variational inference with normalizing flows,” in Proc. ICML , 2015
2015
Earlier work this paper cites.
D. P. Kingma and J. L. Ba, “Adam: A method for stochastic optimization,” in Proc. ICLR , 2015
2015
Earlier work this paper cites.
J. Sohl-Dickstein, E. Weiss, N. Maheswaranathan, and S. Ganguli, “Deep unsupervised learning using nonequilibrium thermodynamic,” in Proc. ICML , 2015
2015
Earlier work this paper cites.
2016
Earlier work this paper cites.
A. Tamamori, T. Hayashi, K. Kobayashi, K. Takeda, and T. Toda, “Speaker-dependent WaveNet vocoder.” in Proc. Interspeech , 2017
2017
Earlier work this paper cites.
H. H. Bauschke and P. L. Combettes, Convex Analysis and Monotone Operator Theory in Hilbert Spaces . Springer, 2017
2017
Earlier work this paper cites.
S. Mehri, K. Kumar, I. Gulrajani, R. Kumar, S. Jain, and J. Sotelo, “SampleRNN: An unconditional end-to-end neural audio,” in Proc. ICLR , 2018
2018
Earlier work this paper cites.
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Skerrv-Ryan, R. A. Saurous, Y. Agiomvrgiannakis, and Y. Wu, “Natural TTS synthesis by conditioning WaveNet on mel spectrogram predictions,” in Proc. ICASSP , 2018
2018
Earlier work this paper cites.
W. B. Kleijn, F. S. C. Lim, A. Luebs, J. Skoglund, F. Stimberg, Q. Wang, and T. C. Walters, “WaveNet based low rate speech coding,” in Proc. ICASSP , 2018
2018
Earlier work this paper cites.
T. Yoshimura, K. Hashimoto, K. Oura, Y. Nankaku, and K. Tokuda, “WaveNet-based zero-delay lossless speech coding,” in Proc. SLT , 2018
2018
Earlier work this paper cites.
N. Kalchbrenner, W. Elsen, K. Simonyan, S. Noury, N. Casagrande, W. Lockhart, F. Stimberg, A. van den Oord, S. Dieleman, and K. Kavukcuoglu, “Efficient neural audio synthesis,” in Proc. ICML , 2018
2018
Earlier work this paper cites.
A. van den Oord, Y. Li, I. Babuschkin, K. Simonyan, O. Vinyals, K. Kavukcuoglu, G. van den Driessche, E. Lockhart, L. C. Cobo, F. Stimberg, N. Casagrande, D. Grewe, S. Noury, S. Dieleman, E. Elsen, N. Kalchbrenner, H. Zen, A. Graves, H. King, T. Walters, D. Belov, and D. Hassabis, “Parallel WaveNet: Fast high-fidelity speech synthesis.” in Proc. ICML , 2018
2018
Earlier work this paper cites.
G. T. Buzzard, S. H. Chan, S. Sreehari, and C. A. Bouman, “Plug-and-play unplugged: Optimization-free reconstruction using consensus equilibrium,” SIAM J. Imaging Sci. , 2018
2018
Earlier work this paper cites.
R. Prenger, R. Valle, and B. Catanzaro, “WaveGlow: A flow-based generative network for speech synthesis,” in Proc. ICASSP , 2019
2019
Earlier work this paper cites.
Y. Ren, Y. Ruan, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T.-Y. Liu, “FastSpeech: Fast, robust and controllable text to speech,” in Proc. NeurIPS , 2019
2019
Earlier work this paper cites.
Y. Jia, R. J. Weiss, F. Biadsy, W. Macherey, M. Johnson, Z. Chen, and Y. Wu, “Direct speech-to-speech translation with a sequence-to-sequence model,” in Proc. Interspeech , 2019
2019
Earlier work this paper cites.
S. Maiti and M. I. Mandel, “Parametric resynthesis with neural vocoders,” in Proc. IEEE WASPAA , 2019
2019
Earlier work this paper cites.
J.-M. Valin and J. Skoglund, “A real-time wideband neural vocoder at 1.6kb/s using LPCNet,” in Proc. Interspeech , 2019
2019
Earlier work this paper cites.
J.-M. Valin and J. Skoglund, “LPCNet: Improving neural speech synthesis through linear prediction,” in Proc. ICASSP , 2019
2019
Earlier work this paper cites.
C. Donahue, J. McAuley, and M. Puckette, “Adversarial audio synthesis,” in Proc. ICLR , 2019
2019
Cited alongside, same era.
K. Kumar, R. Kumar, T. de Boissiere, L. Gestin, W. Z. Teoh, J. Sotelo, A. de Brébisson, Y. Bengio, and A. C. Courville, “MelGAN: Generative adversarial networks for conditional waveform synthesis,” in Proc. Adv. Neural Inf. Process. Syst. (NeurIPS) , 2019
2019
Cited alongside, same era.
E. Ryu, J. Liu, S. Wang, X. Chen, Z. Wang, and W. Yin, “Plug-and-play methods provably converge with properly trained denoisers,” in Proc. ICML , 2019
2019
Cited alongside, same era.
Y. Masuyama, K. Yatabe, Y. Koizumi, Y. Oikawa, and N. Harada, “Deep Griffin-Lim iteration,” in Proc. ICASSP , 2019
2019
Cited alongside, same era.
H. Zen, R. Clark, R. J. Weiss, V. Dang, Y. Jia, Y. Wu, Y. Zhang, and Z. Chen, “LibriTTS: A corpus derived from LibriSpeech for text-to-speech,” in Proc. Interspeech , 2019
W. Jang, D. Lim, J. Yoon, B. Kim, and J. Kim, “UnivNet: A neural vocoder with multi-resolution spectrogram discriminators for high-fidelity waveform generation,” in Proc. Interspeech , 2021
2021
Later among the works it cites.
N. Chen, Y. Zhang, H. Zen, R. J. Weiss, M. Norouzi, and W. Chan, “WaveGrad: Estimating gradients for waveform generation,” in Proc. ICLR , 2021
2021
Later among the works it cites.
Z. Kong, W. Ping, J. Huang, K. Zhao, and B. Catanzaro, “DiffWave: A versatile diffusion model for audio synthesis,” in Proc. ICLR , 2021
2021
Later among the works it cites.
T. Okamoto, T. Toda, Y. Shiga, and H. Kawai, “Noise level limited sub-modeling for diffusion probabilistic vocoders,” in Proc. ICASSP , 2021
2021
Later among the works it cites.
P. L. Combettes and J.-C. Pesquet, “Fixed point strategies in data science,” IEEE Trans. Signal Process. , 2021
2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2019
Cited alongside, same era.
W. Ping, K. Peng, K. Zhao, and Z. Song, “WaveFlow: A compact flow-based model for raw audio,” in Proc. ICML , 2020
2020
Cited alongside, same era.
2020
Cited alongside, same era.
——, “Speaker independence of neural vocoders and their effect on parametric resynthesis speech enhancement,” in Proc. ICASSP , 2020
2020
Cited alongside, same era.
J. Su, Z. Jin, and A. Finkelstein, “HiFi-GAN: High-fidelity denoising and dereverberation based on speech deep features in adversarial networks,” in Proc. Interspeech , 2020
2020
Cited alongside, same era.
J. Kong, J. Kim, and J. Bae, “HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis,” in Proc. NeurIPS , 2020
2020
Cited alongside, same era.
R. Yamamoto, E. Song, and J.-M. Kim, “Parallel WaveGAN: A fast waveform generation model based on generative adversarial networks with multi-resolution spectrogram,” in Proc. ICASSP , 2020
2020
Cited alongside, same era.
J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” in Proc. NeurIPS , 2020
2020
Cited alongside, same era.
Later among the works it cites.
J.-C. Pesquet, A. Repetti, M. Terris, and Y. Wiaux, “Learning maximally monotone operators for image recovery,” SIAM J. Imaging Sci. , 2021
2021
Later among the works it cites.
——, “Deep Griffin–Lim iteration: Trainable iterative phase reconstruction using neural network,” IEEE J. Sel. Top. Signal Process. , 2021
2021
Later among the works it cites.
R. Cohen, M. Elad, and P. Milanfar, “Regularization by denoising via fixed-point projection (RED-PRO),” SIAM J. Imaging Sci. , 2021
2021
Later among the works it cites.
W.-C. Huang, S.-W. Yang, T. Hayashi, and T. Toda, “A comparative study of self-supervised speech representation based voice conversion,” EEE J. Sel. Top. Signal Process. , 2022
2022
Closest in time.
Y. Jia, M. T. Ramanovich, T. Remez, and R. Pomerantz, “Translatotron 2: High-quality direct speech-to-speech translation with voice preservation,” in Proc. ICML , 2022
2022
Closest in time.
A. Lee, P.-J. Chen, C. Wang, J. Gu, S. Popuri, X. Ma, A. Polyak, Y. Adi, Q. He, Y. Tang, J. Pino, and W.-N. Hsu, “Direct speech-to-speech translation with discrete units,” in Proc. 60th Annu. Meet. Assoc. Comput. Linguist. (Vol. 1: Long Pap.) , 2022
2022
Closest in time.
T. Saeki, S. Takamichi, T. Nakamura, N. Tanji, and H. Saruwatari, “SelfRemaster: Self-supervised speech restoration with analysis-by-synthesis approach using channel modeling,” in Proc. Interspeech , 2022
2022
Closest in time.
N. Zeghidour, A. Luebs, A. Omran, J. Skoglund, and M. Tagliasacchi, “SoundStream: An end-to-end neural audio codec,” IEEE/ACM Trans. Audio, Speech and Lang. Proc. , 2022
2022
Closest in time.
T. Kaneko, K. Tanaka, H. Kameoka, and S. Seki, “iSTFTNet: Fast and lightweight mel-spectrogram vocoder incorporating inverse short-time Fourier transform,” in Proc. ICASSP , 2022
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
M. W. Y. Lam, J. Wang, D. Su, and D. Yu, “BDDM: Bilateral denoising diffusion models for fast and high-quality speech synthesis,” in Proc. ICLR , 2022
2022
Closest in time.
S. Lee, H. Kim, C. Shin, X. Tan, C. Liu, Q. Meng, T. Qin, W. Chen, S. Yoon, and T.-Y. Liu, “PriorGrad: Improving conditional denoising diffusion models with data-dependent adaptive prior,” in Proc. ICLR , 2022
2022
Closest in time.
Y. Koizumi, H. Zen, K. Yatabe, N. Chen, and M. Bacchiani, “SpecGrad: Diffusion probabilistic model based neural vocoder with adaptive noise spectral shaping,” in Proc. Interspeech , 2022
2022
Closest in time.
2022
Closest in time.
Z. Chen, X. Tan, K. Wang, S. Pan, D. Mandic, L. He, and S. Zhao, “InferGrad: Improving diffusion models for vocoder by considering inference in training,” in Proc. ICASSP , 2022
2022
Closest in time.
Z. Xiao, K. Kreis, and A. Vahdat, “Tackling the generative learning trilemma with denoising diffusion GANs,” in Proc. ICLR , 2022
2022
Closest in time.