Fetching the paper…
Reading the bibliography…
Most neural vocoders employ band-limited mel-spectrograms to generate waveforms.
C. E. Shannon, “Communication in the Presence of Noise,” Proceedings of the IRE , vol. 37, no. 1, pp. 10–21, 1949
1949
Earlier work this paper cites.
D. Griffin and J. Lim, “Signal Estimation from Modified Short-Time Fourier Transform,” IEEE Transactions on Acoustics, Speech, and Signal Processing , vol. 32, no. 2, pp. 236–243, 1984
1984
Earlier work this paper cites.
L.-J. Liu, Z.-H. Ling, Y. Jiang, M. Zhou, and L.-R. Dai, “WaveNet vocoder with limited training data for voice conversion.” in Interspeech , 2018, pp. 1983–1987
1987
Earlier work this paper cites.
H. Kawahara, I. Masuda-Katsuse, and A. De Cheveigne, “Restructuring speech representations using a pitch-adaptive time-frequency smoothing and an instantaneous-frequency-based f0 extraction: Possible role of a repetitive structure in sounds,” Speech communication , vol. 27, no. 3-4, pp. 187–207, 1999
1999
Earlier work this paper cites.
A. W. Rix, J. G. Beerends, M. P. Hollier, and A. P. Hekstra, “Perceptual Evaluation of Speech Quality (PESQ)-A New Method for Speech Quality Assessment of Telephone Networks and Codecs,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2001
2001
Earlier work this paper cites.
A. L. Maas, A. Y. Hannun, and A. Y. Ng, “Rectifier Nonlinearities Improve Neural Network Acoustic Models,” in International Conference on Machine Learning (ICML) , 2013
2013
Earlier work this paper cites.
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative Adversarial Nets,” in Advances in Neural Information Processing Systems (NeurIPS) , 2014
2014
Earlier work this paper cites.
X.-L. Zhang and D. Wang, “Boosting Contextual Information for Deep Neural Network Based Voice Activity Detection,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 24, no. 2, pp. 252–264, 2015
2015
Earlier work this paper cites.
D. P. Kingma and J. Ba, “Adam: A Method for Stochastic Optimization,” in International Conference on Learning Representations (ICLR) , 2015
2015
Earlier work this paper cites.
M. Morise, F. Yokomori, and K. Ozawa, “WORLD: A Vocoder-Based High-Quality Speech Synthesis System for Real-Time Applications,” IEICE TRANSACTIONS on Information and Systems , vol. 99, no. 7, pp. 1877–1884, 2016
2016
Earlier work this paper cites.
A. van den Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. Senior, and K. Kavukcuoglu, “WaveNet: A Generative Model for Raw Audio,” in 9th ISCA Speech Synthesis Workshop , 2016
2016
Earlier work this paper cites.
A. Van den Oord, N. Kalchbrenner, L. Espeholt, O. Vinyals, A. Graves et al. , “Conditional Image Generation with PixelCNN Decoders,” in Advances in Neural Information Processing Systems (NeurIPS) , 2016
2016
Cited alongside, same era.
T. Salimans and D. P. Kingma, “Weight Normalization: A Simple Reparameterization to Accelerate Training of Deep Neural Networks,” in Advances in Neural Information Processing Systems (NeurIPS) , 2016
2016
Cited alongside, same era.
X. Mao, Q. Li, H. Xie, R. Y. Lau, Z. Wang, and S. Paul Smolley, “Least Squares Generative Adversarial Networks,” in Proceedings of the IEEE international conference on computer vision (ICCV) , 2017
2017
Cited alongside, same era.
K. Ito et al. , “The LJ speech dataset,” https://keithito.com/LJ-Speech-Dataset, 2017
2017
Cited alongside, same era.
H. Zen, V. Dang, R. Clark, Y. Zhang, R. J. Weiss, Y. Jia, Z. Chen, and Y. Wu, “LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech,” in Proceedings of the Annual Conference of the International Speech Communication Association (INTERSPEECH) , 2019
2019
Later among the works it cites.
D. Lim, W. Jang, G. O, H. Park, B. Kim, and J. Yoon, “JDI-T: Jointly trained Duration Informed Transformer for Text-To-Speech without Explicit Alignment,” in Proceedings of the Annual Conference of the International Speech Communication Association (INTERSPEECH) , 2020
2020
Later among the works it cites.
R. Yamamoto, E. Song, and J.-M. Kim, “Parallel WaveGAN: A Fast Waveform Generation Model based on Generative Adversarial Networks with Multi-Resolution Spectrogram,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2020
2020
Later among the works it cites.
J. Kong, J. Kim, and J. Bae, “HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis,” in Advances in Neural Information Processing Systems (NeurIPS) , 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Skerry-Ryan et al. , “Natural TTS Synthesis by Conditioning WaveNet on Mel Spectrogram Predictions,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2018
2018
Cited alongside, same era.
N. Kalchbrenner, E. Elsen, K. Simonyan, S. Noury, N. Casagrande, E. Lockhart, F. Stimberg, A. van den Oord, S. Dieleman, and K. Kavukcuoglu, “Efficient Neural Audio Synthesis,” in International Conference on Machine Learning (ICML) , 2018
2018
Cited alongside, same era.
W. Ping, K. Peng, and J. Chen, “ClariNet: Parallel Wave Generation in End-to-End Text-to-Speech,” in International Conference on Learning Representations (ICLR) , 2018
2018
Cited alongside, same era.
Y. Jia, R. J. Weiss, F. Biadsy, W. Macherey, M. Johnson, Z. Chen, and Y. Wu, “Direct speech-to-speech translation with a sequence-to-sequence model,” in Proceedings of the Annual Conference of the International Speech Communication Association (INTERSPEECH) , 2019
2019
Cited alongside, same era.
K. Kumar, R. Kumar, T. de Boissiere, L. Gestin, W. Z. Teoh, J. Sotelo, A. de Brébisson, Y. Bengio, and A. Courville, “MelGAN: Generative Adversarial Networks for Conditional Waveform Synthesis,” in Advances in Neural Information Processing Systems (NeurIPS) , 2019
2019
Cited alongside, same era.
R. Prenger, R. Valle, and B. Catanzaro, “WaveGlow: A Flow-based Generative Network for Speech Synthesis,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2019
2019
Cited alongside, same era.
2020
Later among the works it cites.
W. Ping, K. Peng, K. Zhao, and Z. Song, “WaveFlow: A Compact Flow-based Model for Raw Audio,” in International Conference on Machine Learning (ICML) , 2020
2020
Later among the works it cites.
Y. Ren, C. Hu, T. Qin, S. Zhao, Z. Zhao, and T.-Y. Liu, “FastSpeech 2: Fast and high-quality end-to-end text-to-speech,” in International Conference on Learning Representations (ICLR) , 2021
2021
Closest in time.
A. Lańcucki, “FastPitch: Parallel Text-to-Speech with Pitch Prediction,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2021
2021
Closest in time.
Z. Zeng, J. Wang, N. Cheng, and J. Xiao, “LVCNet: Efficient Condition-Dependent Modeling Network for Waveform Generation,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2021
2021
Closest in time.
A. Mustafa, N. Pia, and G. Fuchs, “StyleMelGAN: An Efficient High-Fidelity Adversarial Vocoder with Temporal Adaptive Normalization,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2021
2021
Closest in time.
G. Yang, S. Yang, K. Liu, P. Fang, W. Chen, and L. Xie, “Multi-band MelGAN: Faster Waveform Generation for High-Quality Text-to-Speech,” in IEEE Spoken Language Technology Workshop (SLT) , 2021
2021
Closest in time.