Fetching the paper…
Reading the bibliography…
Current text to speech (TTS) systems usually leverage a cascaded acoustic model and vocoder pipeline with mel-spectrograms as the intermediate representations, which suffer from two limitations: 1) the acoustic model and vocoder are separately trained instead of jointly optimized, which incurs cascaded errors; 2) the intermediate speech representations (e.g., mel-spectrogram) are pre-designed and lose phase information, which are sub-optimal.
B.-H. Juang and A. Gray, “Multiple stage vector quantization for speech coding,” in ICASSP’82. IEEE International Conference on Acoustics, Speech, and Signal Processing , vol. 7. IEEE, 1982, pp. 597–600
1982
Earlier work this paper cites.
G. P. Nason and B. W. Silverman, “The discrete wavelet transform in s,” Journal of Computational and Graphical statistics , vol. 3, no. 2, pp. 163–191, 1994
1994
Earlier work this paper cites.
Z. Wang, A. C. Bovik, H. R. Sheikh, and E. P. Simoncelli, “Image quality assessment: from error visibility to structural similarity,” IEEE transactions on image processing , vol. 13, no. 4, pp. 600–612, 2004
2004
Earlier work this paper cites.
2014
Earlier work this paper cites.
A. Hines, J. Skoglund, A. C. Kokaram, and N. Harte, “Visqol: an objective speech quality model,” EURASIP Journal on Audio, Speech, and Music Processing , vol. 2015, no. 1, pp. 1–18, 2015
2015
Earlier work this paper cites.
2018
Earlier work this paper cites.
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Skerrv-Ryan et al. , “Natural TTS synthesis by conditioning wavenet on mel spectrogram predictions,” in 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2018, pp. 4779–4783
2018
Earlier work this paper cites.
N. Li, S. Liu, Y. Liu, S. Zhao, and M. Liu, “Neural speech synthesis with transformer network,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 33, no. 01, 2019, pp. 6706–6713
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2020
Cited alongside, same era.
2020
Cited alongside, same era.
J. Kong, J. Kim, and J. Bae, “Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,” Advances in Neural Information Processing Systems , vol. 33, 2020
2020
Cited alongside, same era.
R. Liu, B. Sisman, J. Li, F. Bao, G. Gao, and H. Li, “Teacher-student training for robust tacotron-based TTS,” in ICASSP 2020-2020 IEEE international conference on acoustics, speech and signal processing (ICASSP) . IEEE, 2020, pp. 6274–6278
2020
Cited alongside, same era.
2021
Later among the works it cites.
2021
Later among the works it cites.
R. J. Weiss, R. Skerry-Ryan, E. Battenberg, S. Mariooryad, and D. P. Kingma, “Wave-tacotron: Spectrogram-free end-to-end text-to-speech synthesis,” in ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2021, pp. 5679–5683
2021
Later among the works it cites.
A. Mustafa, J. Büthe, S. Korse, K. Gupta, G. Fuchs, and N. Pia, “A streamwise gan vocoder for wideband speech coding at very low bit rate,” in 2021 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA) . IEEE, 2021, pp. 66–70
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
R. Yamamoto, E. Song, and J.-M. Kim, “Parallel wavegan: A fast waveform generation model based on generative adversarial networks with multi-resolution spectrogram,” in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 6199–6203
2020
Cited alongside, same era.
Microsoft, “Azure neural TTS upgraded with hifinet, achieving higher audio fidelity and faster synthesis speed,” p. 0, Nov 2020, hiFiNet. [Online]. Available: https://techcommunity.microsoft.com/t5/azure-ai/azure-neural-tts-upgraded-with-hifinet-achieving-higher-audio/ba-p/1847860
2020
Cited alongside, same era.
2021
Cited alongside, same era.
Y. Ren, C. Hu, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T.-Y. Liu, “Fastspeech 2: Fast and high-quality end-to-end text to speech,” in International Conference on Learning Representations , 2021. [Online]. Available: https://openreview.net/forum?id=piLPYqxtWuA
2021
Cited alongside, same era.
2021
Cited alongside, same era.
2021
Later among the works it cites.
N. Zeghidour, A. Luebs, A. Omran, J. Skoglund, and M. Tagliasacchi, “Soundstream: An end-to-end neural audio codec,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
T. Srikotr and K. Mano, “Vector quantization of speech spectrum based on the vq-vae embedding space learning by gan technique,” IEICE Transactions on Fundamentals of Electronics, Communications and Computer Sciences , 2021
2021
Later among the works it cites.