Fetching the paper…
Reading the bibliography…
Text-to-Speech (TTS) services that run on edge devices have many advantages compared to cloud TTS, e.g., latency and privacy issues.
ITU-T. Recommendation G. 711. , Pulse Code Modulation (PCM) of voice frequencies, 1988
1988
Earlier work this paper cites.
R. Kubichek, “Mel-cepstral distance measure for objective speech quality assessment,” in Proc. of IEEE Pacific Rim Conference on Communications Computers and Signal Processing , vol. 1, 1993, pp. 125–128
1993
Earlier work this paper cites.
ITU-T. Recommendation P. 800 , Methods for subjective determination of transmission quality, 1996
1996
Earlier work this paper cites.
J.-M. Valin, G. Maxwell, T. B. Terriberry, and K. Vos, “High-quality low-delay music coding in the opus codec,” Audio Eng. Soc. Conv. 135 , p. 8942, 2013
2013
Earlier work this paper cites.
2016
Earlier work this paper cites.
M. Morise, F. Yokomori, and K. Ozawa, “World: A vocoder-based high-quality speech synthesis system for real-time applications,” IEICE Transactions on Information and Systems , vol. E99.D, no. 7, pp. 1877–1884, 2016
2016
Earlier work this paper cites.
T. Salimans, A. Karpathy, X. Chen, and D. P. Kingma, “Improving the pixelcnn with discretized logistic mixture likelihood and other modifications,” in Proc. of International Conference on Learning Representations (ICLR) , 2017
2017
Earlier work this paper cites.
N. Kalchbrenner, E. Elsen, K. Simonyan, S. Noury, N. Casagrande, E. Lockhart, F. Stimberg, A. Oord, S. Dieleman, and K. Kavukcuoglu, “Efficient neural audio synthesis,” in Proc. of International Conference on Machine Learning (ICML) , 2018, pp. 2410–2419
2018
Cited alongside, same era.
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Skerrv-Ryan, R. A. Saurous, Y. Agiomvrgiannakis, and Y. Wu, “Natural tts synthesis by conditioning wavenet on mel spectrogram predictions,” in Proc. of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2018, pp. 4779–4783
2018
Cited alongside, same era.
A. van den Oord, Y. Li, I. Babuschkin, K. Simonyan, O. Vinyals, K. Kavukcuoglu, G. van den Driessche, E. Lockhart, L. Cobo, F. Stimberg, N. Casagrande, D. Grewe, S. Noury, S. Dieleman, E. Elsen, N. Kalchbrenner, H. Zen, A. Graves, H. King, T. Walters, D. Belov, and D. Hassabis, “Parallel WaveNet: Fast high-fidelity speech synthesis,” in Proc. of International Conference on Machine Learning (ICML) , 2018, pp. 3918–3926
2018
Cited alongside, same era.
J. Kong, J. Kim, and J. Bae, “Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,” in Proc. of Advances in Neural Information Processing Systems , 2020
2020
Later among the works it cites.
M.-J. Hwang, E. Song, R. Yamamoto, F. Soong, and H.-G. Kang, “Improving lpcnet-based text-to-speech with linear prediction-structured mixture density network,” in Proc. of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2020, pp. 7219–7223
2020
Later among the works it cites.
R. Vipperla, S. Park, K. Choo, S. Ishtiaq, K. Min, S. Bhattacharya, A. Mehrotra, A. G. C. P. Ramos, and N. D. Lane, “Bunched lpcnet : Vocoder for low-cost neural text-to-speech systems,” in Proc. of INTERSPEECH , 2020, pp. 3565–3569
2020
Later among the works it cites.
V. Popov, M. Kudinov, and T. Sadekova, “Gaussian lpcnet for multisample speech synthesis,” in Proc. of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2020, pp. 6204–6208
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
R. Prenger, R. Valle, and B. Catanzaro, “Waveglow: A flow-based generative network for speech synthesis,” in Proc. of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2019, pp. 3617–3621
2019
Cited alongside, same era.
K. Kumar, R. Kumar, T. de Boissiere, L. Gestin, W. Z. Teoh, J. Sotelo, A. de Brébisson, Y. Bengio, and A. C. Courville, “Melgan: Generative adversarial networks for conditional waveform synthesis,” in Proc. of Advances in Neural Information Processing Systems , 2019
2019
Cited alongside, same era.
J.-M. Valin and J. Skoglund, “LPCNet: Improving Neural Speech Synthesis through Linear Prediction,” in Proc. of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2019, pp. 5891–5895
2019
Cited alongside, same era.
2020
Cited alongside, same era.
2020
Later among the works it cites.
Y. Cui, X. Wang, L. He, and F. K. Soong, “An efficient subband linear prediction for lpcnet-based neural synthesis.” in Proc. of INTERSPEECH , 2020, pp. 3555–3559
2020
Later among the works it cites.
K. Matsubara, T. Okamoto, R. Takashima, T. Takiguchi, T. Toda, Y. Shiga, and H. Kawai, “Full-band lpcnet: A real-time neural vocoder for 48 khz audio with a cpu,” IEEE Access , vol. 9, pp. 94 923–94 933, 2021
2021
Later among the works it cites.
N. Ellinas, G. Vamvoukakis, K. Markopoulos, A. Chalamandaris, G. Maniati, P. Kakoulidis, S. Raptis, J. S. Sung, H. Park, and P. Tsiakoulis, “High quality streaming speech synthesis with low, sentence-length-independent latency,” in Proc. of INTERSPEECH , 2020, pp. 2022–2026
2026
Closest in time.