Fetching the paper…
Reading the bibliography…
The multi-codebook speech codec enables the application of large language models (LLM) in TTS but bottlenecks efficiency and robustness due to multi-sequence prediction.
R. F. Kubichek, “Mel-cepstral distance measure for objective speech quality assessment,” Proceedings of IEEE Pacific Rim Conference on Communications Computers and Signal Processing , vol. 1, pp. 125–128 vol.1, 1993
1993
Earlier work this paper cites.
M. Schuster and K. K. Paliwal, “Bidirectional recurrent neural networks,” IEEE Trans. Signal Process. , vol. 45, pp. 2673–2681, 1997
1997
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural Comput. , vol. 9, no. 8, pp. 1735–1780, 1997
1997
Earlier work this paper cites.
I.-T. Recommendation, “Perceptual evaluation of speech quality (pesq): An objective method for end-to-end speech quality assessment of narrow-band telephone networks and speech codecs,” Rec. ITU-T P. 862 , 2001
2001
Earlier work this paper cites.
E. Vincent, R. Gribonval, and C. Févotte, “Performance measurement in blind audio source separation,” IEEE Trans. Speech Audio Process. , vol. 14, no. 4, pp. 1462–1469, 2006
2006
Earlier work this paper cites.
C. H. Taal, R. C. Hendriks, R. Heusdens, and J. Jensen, “A short-time objective intelligibility measure for time-frequency weighted noisy speech,” in ICASSP . IEEE, 2010, pp. 4214–4217
2010
Earlier work this paper cites.
2014
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in CVPR . IEEE Computer Society, 2016, pp. 770–778
2016
Earlier work this paper cites.
2017
Earlier work this paper cites.
H. Bu, J. Du, X. Na, B. Wu, and H. Zheng, “AISHELL-1: an open-source mandarin speech corpus and a speech recognition baseline,” in O-COCOSDA . IEEE, 2017, pp. 1–5
2017
Earlier work this paper cites.
H. Zen, V. Dang, R. Clark, Y. Zhang, R. J. Weiss, Y. Jia, Z. Chen, and Y. Wu, “Libritts: A corpus derived from librispeech for text-to-speech,” in INTERSPEECH . ISCA, 2019, pp. 1526–1530
2019
Earlier work this paper cites.
Z. Liu and B. K.-W. Mak, “Cross-lingual multi-speaker text-to-speech synthesis for voice cloning without using parallel corpus for unseen speakers,” arXiv: Audio and Speech Processing , 2019
2019
Cited alongside, same era.
2020
Cited alongside, same era.
A. Gulati, J. Qin, C. Chiu, N. Parmar, Y. Zhang, J. Yu, W. Han, S. Wang, Z. Zhang, Y. Wu, and R. Pang, “Conformer: Convolution-augmented transformer for speech recognition,” in INTERSPEECH . ISCA, 2020, pp. 5036–5040
2020
Cited alongside, same era.
P. Esser, R. Rombach, and B. Ommer, “Taming transformers for high-resolution image synthesis,” in CVPR . Computer Vision Foundation / IEEE, 2021, pp. 12 873–12 883
2021
Cited alongside, same era.
2023
Later among the works it cites.
2023
Later among the works it cites.
J. Betker, “Better speech synthesis through scaling,” CoRR , vol. abs/2305.07243, 2023
2023
Later among the works it cites.
S. Lee, W. Ping, B. Ginsburg, B. Catanzaro, and S. Yoon, “Bigvgan: A universal neural vocoder with large-scale training,” in ICLR . OpenReview.net, 2023
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
E. Bakhturina, V. Lavrukhin, B. Ginsburg, and Y. Zhang, “Hi-fi multi-speaker english TTS dataset,” in Interspeech . ISCA, 2021, pp. 2776–2780
2021
Cited alongside, same era.
2022
Cited alongside, same era.
T. Saeki, D. Xin, W. Nakata, T. Koriyama, S. Takamichi, and H. Saruwatari, “UTMOS: utokyo-sarulab system for voicemos challenge 2022,” in INTERSPEECH . ISCA, 2022, pp. 4521–4525
2022
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
H. Wang, C. Liang, S. Wang, Z. Chen, B. Zhang, X. Xiang, Y. Deng, and Y. Qian, “Wespeaker: A research and production oriented speaker embedding learning toolkit,” in ICASSP . IEEE, 2023, pp. 1–5
2023
Later among the works it cites.
A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, “Robust speech recognition via large-scale weak supervision,” in ICML , ser. Proceedings of Machine Learning Research, vol. 202. PMLR, 2023, pp. 28 492–28 518
2023
Later among the works it cites.
2024
Closest in time.