Fetching the paper…
Reading the bibliography…
Recent advancements in speech generation have been driven by large-scale training datasets.
K. P. F.R.S., “Liii. on lines and planes of closest fit to systems of points in space,” The London, Edinburgh, and Dublin Philosophical Magazine and Journal of Science , vol. 2, no. 11, pp. 559–572, 1901
1901
Earlier work this paper cites.
K. Ito and L. Johnson, “The lj speech dataset,” keithito.com/LJ-Speech-Dataset/ , 2017
2017
Earlier work this paper cites.
Y. Wang, R. Skerry-Ryan, D. Stanton, Y. Wu, R. J. Weiss, N. Jaitly, Z. Yang, Y. Xiao, Z. Chen, S. Bengio et al. , “Tacotron: Towards end-to-end speech synthesis,” in INTERSPEECH , 2017
2017
Earlier work this paper cites.
J. Yamagishi, C. Veaux, and K. MacDonald, “Cstr vctk corpus: English multi-speaker corpus for cstr voice cloning toolkit (version 0.92),” datashare.ed.ac.uk/handle/10283/3443 , 2019
2019
Earlier work this paper cites.
H. Zen, V. Dang, R. Clark, Y. Zhang, R. J. Weiss, Y. Jia, Z. Chen, and Y. Wu, “Libritts: A corpus derived from librispeech for text-to-speech,” in INTERSPEECH , 2019
2019
Earlier work this paper cites.
Y. Ren, Y. Ruan, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T.-Y. Liu, “Fastspeech: Fast, robust and controllable text to speech,” in NeurIPS , 2019
2019
Earlier work this paper cites.
N. Reimers and I. Gurevych, “Sentence-bert: Sentence embeddings using siamese bert-networks,” in EMNLP , 2019
2019
Earlier work this paper cites.
J. Kahn, M. Rivière, W. Zheng, E. Kharitonov, Q. Xu, P.-E. Mazaré, J. Karadayi, V. Liptchinsky, R. Collobert, C. Fuegen, T. Likhomanenko, G. Synnaeve, A. Joulin, A. Mohamed, and E. Dupoux, “Libri-light: A benchmark for asr with limited or no supervision,” in ICASSP , 2020
2020
Earlier work this paper cites.
V. Pratap, Q. Xu, A. Sriram, G. Synnaeve, and R. Collobert, “Mls: A large-scale multilingual dataset for speech research,” in INTERSPEECH , 2020
2020
Earlier work this paper cites.
R. Ardila, M. Branson, K. Davis, M. Kohler, J. Meyer, M. Henretty, R. Morais, L. Saunders, F. Tyers, and G. Weber, “Common voice: A massively-multilingual speech corpus,” in LREC , 2020
2020
Earlier work this paper cites.
2021
Earlier work this paper cites.
T.-h. Huang, J.-h. Lin, and H.-y. Lee, “How far are we from robust voice conversion: A survey,” in SLT , 2021
2021
Earlier work this paper cites.
S. Yao, H. Bu, X. Xu, S. Zhang, and M. Li, “Aishell-3: A multi-speaker mandarin tts corpus,” in INTERSPEECH , 2021
2021
Earlier work this paper cites.
G. Chen, S. Chai, G. Wang, J. Du, W.-Q. Zhang, C. Weng, D. Su, D. Povey, J. Trmal, J. Zhang, M. Jin, S. Khudanpur, S. Watanabe, S. Zhao, W. Zou, X. Li, X. Yao, Y. Wang, Y. Wang, Z. You, and Z. Yan, “Gigaspeech: An evolving, multi-domain asr corpus with 10,000 hours of transcribed audio,” in INTERSPEECH , 2021
2021
Earlier work this paper cites.
M. Kim, W. Choi, J. Chung, D. Lee, and S. Jung, “Kuielab-mdx-net: A two-stream neural network for music demixing,” in ISMIR MDX Workshop , 2021
2021
Earlier work this paper cites.
B. Zhang, H. Lv, P. Guo, Q. Shao, C. Yang, L. Xie, X. Xu, H. Bu, X. Chen, C. Zeng, D. Wu, and Z. Peng, “Wenetspeech: A 10000+ hours multi-domain mandarin corpus for speech recognition,” in ICASSP , 2022
2022
Cited alongside, same era.
C. K. Reddy, V. Gopal, and R. Cutler, “Dnsmos p.835: A non-intrusive perceptual objective speech quality metric to evaluate noise suppressors,” in ICASSP , 2022
2022
Cited alongside, same era.
S. Chen, C. Wang, Z. Chen, Y. Wu, S. Liu, Z. Chen, J. Li, N. Kanda, T. Yoshioka, X. Xiao et al. , “Wavlm: Large-scale self-supervised pre-training for full stack speech processing,” IEEE Journal of Selected Topics in Signal Processing , vol. 16, no. 6, pp. 1505–1518, 2022
2022
Cited alongside, same era.
2023
Cited alongside, same era.
2024
Later among the works it cites.
X. Zhang, L. Xue, Y. Gu, Y. Wang, J. Li, H. He, C. Wang, T. Song, X. Chen, Z. Fang, H. Chen, J. Zhang, T. Y. Tang, L. Zou, M. Wang, J. Han, K. Chen, H. Li, and Z. Wu, “Amphion: An open-source audio, music and speech generation toolkit,” in SLT , 2024
2024
Later among the works it cites.
J. Yu, H. Chen, Y. Bian, X. Li, Y. Luo, J. Tian, M. Liu, J. Jiang, and S. Wang, “Autoprep: An automatic preprocessing framework for in-the-wild speech data,” in ICASSP , 2024
2024
Later among the works it cites.
L. Ma, D. Guo, K. Song, Y. Jiang, S. Wang, L. Xue, W. Xu, H. Zhao, B. Zhang, and L. Xie, “Wenetspeech4tts: A 12,800-hour mandarin tts corpus for large speech generation model benchmark,” in INTERSPEECH , 2024
2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
H. Bredin, “Pyannote.audio 2.1 speaker diarization pipeline: Principle, benchmark, and recipe,” in INTERSPEECH , 2023
2023
Cited alongside, same era.
A. Plaquet and H. Bredin, “Powerset multi-class cross entropy loss for neural speaker diarization,” in INTERSPEECH , 2023
2023
Cited alongside, same era.
A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, “Robust speech recognition via large-scale weak supervision,” in ICML , 2023
2023
Cited alongside, same era.
M. Bain, J. Huh, T. Han, and A. Zisserman, “Whisperx: Time-accurate speech transcription of long-form audio,” in INTERSPEECH , 2023
2023
Cited alongside, same era.
X. Li, S. Takamichi, T. Saeki, W. Chen, S. Shiota, and S. Watanabe, “Yodas: Youtube-oriented dataset for audio and speech,” in ASRU , 2023
2023
Cited alongside, same era.
W. Kang, X. Yang, Z. Yao, F. Kuang, Y. Yang, L. Guo, L. Lin, and D. Povey, “Libriheavy: a 50,000 hours asr corpus with punctuation casing and context,” in ICASSP , 2024
2024
Later among the works it cites.
H. He, Z. Shang, C. Wang, X. Li, Y. Gu, H. Hua, L. Liu, C. Yang, J. Li, P. Shi, Y. Wang, K. Chen, P. Zhang, and Z. Wu, “Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation,” in SLT , 2024
2024
Later among the works it cites.
X. Tan, J. Chen, H. Liu, J. Cong, C. Zhang, Y. Liu, X. Wang, Y. Leng, Y. Yi, L. He, S. Zhao, T. Qin, F. Soong, and T.-Y. Liu, “Naturalspeech: End-to-end text-to-speech synthesis with human-level quality,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 46, no. 6, pp. 4234–4245, 2024
2024
Later among the works it cites.
Z. Ma, Z. Zheng, J. Ye, J. Li, Z. Gao, S. Zhang, and X. Chen, “Emotion2vec: Self-supervised pre-training for speech emotion representation,” Findings of ACL , 2024
2024
Later among the works it cites.
P. Liu, L. Wang, R. He, H. He, L. Wang, H. Zheng, J. Shi, T. Xiao, and Z. Wu, “Spmis: An investigation of synthetic spoken misinformation detection,” in SLT , 2024
2024
Later among the works it cites.
X. Li, Z. Shang, H. Hua, P. Shi, C. Yang, L. Wang, and P. Zhang, “Sf-speech: Straightened flow for zero-shot voice clone,” IEEE Transactions on Audio, Speech and Language Processing , vol. 33, pp. 1706–1718, 2025
2025
Closest in time.
International Telecommunication Union, “P.835: Subjective test methodology for evaluating speech communication systems that include noise suppression algorithm,” https://www.itu.int/rec/T-REC-P.835 , 2003, accessed: 2025-09-05
2025
Closest in time.
Y. Wang, H. Zhan, L. Liu, R. Zeng, H. Guo, J. Zheng, Q. Zhang, X. Zhang, S. Zhang, and Z. Wu, “Maskgct: Zero-shot text-to-speech with masked generative codec transformer,” in ICLR , 2025
2025
Closest in time.
2025
Closest in time.