Fetching the paper…
Reading the bibliography…
This paper introduces a high-quality open-source speech synthesis dataset for Kazakh, a low-resource language spoken by over 13 million people worldwide.
P. Boersma, “Praat, a system for doing phonetics by computer,” Glot International , vol. 5, no. 9, pp. 341–345, 2001
2001
Earlier work this paper cites.
A. W. Black and K. Tokuda, “The Blizzard Challenge - 2005: Evaluating corpus-based speech synthesis on common datasets,” in Proc. Eurospeech European Conference on Speech Communication and Technology (Interspeech) . ISCA, 2005, pp. 77–80
2005
Earlier work this paper cites.
B. Dave, Kazakhstan-ethnicity, language and power . Routledge, 2007
2007
Earlier work this paper cites.
P. Taylor, Text-to-speech synthesis . Cambridge University Press, 2009
2009
Earlier work this paper cites.
2009
Earlier work this paper cites.
2010
Earlier work this paper cites.
O. Makhambetov, A. Makazhanov, Z. Yessenbayev, B. Matkarimov, I. Sabyrgaliyev, and A. Sharafudinov, “Assembling the Kazakh language corpus,” in Proc. Conference on Empirical Methods in Natural Language Processing (EMNLP) . ACL, 2013, pp. 1022–1031
2013
Earlier work this paper cites.
N. Perraudin, P. Balázs, and P. L. Søndergaard, “A fast Griffin-Lim algorithm,” in Proc. IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA) . IEEE, 2013, pp. 1–4
2013
Earlier work this paper cites.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in Proc. International Conference on Learning Representations (ICLR) , 2015
2015
Earlier work this paper cites.
C. Veaux, J. Yamagishi, K. MacDonald et al. , “Superseded-CSTR VCTK corpus: English multi-speaker corpus for CSTR voice cloning toolkit,” 2016
2016
Cited alongside, same era.
A. van den Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. W. Senior, and K. Kavukcuoglu, “WaveNet: A generative model for raw audio,” in Proc. ISCA Speech Synthesis Workshop . ISCA, 2016, p. 125
2016
Cited alongside, same era.
Y. Wang, R. J. Skerry-Ryan, D. Stanton, Y. Wu, R. J. Weiss, N. Jaitly, Z. Yang, Y. Xiao, Z. Chen, S. Bengio, Q. V. Le, Y. Agiomyrgiannakis, R. Clark, and R. A. Saurous, “Tacotron: Towards end-to-end speech synthesis,” in Proc. Annual Conference of the International Speech Communication Association (Interspeech) . ISCA, 2017, pp. 4006–4010
2017
Cited alongside, same era.
J. Sotelo, S. Mehri, K. Kumar, J. F. Santos, K. Kastner, A. C. Courville, and Y. Bengio, “Char2Wav: End-to-end speech synthesis,” in Proc. International Conference on Learning Representations (ICLR) , 2017
N. Li, S. Liu, Y. Liu, S. Zhao, and M. Liu, “Neural speech synthesis with transformer network,” in Proc. AAAI Conference on Artificial Intelligence (AAAI) . AAAI Press, 2019, pp. 6706–6713
2019
Later among the works it cites.
H. Zen, V. Dang, R. Clark, Y. Zhang, R. J. Weiss, Y. Jia, Z. Chen, and Y. Wu, “LibriTTS: A corpus derived from LibriSpeech for text-to-speech,” in Proc. Annual Conference of the International Speech Communication Association (Interspeech) . ISCA, 2019, pp. 1526–1530
2019
Later among the works it cites.
Y. Ren, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T. Liu, “Almost unsupervised text to speech and automatic speech recognition,” in Proc. International Conference on Machine Learning (ICML) , vol. 97. PMLR, 2019, pp. 5410–5419
2019
Later among the works it cites.
Y. Chung, Y. Wang, W. Hsu, Y. Zhang, and R. J. Skerry-Ryan, “Semi-supervised training for improving data efficiency in end-to-end speech synthesis,” in Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2019, pp. 6940–6944
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2017
Cited alongside, same era.
S. Ö. Arik, M. Chrzanowski, A. Coates, G. F. Diamos, A. Gibiansky, Y. Kang, X. Li, J. Miller, A. Y. Ng, J. Raiman, S. Sengupta, and M. Shoeybi, “Deep Voice: Real-time neural text-to-speech,” in Proc. International Conference on Machine Learning (ICML) , vol. 70. PMLR, 2017, pp. 195–204
2017
Cited alongside, same era.
K. Ito and L. Johnson, “The LJ speech dataset,” https://keithito.com/LJ-Speech-Dataset/ , 2017
2017
Cited alongside, same era.
R. J. Skerry-Ryan, E. Battenberg, Y. Xiao, Y. Wang, D. Stanton, J. Shor, R. J. Weiss, R. Clark, and R. A. Saurous, “Towards end-to-end prosody transfer for expressive speech synthesis with tacotron,” in Proc. International Conference on Machine Learning (ICML) , vol. 80. PMLR, 2018, pp. 4700–4709
2018
Cited alongside, same era.
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Ryan, R. A. Saurous, Y. Agiomyrgiannakis, and Y. Wu, “Natural TTS synthesis by conditioning Wavenet on MEL spectrogram predictions,” in Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2018, pp. 4779–4783
2018
Cited alongside, same era.
Y. Ren, Y. Ruan, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T. Liu, “FastSpeech: Fast, robust and controllable text to speech,” in Proc. Annual Conference on Neural Information Processing Systems (NeurIPS) , 2019, pp. 3165–3174
2019
Cited alongside, same era.
Telegram FZ LLC and Telegram Messenger Inc., “Telegram.” [Online]. Available: https://telegram.org
Cited in the paper.
Amazon.com, Inc., “Amazon Mechanical Turk (MTurk).” [Online]. Available: https://www.mturk.com
Cited in the paper.
2019
Later among the works it cites.
O. Mamyrbayev, K. Alimhan, B. Zhumazhanov, T. Turdalykyzy, and F. Gusmanova, “End-to-end speech recognition in agglutinative languages,” in Proc. Asian Conference on Intelligent Information and Database Systems (ACIIDS) , vol. 12034, 2020, pp. 391–401
2020
Later among the works it cites.
T. Hayashi, R. Yamamoto, K. Inoue, T. Yoshimura, S. Watanabe, T. Toda, K. Takeda, Y. Zhang, and X. Tan, “ESPnet-TTS: Unified, reproducible, and integratable open source end-to-end text-to-speech toolkit,” in Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 7654–7658
2020
Later among the works it cites.
R. Yamamoto, E. Song, and J. Kim, “Parallel Wavegan: A fast waveform generation model based on generative adversarial networks with multi-resolution spectrogram,” in Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 6199–6203
2020
Later among the works it cites.
Y. Chen, T. Tu, C. Yeh, and H. Lee, “End-to-end text-to-speech for low-resource languages by cross-lingual transfer learning,” in Proc. Annual Conference of the International Speech Communication Association (Interspeech) . ISCA, 2019, pp. 2075–2079
2079
Closest in time.