Fetching the paper…
Reading the bibliography…
Although there are more than 6,500 languages in the world, the pronunciations of many phonemes sound similar across the languages.
E. H Rothauser, W. D Chapman, N. Guttman, H. R Silbiger, M. Hecker, G. E Urbanek, K. S Nordby, and M. Weinstock, “Ieee recommended pratice for speech quality measurements,”
1969
Earlier work this paper cites.
L. Rice, “Hardware and software for speech synthesis,”
1976
Earlier work this paper cites.
D. W. Griffin and J. S. Lim, “Signal estimation from modified short-time fourier transform,” in
1983
Earlier work this paper cites.
I. P. Association, C. PRESS, and D. Decker,
1999
Earlier work this paper cites.
J. E. Flege, C. Schirru, and I. R. MacKay, “Interaction between the native and second language phonetic subsystems,”
2003
Earlier work this paper cites.
J. Kominek, A. W. Black, and V. Ver, “Cmu arctic databases for speech synthesis,” Tech. Rep., 2003
2003
Earlier work this paper cites.
H. Zen, N. Braunschweiler, S. Buchholz, M. J. F. Gales, K. Knill, S. Krstulovic, and J. Latorre, “Statistical parametric speech synthesis based on speaker and language factorization,”
2012
Earlier work this paper cites.
A. I. Rudnicky, “The cmu pronouncing dictionary,” http://www.speech.cs.cmu.edu/cgi-bin/cmudict, 2015
2015
Earlier work this paper cites.
B. Li and H. Zen, “Multi-language multi-speaker acoustic modeling for lstm-rnn based statistical parametric speech synthesis,” 2016
2016
Cited alongside, same era.
C. Veaux, J. Yamagishi, K. MacDonald
2016
Cited alongside, same era.
Y. Wang, R. Skerry-Ryan, D. Stanton, Y. Wu, R. J. Weiss, N. Jaitly, Z. Yang, Y. Xiao, Z. Chen, S. Bengio, Q. Le, Y. Agiomyrgiannakis, R. Clark, and R. A. Saurous, “Tacotron: Towards end-to-end speech synthesis,” in
2017
Cited alongside, same era.
S. Ö. Arik, M. Chrzanowski, A. Coates, G. Diamos, A. Gibiansky, Y. Kang, X. Li, J. Miller, A. Ng, J. Raiman
2017
Cited alongside, same era.
A. Gibiansky, S. Arik, G. Diamos, J. Miller, K. Peng, W. Ping, J. Raiman, and Y. Zhou, “Deep voice 2: Multi-speaker neural text-to-speech,” in
2017
Cited alongside, same era.
Y. Cho, “Kog2p,” https://github.com/scarletcho/KoG2P, 2017
2017
Later among the works it cites.
N. Li, S. Liu, Y. Liu, S. Zhao, M. Liu, and M. Zhou, “Neural speech synthesis with transformer network,”
2018
Closest in time.
Y.-A. Chung, Y. Wang, W.-N. Hsu, Y. Zhang, and R. Skerry-Ryan, “Semi-supervised training for improving data efficiency in end-to-end speech synthesis,”
2018
Closest in time.
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Ryan, R. A. Saurous, Y. Agiomyrgiannakis, and Y. Wu, “Natural TTS synthesis by conditioning wavenet on MEL spectrogram predictions,” in
2018
Closest in time.
Y. Jia, Y. Zhang, R. J. Weiss, Q. Wang, J. Shen, F. Ren, Z. Chen, P. Nguyen, R. Pang, I. L. Moreno
2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Gutkin, “Uniform multilingual multi-speaker acoustic model for statistical parametric speech synthesis of low-resourced languages,” in
2017
Cited alongside, same era.
K. Ito, “The lj speech dataset,” https://keithito.com/LJ-Speech-Dataset/, 2017
2017
Cited alongside, same era.
K. Park and T. Mulc, “Css10: A collection of single speaker speech datasets for 10 languages,” https://github.com/Kyubyong/css10/, 2018
2018
Closest in time.
K. Park and J. Kim, “g2p-en,” https://github.com/Kyubyong/g2p, 2018
2018
Closest in time.