Fetching the paper…
Reading the bibliography…
With the recent developments in cross-lingual Text-to-Speech (TTS) systems, L2 (second-language, or foreign) accent problems arise.
B. M. Lobanov, “Classification of russian vowels spoken by different speakers,” The Journal of the Acoustical Society of America , vol. 49, no. 2B, pp. 606–608, 1971
1971
Earlier work this paper cites.
G. Fant, “Speech sounds and features.” 1973
1973
Earlier work this paper cites.
J. Hillenbrand, L. A. Getty, M. J. Clark, and K. Wheeler, “Acoustic characteristics of american english vowels,” The Journal of the Acoustical society of America , vol. 97, no. 5, pp. 3099–3111, 1995
1995
Earlier work this paper cites.
O.-S. B.-J. E. Flege, “Perpeption and production of a new vowel category,” Second-language speech: Structure and process , vol. 13, p. 53, 1997
1997
Earlier work this paper cites.
S. H. Weinberger, “Minimal segments in second language phonology,” Second language speech: Structure and process , pp. 263–312, 1997
1997
Earlier work this paper cites.
A. Dowd, J. Smith, and J. Wolfe, “Learning to pronounce vowel sounds in a foreign language using acoustic measurements of the vocal tract as feedback in real time,” Language and Speech , vol. 41, no. 1, pp. 1–20, 1998
1998
Earlier work this paper cites.
I. P. Association, Handbook of the International Phonetic Association: A guide to the use of the International Phonetic Alphabet . Cambridge University Press, 1999
1999
Earlier work this paper cites.
C.-G. Kwak, “The vowel system of contemporary korean and direction of change,” Journal of Korean Linguistics , vol. 41, pp. 59–91, 2003
2003
Earlier work this paper cites.
J. E. Flege, “Assessing constraints on second-language segmental production and perception,” Phonetics and phonology in language comprehension and production: Differences and similarities , vol. 6, pp. 319–355, 2003
2003
Earlier work this paper cites.
P. Adank, R. Smits, and R. Van Hout, “A comparison of vowel normalization procedures for language variation research,” The Journal of the Acoustical Society of America , vol. 116, no. 5, pp. 3099–3107, 2004
2004
Earlier work this paper cites.
W. Baker and P. Trofimovich, “Interaction of native-and second-language vowel system (s) in early and late bilinguals,” Language and speech , vol. 48, no. 1, pp. 1–27, 2005
2005
Earlier work this paper cites.
T. J. Vance, The sounds of Japanese with audio CD . Cambridge University Press, 2008
2008
Earlier work this paper cites.
A. H. Fabricius, D. Watt, and D. E. Johnson, “A comparison of three speaker-intrinsic vowel formant frequency normalization algorithms for sociophonetics,” Language Variation and Change , vol. 21, no. 3, pp. 413–435, 2009
2009
Cited alongside, same era.
S. Sandoval, V. Berisha, R. L. Utianski, J. M. Liss, and A. Spanias, “Automatic assessment of vowel space area,” The Journal of the Acoustical Society of America , vol. 134, no. 5, pp. EL477–EL483, 2013
2013
Cited alongside, same era.
J. Nycz and L. Hall-Lew, “Best practices in measuring vowel merger,” in Proc. of Meetings on Acoustics 166ASA , vol. 20, no. 1. Acoustical Society of America, 2013, p. 060008
2013
Cited alongside, same era.
S. Dimov and A. Bradlow, “Non-native vowel production accuracy and variability in relation to overall intelligibility,” The Journal of the Acoustical Society of America , vol. 134, no. 5, pp. 4107–4107, 2013
2013
Cited alongside, same era.
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Skerrv-Ryan, R. A. Saurous, Y. Agiomvrgiannakis, and Y. Wu, “Natural tts synthesis by conditioning wavenet on mel spectrogram predictions,” in Proc. ICASSP , 2018, pp. 4779–4783
2018
Later among the works it cites.
E. Nachmani and L. Wolf, “Unsupervised polyglot text-to-speech,” in Proc. ICASSP . IEEE, 2019, pp. 7055–7059
2019
Later among the works it cites.
K. Park and T. Mulc, “Css10: A collection of single speaker speech datasets for 10 languages,” Interspeech , 2019
2019
Later among the works it cites.
J.-M. Valin and J. Skoglund, “Lpcnet: Improving neural speech synthesis through linear prediction,” in ICASSP . IEEE, 2019, pp. 5891–5895
2019
Later among the works it cites.
M. K. Huffman and K. S. Schuhmann, “The relation between l1 and l2 category compactness and l2 vot learning,” in Proc. of Meetings on Acoustics 179ASA , vol. 42, no. 1. Acoustical Society of America, 2020, p. 060011
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
B. Collins and I. M. Mees, Practical phonetics and phonology: A resource book for students . Routledge, 2013
2013
Cited alongside, same era.
N. Kartushina and U. H. Frauenfelder, “On the effects of l2 perception and of individual differences in l1 production on l2 pronunciation,” Frontiers in psychology , vol. 5, p. 1246, 2014
2014
Cited alongside, same era.
P. Ladefoged and K. Johnson, A course in phonetics . Cengage learning, 2014
2014
Cited alongside, same era.
S. Kleiner, “Duden–das aussprachewörterbuch. bearbeitet von stefan kleiner und ralf knöbl in zusammenarbeit mit der dudenredaktion. 7., komplett überarb. und aktual,” Aufl. Berlin , 2015
2015
Cited alongside, same era.
N. Kartushina, A. Hervais-Adelman, U. H. Frauenfelder, and N. Golestani, “Mutual influences between native and non-native vowels in production: Evidence from short-term visual articulatory feedback training,” Journal of Phonetics , vol. 57, pp. 21–39, 2016
2016
Cited alongside, same era.
Y. Wang, R. Skerry-Ryan, D. Stanton, Y. Wu, R. J. Weiss, N. Jaitly, Z. Yang, Y. Xiao, Z. Chen, S. Bengio, Q. Le, Y. Agiomyrgiannakis, R. Clark, and R. A. Saurous, “Tacotron: Towards End-to-End Speech Synthesis,” in Proc. Interspeech , 2017, pp. 4006–4010
2017
Cited alongside, same era.
K. Ito and L. Johnson, “The lj speech dataset,” https://keithito.com/LJ-Speech-Dataset/
2017
Cited alongside, same era.
2020
Later among the works it cites.
J. Kim, S. Kim, J. Kong, and S. Yoon, “Glow-tts: A generative flow for text-to-speech via monotonic alignment search,” NeurIPS , vol. 33, pp. 8067–8077, 2020
2020
Later among the works it cites.
J. Ye, H. Zhou, Z. Su, W. He, K. Ren, L. Li, and H. Lu, “Improving cross-lingual speech synthesis with triplet training scheme,” in ICASSP . IEEE, 2022, pp. 6072–6076
2022
Closest in time.
B. N. Abeysinghe, J. James, C. Watson, and F. Marattukalam, “Visualising Model Training via Vowel Space for Text-To-Speech Systems,” in Proc. Interspeech , 2022, pp. 511–515
2022
Closest in time.
S. Park, K. Choo, J. Lee, A. V. Porov, K. Osipov, and J. S. Sung, “Bunched LPCNet2: Efficient Neural Vocoders Covering Devices from Cloud to Edge,” in Proc. Interspeech , 2022, pp. 808–812
2022
Closest in time.
N. Ellinas, G. Vamvoukakis, K. Markopoulos, A. Chalamandaris, G. Maniati, P. Kakoulidis, S. Raptis, J. S. Sung, H. Park, and P. Tsiakoulis, “High Quality Streaming Speech Synthesis with Low, Sentence-Length-Independent Latency,” in Proc. Interspeech , 2020, pp. 2022–2026
2026
Closest in time.
Y. Zhang, R. J. Weiss, H. Zen, Y. Wu, Z. Chen, R. Skerry-Ryan, Y. Jia, A. Rosenberg, and B. Ramabhadran, “Learning to Speak Fluently in a Foreign Language: Multilingual Speech Synthesis and Cross-Language Voice Cloning,” in Proc. Interspeech , 2019, pp. 2080–2084
2084
Closest in time.