Fetching the paper…
Reading the bibliography…
Although end-to-end text-to-speech (TTS) models can generate natural speech, challenges still remain when it comes to estimating sentence-level phonetic and prosodic information from raw text in Japanese TTS systems.
“Accentuation rules for Japanese word concatenation,”
Y. Sagisaka and H. Sato, · 1983
Earlier work this paper cites.
“Long short-term memory,”
S. Hochreiter and J. Schmidhuber, · 1997
Earlier work this paper cites.
“Conditional random fields: Probabilistic models for segmenting and labeling sequence data,”
J. D. Lafferty, A. McCallum, and F. C. N. Pereira, · 2001
Earlier work this paper cites.
“Corpus of spontaneous Japanese: its design and evaluation,”
Kikuo Maekawa, · 2003
Earlier work this paper cites.
“Word-based partial annotation for efficient corpus construction,”
G. Neubig and S. Mori, · 2010
Earlier work this paper cites.
“Japanese pronunciation prediction as phrasal statistical machine translation,”
J. Hatori and H. Suzuki, · 2011
Earlier work this paper cites.
“Accent sandhi estimation of Tokyo dialect of Japanese using conditional random fields,”
M. Suzuki, R. Kuroiwa, K. Innami, S. Kobayashi, S. Shimizu, N. Minematsu, and K. Hirose, · 2017
Earlier work this paper cites.
“Attention is all you need,”
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, · 2017
Earlier work this paper cites.
“JSUT corpus: free large-scale Japanese speech corpus for end-to-end speech synthesis,”
R. Sonobe, S. Takamichi, and H. Saruwatari, · 2017
Earlier work this paper cites.
“Natural TTS synthesis by conditioning WaveNet on mel spectrogram predictions,”
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Skerry-Ryan, R. A. Saurous, Y. Agiomvrgiannakis, and Y. Wu, · 2018
Earlier work this paper cites.
“Contextual string embeddings for sequence labeling,”
A. Akbik, D. Blythe, and R. Vollgraf, · 2018
Earlier work this paper cites.
“Style tokens: Unsupervised style modeling, control and transfer in end-to-end speech synthesis,”
Y. Wang, D. Stanton, Y. Zhang, R. Skerry-Ryan, E. Battenberg, J. Shor, Y. Xiao, Y. Jia, F. Ren, and R. A. Saurous, · 2018
Cited alongside, same era.
“Impacts of input linguistic feature representation on Japanese end-to-end speech synthesis,”
T. Fujimoto, K. Hashimoto, K. Oura, Y. Nankaku, and K. Tokuda, · 2019
Cited alongside, same era.
“Predicting prosodic prominence from text with pre-trained contextualized word representations,”
A. Talman, A. Suni, H. Celikkanat, S. Kakouros, J. Tiedemann, and M. Vainio, · 2019
Cited alongside, same era.
“Pre-trained text representations for improving front-end text processing in Mandarin text-to-speech synthesis,”
B. Yang, J. Zhong, and S. Liu, · 2019
Cited alongside, same era.
“Pre-trained text embeddings for enhanced text-to-speech synthesis,”
T. Hayashi, S. Watanabe, T. Toda, K. Takeda, S. Toshniwal, and K. Livescu, · 2019
Cited alongside, same era.
“Improving the prosody of RNN-based English text-to-speech synthesis by incorporating a BERT model,”
T. Kenter, M. Sharma, and R. Clark, · 2020
Later among the works it cites.
“g2pM: A neural grapheme-to-phoneme conversion package for Mandarin Chinese based on a new open benchmark dataset,”
K. Park and S. Lee, · 2020
Later among the works it cites.
R. Yamamoto, E. Song, and J. M. Kim, · 2020
Later among the works it cites.
“FastSpeech 2: Fast and high-quality end-to-end text to speech,”
Y. Ren, C. Hu, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T.-Y. Liu, · 2021
Later among the works it cites.
“Prosodic features control by symbols as input of sequence-to-sequence acoustic modeling for neural TTS,”
K. Kurihara, N. Seiyama, and T. Kumano, · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Knowledge distillation from BERT in pre-training and fine-tuning for polyphone disambiguation,”
H. Sun, X. Tan, J. W. Gan, S. Zhao, D. Han, H. Liu, T. Qin, and T. Y. Liu, · 2019
Cited alongside, same era.
“Polyphone disambiguation for Mandarin Chinese using conditional neural network with multi-level embedding features,”
Z. Cai, Y. Yang, C. Zhang, X. Qin, and M. Li, · 2019
Cited alongside, same era.
“BERT: Pre-training of deep bidirectional transformers for language understanding,”
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, · 2019
Cited alongside, same era.
“A structural probe for finding syntax in word representations,”
J. Hewitt and C. D. Manning, · 2019
Cited alongside, same era.
“Does BERT make any sense? interpretable word sense disambiguation with contextualized embeddings,”
G. Wiedemann, S. Remus, A. Chawla, and C. Biemann, · 2019
Cited alongside, same era.
“FLAIR: An easy-to-use framework for state-of-the-art NLP,”
A. Akbik, T. Bergmann, D. Blythe, K. Rasul, S. Schweter, and R. Vollgraf, · 2019
Cited alongside, same era.
ASJ Japanese Newspaper Article Sentences Read Speech Corpus (JNAS), http://research.nii.ac.jp/src/JNAS.html
Cited in the paper.
“Phrase break prediction with bidirectional encoder representations in Japanese text-to-speech synthesis,”
K. Futamata, B. Park, R. Yamamoto, and K. Tachibana, · 2021
Later among the works it cites.
“A universal BERT-based front-end model for Mandarin text-to-speech synthesis,”
Z. Bai and B. Hu, · 2021
Later among the works it cites.
“CAMP: A two-stage approach to modelling prosody in context,”
Z. Hodari, A. Moinet, S. Karlapati, J. Lorenzo-Trueba, T. Merritt, A. Joly, A. Abbas, P. Karanasou, and T. Drugman, · 2021
Later among the works it cites.
“PnG BERT: Augmented BERT on Phonemes and Graphemes for Neural TTS,”
Y. Jia, H. Zen, J. Shen, Y. Zhang, and Y. Wu, · 2021
Later among the works it cites.
“Phonetic and prosodic information estimation from texts for genuine Japanese end-to-end text-to-speech,”
N. Kakegawa, S. Hara, M. Abe, and Y. Ijima, · 2021
Later among the works it cites.