Fetching the paper…
Reading the bibliography…
End-to-end TTS requires a large amount of speech/text paired data to cover all necessary knowledge, particularly how to pronounce different words in diverse contexts, so that a neural model may learn such knowledge accordingly.
M. J. Seo, A. Kembhavi, A. Farhadi, and H. Hajishirzi, “Bidirectional attention flow for machine comprehension,” in Proc. ICLR 2017 . OpenReview.net, 2017
2017
Earlier work this paper cites.
D. Chen, A. Fisch, J. Weston, and A. Bordes, “Reading Wikipedia to answer open-domain questions,” in Proc. ACL 2017 . ACL, 2017, pp. 1870–1879
2017
Earlier work this paper cites.
Y. Feng, S. Zhang, A. Zhang, D. Wang, and A. Abel, “Memory-augmented neural machine translation,” in Proc. EMNLP 2017 . ACL, 2017, pp. 1390–1399
2017
Earlier work this paper cites.
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Ryan, R. A. Saurous, Y. Agiomyrgiannakis, and Y. Wu, “Natural TTS synthesis by conditioning wavenet on MEL spectrogram predictions,” in Proc. ICASSP 2018 . IEEE, 2018, pp. 4779–4783
2018
Earlier work this paper cites.
L. Bauer, Y. Wang, and M. Bansal, “Commonsense for generative multi-hop question answering tasks,” in Proc. EMNLP 2018 . ACL, 2018, pp. 4220–4230
2018
Earlier work this paper cites.
J. Ni, Y. Shiga, and H. Kawai, “Multilingual grapheme-to-phoneme conversion with global character vectors,” in Proc. Interspeech 2018 . ISCA, 2018, pp. 2823–2827
2018
Earlier work this paper cites.
K. Gorman, G. Mazovetskiy, and V. Nikolaev, “Improving homograph disambiguation with supervised machine learning,” in Proc. LREC 2018 . European Language Resources Association (ELRA), 2018
2018
Earlier work this paper cites.
A. Bruguier, A. Bakhtin, and D. Sharma, “Dictionary augmented sequence-to-sequence neural network for grapheme to phoneme prediction,” in Proc. Interspeech 2018 . ISCA, 2018, pp. 3733–3737
2018
Earlier work this paper cites.
J. Fong, J. Taylor, K. Richmond, and S. King, “A comparison between letters and phones as input to sequence-to-sequence models for speech synthesis,” in The 10th ISCA Speech Synthesis Workshop . ISCA, 2019, pp. 223–227
2019
Earlier work this paper cites.
F. Petroni, T. Rocktäschel, S. Riedel, P. S. H. Lewis, A. Bakhtin, Y. Wu, and A. H. Miller, “Language models as knowledge bases?” in Proc. EMNLP/IJCNLP 2019 . ACL, 2019, pp. 2463–2473
2019
Earlier work this paper cites.
C. Wang and H. Jiang, “Explicit utilization of general knowledge in machine reading comprehension,” in Proc. ACL 2019 . ACL, 2019, pp. 2263–2272
2019
Cited alongside, same era.
L. Huang, C. Sun, X. Qiu, and X. Huang, “GlossBERT: BERT for word sense disambiguation with gloss knowledge,” in Proc. EMNLP-IJCNLP 2019 . ACL, 2019, pp. 3507–3512
2019
Cited alongside, same era.
S. Alqahtani, H. Aldarmaki, and M. T. Diab, “Homograph disambiguation through selective diacritic restoration,” in Proc. Fourth Arabic Natural Language Processing Workshop, WANLP@ACL 2019 . ACL, 2019, pp. 49–59
2019
Cited alongside, same era.
T. Hayashi, S. Watanabe, T. Toda, K. Takeda, S. Toshniwal, and K. Livescu, “Pre-trained text embeddings for enhanced text-to-speech synthesis,” in Proc. Interspeech 2019 . ISCA, 2019, pp. 4430–4434
2019
Cited alongside, same era.
2021
Closest in time.
J. Taylor and K. Richmond, “Confidence intervals for ASR-based TTS evaluation,” in Proc. Interspeech 2021 . ISCA, 2021, pp. 2791–2795
2021
Closest in time.
W. Liu, X. Fu, Y. Zhang, and W. Xiao, “Lexicon enhanced Chinese sequence labeling using BERT adapter,” in Proc. ACL/IJCNLP 2021 . ACL, 2021, pp. 5847–5858
2021
Closest in time.
Y. Shi, C. Wang, Y. Chen, and B. Wang, “Polyphone disambiguation in Mandarin Chinese with semi-supervised learning,” in Proc. Interspeech 2021 . ISCA, 2021, pp. 4109–4113
2021
Closest in time.
Y. Jia, H. Zen, J. Shen, Y. Zhang, and Y. Wu, “PnG BERT: Augmented BERT on phonemes and graphemes for neural TTS,” in Proc. Interspeech 2021 . ISCA, 2021, pp. 151–155
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Conneau, K. Khandelwal, N. Goyal, V. Chaudhary, G. Wenzek, F. Guzmán, E. Grave, M. Ott, L. Zettlemoyer, and V. Stoyanov, “Unsupervised cross-lingual representation learning at scale,” in Proc. ACL 2020 . ACL, 2020, pp. 8440–8451
2020
Cited alongside, same era.
2020
Cited alongside, same era.
M. Yu, H. D. Nguyen, A. Sokolov, J. Lepird, K. M. Sathyendra, S. Choudhary, A. Mouchtaris, and S. Kunzmann, “Multilingual grapheme-to-phoneme conversion with byte representation,” in Proc. ICASSP 2020 . IEEE, 2020, pp. 8234–8238
2020
Cited alongside, same era.
H. Zhang, H. Pan, and X. Li, “A mask-based model for Mandarin Chinese polyphone disambiguation,” in Proc. Interspeech 2020 . ISCA, 2020, pp. 1728–1732
2020
Cited alongside, same era.
Y. Xiao, L. He, H. Ming, and F. K. Soong, “Improving prosody with linguistic and BERT derived features in multi-speaker based Mandarin Chinese neural TTS,” in Proc. ICASSP 2020 . IEEE, 2020, pp. 6704–6708
2020
Cited alongside, same era.
Y. Yasuda, X. Wang, and J. Yamagishi, “Investigation of learning abilities on linguistic features in sequence-to-sequence text-to-speech synthesis,” Comput. Speech Lang. , vol. 67, p. 101183, 2021
2021
Cited alongside, same era.
2021
Closest in time.
M. Nicolis and V. Klimkov, “Homograph disambiguation with contextual word embeddings for TTS systems,” in The 11th ISCA Speech Synthesis Workshop . ISCA, 2021, pp. 222–226
2021
Closest in time.
2022
Closest in time.
Y. Zou, L. Dong, and B. Xu, “Boosting character-based Chinese speech synthesis via multi-task learning and dictionary tutoring,” in Proc. Interspeech 2019 . ISCA, 2019, pp. 2055–2059
2059
Closest in time.
J. Taylor and K. Richmond, “Analysis of pronunciation learning in end-to-end speech synthesis,” in Proc. Interspeech 2019 . ISCA, 2019, pp. 2070–2074
2074
Closest in time.
D. Dai, Z. Wu, S. Kang, X. Wu, J. Jia, D. Su, D. Yu, and H. Meng, “Disambiguation of Chinese polyphones in an end-to-end framework with semantic features extracted by pre-trained BERT,” in Proc. Interspeech 2019 . ISCA, 2019, pp. 2090–2094
2094
Closest in time.