Fetching the paper…
Reading the bibliography…
Polyphone disambiguation aims to capture accurate pronunciation knowledge from natural text sequences for reliable Text-to-speech (TTS) systems.
Heteronyms and polyphones: Categories of words with multiple phonemic representations
Maryanne Martin, Gregory V Jones, Douglas L Nelson, and Louise Nelson · 1981
Earlier work this paper cites.
Issues in building general letter to sound rules
Alan W Black, Kevin Lenzo, and Vincent Pagel · 1998
Earlier work this paper cites.
Disambiguation of chinese polyphonic characters
Hong Zhang, JiangSheng Yu, WeiDong Zhan, and ShiWen Yu · 2001
Earlier work this paper cites.
Principles of generative phonology: An introduction
John T Jensen · 2004
Earlier work this paper cites.
Dynamic time warping
Meinard Müller · 2007
Earlier work this paper cites.
Disambiguating effectively chinese polyphonic ambiguity based on unify approach
Feng-Long Huang · 2008
Earlier work this paper cites.
Polyphonic word disambiguation with machine learning approaches
Jinke Liu, Weiguang Qu, Xuri Tang, Yizhe Zhang, and Yuxia Sun · 2010
Earlier work this paper cites.
Sequence-to-sequence neural net models for grapheme-to-phoneme conversion
Kaisheng Yao and Geoffrey Zweig · 2015
Earlier work this paper cites.
Categorical reparameterization with gumbel-softmax
Eric Jang, Shixiang Gu, and Ben Poole · 2016
Earlier work this paper cites.
A bi-directional lstm approach for polyphone disambiguation in mandarin chinese
Changhao Shan, Lei Xie, and Kaisheng Yao · 2016
Earlier work this paper cites.
Wavenet: A generative model for raw audio
Aäron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew W. Senior, and Koray Kavukcuoglu · 2016
Earlier work this paper cites.
Deep voice: Real-time neural text-to-speech
Sercan Ö Arık, Mike Chrzanowski, Adam Coates, Gregory Diamos, Andrew Gibiansky, Yongguo Kang, Xian Li, John Miller, Andrew Ng, Jonathan Raiman, et al · 2017
Earlier work this paper cites.
Chinese standard mandarin speech corpus, 2017
Data Baker · 2017
Earlier work this paper cites.
Jsut corpus: free large-scale japanese speech corpus for end-to-end speech synthesis
Ryosuke Sonobe, Shinnosuke Takamichi, and Hiroshi Saruwatari · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Tacotron: Towards end-to-end speech synthesis
Yuxuan Wang, RJ Skerry-Ryan, Daisy Stanton, Yonghui Wu, Ron J Weiss, Navdeep Jaitly, Zongheng Yang, Ying Xiao, Zhifeng Chen, Samy Bengio, et al · 2017
Earlier work this paper cites.
Convolutional sequence to sequence model with non-sequential greedy decoding for grapheme to phoneme conversion
Moon-Jung Chae, Kyubyong Park, Jinhyun Bang, Soobin Suh, Jonghyuk Park, Namju Kim, and Longhun Park · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Learning pronunciation from a foreign language in speech synthesis networks
Younggun Lee, Suwon Shon, and Taesu Kim · 2018
Earlier work this paper cites.
Dissecting contextual word embeddings: Architecture and representation
Matthew E Peters, Mark Neumann, Luke Zettlemoyer, and Wen-tau Yih · 2018
Earlier work this paper cites.
Self-attention with relative position representations
Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani · 2018
Earlier work this paper cites.
Natural tts synthesis by conditioning wavenet on mel spectrogram predictions
Jonathan Shen, Ruoming Pang, Ron J Weiss, Mike Schuster, Navdeep Jaitly, Zongheng Yang, Zhifeng Chen, Yu Zhang, Yuxuan Wang, Rj Skerrv-Ryan, et al · 2018
Earlier work this paper cites.
Polyphone disambiguation for mandarin chinese using conditional neural network with multi-level embedding features
Zexin Cai, Yaogen Yang, Chuxiong Zhang, Xiaoyi Qin, and Ming Li · 2019
Cited alongside, same era.
What does bert look at? an analysis of bert’s attention
Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D Manning · 2019
Cited alongside, same era.
Disambiguation of chinese polyphones in an end-to-end framework with semantic features extracted by pre-trained bert
Dongyang Dai, Zhiyong Wu, Shiyin Kang, Xixin Wu, Jia Jia, Dan Su, Dong Yu, and Helen Meng · 2019
Cited alongside, same era.
Pre-trained text embeddings for enhanced text-to-speech synthesis
Tomoki Hayashi, Shinji Watanabe, Tomoki Toda, Kazuya Takeda, Shubham Toshniwal, and Karen Livescu · 2019
Cited alongside, same era.
A structural probe for finding syntax in word representations
John Hewitt and Christopher D Manning · 2019
Cited alongside, same era.
Unified mandarin tts front-end based on distilled bert model
Yang Zhang, Liqun Deng, and Yasheng Wang · 2020
Later among the works it cites.
End-to-end adversarial text-to-speech
Jeff Donahue, Sander Dieleman, Mikolaj Binkowski, Erich Elsen, and Karen Simonyan · 2021
Later among the works it cites.
Parallel tacotron 2: A non-autoregressive neural TTS model with differentiable duration modeling
Isaac Elias, Heiga Zen, Jonathan Shen, Yu Zhang, Ye Jia, R. J. Skerry-Ryan, and Yonghui Wu · 2021
Later among the works it cites.
Mutian He, Jingzhou Yang, Lei He, and Frank K Soong · 2021
Later among the works it cites.
Png BERT: augmented BERT on phonemes and graphemes for neural TTS
Ye Jia, Heiga Zen, Jonathan Shen, Yu Zhang, and Yonghui Wu · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Generalization through memorization: Nearest neighbor language models
Urvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer, and Mike Lewis · 2019
Cited alongside, same era.
Neural speech synthesis with transformer network
Naihan Li, Shujie Liu, Yanqing Liu, Sheng Zhao, and Ming Liu · 2019
Cited alongside, same era.
Fastspeech: Fast, robust and controllable text to speech
Yi Ren, Yangjun Ruan, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu · 2019
Cited alongside, same era.
Token-level ensemble distillation for grapheme-to-phoneme conversion
Hao Sun, Xu Tan, Jun-Wei Gan, Hongzhi Liu, Sheng Zhao, Tao Qin, and Tie-Yan Liu · 2019
Cited alongside, same era.
Knowledge distillation from bert in pre-training and fine-tuning for polyphone disambiguation
Hao Sun, Xu Tan, Jun-Wei Gan, Sheng Zhao, Dongxu Han, Hongzhi Liu, Tao Qin, and Tie-Yan Liu · 2019
Cited alongside, same era.
Bert rediscovers the classical nlp pipeline
Ian Tenney, Dipanjan Das, and Ellie Pavlick · 2019
Cited alongside, same era.
Pre-trained text representations for improving front-end text processing in mandarin text-to-speech synthesis
Bing Yang, Jiaqi Zhong, and Shan Liu · 2019
Cited alongside, same era.
Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech
Jaehyeon Kim, Jungil Kong, and Juhee Son · 2021
Later among the works it cites.
Lexicon enhanced chinese sequence labeling using bert adapter
Wei Liu, Xiyan Fu, Yue Zhang, and Wenming Xiao · 2021
Later among the works it cites.
Fastspeech 2: Fast and high-quality end-to-end text to speech
Yi Ren, Chenxu Hu, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu · 2021
Later among the works it cites.
Portaspeech: Portable and high-quality generative text-to-speech
Yi Ren, Jinglin Liu, and Zhou Zhao · 2021
Later among the works it cites.
A survey on neural speech synthesis
Xu Tan, Tao Qin, Frank Soong, and Tie-Yan Liu · 2021
Later among the works it cites.
Joint alignment learning-attention based model for grapheme-to-phoneme conversion
Yonghe Wang, Feilong Bao, Hui Zhang, and Guanglai Gao · 2021
Later among the works it cites.
Using prior knowledge to guide bert’s attention in semantic textual matching tasks
Tingyu Xia, Yue Wang, Yuan Tian, and Yi Chang · 2021
Later among the works it cites.
g2pw: A conditional weighted softmax bert for polyphone disambiguation in mandarin
Yi-Chang Chen, Yu-Chuan Chang, Yen-Cheng Chang, and Yi-Ren Yeh · 2022
Closest in time.
Rem Hida, Masaki Hamada, Chie Kamada, Emiru Tsunoo, Toshiyuki Sekiya, and Toshiyuki Kumakura · 2022
Closest in time.
Fastdiff: A fast conditional diffusion model for high-quality speech synthesis
Rongjie Huang, Max WY Lam, Jun Wang, Dan Su, Dong Yu, Yi Ren, and Zhou Zhao · 2022
Closest in time.
Generspeech: Towards style transfer for generalizable out-of-domain text-to-speech synthesis
Rongjie Huang, Yi Ren, Jinglin Liu, Chenye Cui, and Zhou Zhao · 2022
Closest in time.
Prodiff: Progressive fast diffusion model for high-quality text-to-speech
Rongjie Huang, Zhou Zhao, Huadai Liu, Jinglin Liu, Chenye Cui, and Yi Ren · 2022
Closest in time.
Diffsinger: Singing voice synthesis via shallow diffusion mechanism
Jinglin Liu, Chengxi Li, Yi Ren, Feiyang Chen, and Zhou Zhao · 2022
Closest in time.
SyntaSpeech: Syntax-Aware Generative Adversarial Text-to-Speech
Zhenhui Ye, Zhou Zhao, Yi Ren, and Fei Wu · 2022
Closest in time.
Wenetspeech: A 10000+ hours multi-domain mandarin corpus for speech recognition
Binbin Zhang, Hang Lv, Pengcheng Guo, Qijie Shao, Chao Yang, Lei Xie, Xin Xu, Hui Bu, Xiaoyu Chen, Chenchen Zeng, et al · 2022
Closest in time.
Guangyan Zhang, Kaitao Song, Xu Tan, Daxin Tan, Yuzi Yan, Yanqing Liu, Gang Wang, Wei Zhou, Tao Qin, Tan Lee, et al · 2022
Closest in time.