Fetching the paper…
Reading the bibliography…
In spoken conversations, spontaneous behaviors like filled pause and prolongations always happen.
K. Tokuda, T. Yoshimura, T. Masuko, T. Kobayashi, and T. Kitamura, “Speech parameter generation algorithms for HMM-based speech synthesis,” in
2000
Earlier work this paper cites.
A. W. Black, H. Zen, and K. Tokuda, “Statistical parametric speech synthesis,” in
2007
Earlier work this paper cites.
T. Koriyama, T. Nose, and T. Kobayashi, “On the use of extended context for hmm-based spontaneous conversational speech synthesis,” in
2011
Earlier work this paper cites.
H. Ze, A. Senior, and M. Schuster, “Statistical parametric speech synthesis using deep neural networks,” in
2013
Earlier work this paper cites.
R. Dall, M. Tomalin, M. Wester, W. J. Byrne, and S. King, “Investigating automatic & human filled pause insertion for speech synthesis,” in
2014
Earlier work this paper cites.
Z.-H. Ling, S.-Y. Kang, H. Zen, A. Senior, M. Schuster, X.-J. Qian, H. M. Meng, and L. Deng, “Deep learning for acoustic modeling in parametric speech generation: A systematic review of existing techniques and future trends,”
2015
Earlier work this paper cites.
M. Tomalin, M. Wester, R. Dall, W. Byrne, and S. King, “A lattice-based approach to automatic filled pause insertion,” in
2015
Earlier work this paper cites.
R. Levitan, S. Benus, A. Gravano, and J. Hirschberg, “Entrainment and turn-taking in human-human dialogue.” in
2015
Earlier work this paper cites.
S. Ö. Arik, M. Chrzanowski, A. Coates, G. F. Diamos, A. Gibiansky, Y. Kang, X. Li, J. Miller, A. Y. Ng, J. Raiman, S. Sengupta, and M. Shoeybi, “Deep voice: Real-time neural text-to-speech,” in
2017
Earlier work this paper cites.
Y. Wang, R. J. Skerry-Ryan, D. Stanton, Y. Wu, R. J. Weiss, N. Jaitly, Z. Yang, Y. Xiao, Z. Chen, S. Bengio, Q. V. Le, Y. Agiomyrgiannakis, R. Clark, and R. A. Saurous, “Tacotron: Towards end-to-end speech synthesis,” in
2017
Cited alongside, same era.
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Ryan, R. A. Saurous, Y. Agiomyrgiannakis, and Y. Wu, “Natural TTS synthesis by conditioning wavenet on MEL spectrogram predictions,” in
2018
Cited alongside, same era.
R. J. Skerry-Ryan, E. Battenberg, Y. Xiao, Y. Wang, D. Stanton, J. Shor, R. J. Weiss, R. Clark, and R. A. Saurous, “Towards end-to-end prosody transfer for expressive speech synthesis with tacotron,” in
2018
Cited alongside, same era.
Y. Wang, D. Stanton, Y. Zhang, R. J. Skerry-Ryan, E. Battenberg, J. Shor, Y. Xiao, Y. Jia, F. Ren, and R. A. Saurous, “Style tokens: Unsupervised style modeling, control and transfer in end-to-end speech synthesis,” in
2018
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” in
2019
Later among the works it cites.
2019
Later among the works it cites.
T. Hayashi, S. Watanabe, T. Toda, K. Takeda, S. Toshniwal, and K. Livescu, “Pre-trained text embeddings for enhanced text-to-speech synthesis,” in
2019
Later among the works it cites.
Y. Yamashita, T. Koriyama, Y. Saito, S. Takamichi, Y. Ijima, R. Masumura, and H. Saruwatari, “Investigating effective additional contextual factors in dnn-based spontaneous speech synthesis,” in
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
N. Kalchbrenner, E. Elsen, K. Simonyan, S. Noury, N. Casagrande, E. Lockhart, F. Stimberg, A. van den Oord, S. Dieleman, and K. Kavukcuoglu, “Efficient neural audio synthesis,” in
2018
Cited alongside, same era.
N. Li, S. Liu, Y. Liu, S. Zhao, M. Liu, and M. Zhou, “Close to Human Quality TTS with Transformer,” in
2019
Cited alongside, same era.
Y. Zhang, S. Pan, L. He, and Z. Ling, “Learning latent representations for style control and transfer in end-to-end speech synthesis,” in
2019
Cited alongside, same era.
2019
Cited alongside, same era.
É. Székely, G. E. Henter, and J. Gustafson, “Casting to corpus: Segmenting and selecting spontaneous dialogue for tts with a cnn-lstm speaker-dependent breath detector,” in
2019
Cited alongside, same era.
2020
Later among the works it cites.
E. Battenberg, R. J. Skerry-Ryan, S. Mariooryad, D. Stanton, D. Kao, M. Shannon, and T. Bagby, “Location-relative attention mechanisms for robust long-form speech synthesis,” in
2020
Later among the works it cites.
Y. Lei, S. Yang, and L. Xie, “Fine-grained emotion strength transfer, control and prediction for emotional speech synthesis,” in
2021
Closest in time.
H. Guo, S. Zhang, F. K. Soong, L. He, and L. Xie, “Conversational end-to-end TTS for voice agents,” in
2021
Closest in time.
C. Yu, H. Lu, N. Hu, M. Yu, C. Weng, K. Xu, P. Liu, D. Tuo, S. Kang, G. Lei, D. Su, and D. Yu, “Durian: Duration informed attention network for speech synthesis,” in
2031
Closest in time.