Fetching the paper…
Reading the bibliography…
While recent text to speech (TTS) models perform very well in synthesizing reading-style (e.g., audiobook) speech, it is still challenging to synthesize spontaneous-style speech (e.g., podcast or conversation), mainly because of two reasons: 1) the lack of training data for spontaneous speech; 2) the difficulty in modeling the filled pauses (um and uh) and diverse rhythms in spontaneous speech.
A. L. Winkworth, P. J. Davis, E. Ellis, and R. D. Adams, “Variability and consistency in speech breathing during reading: Lung volumes, speech intensity, and linguistic factors,” Journal of Speech, Language, and Hearing Research , vol. 37, no. 3, pp. 535–556, 1994
1994
Earlier work this paper cites.
R. Avnimelech and N. Intrator, “Boosted mixture of experts: An ensemble learning scheme,” Neural computation , vol. 11, no. 2, pp. 483–497, 1999
1999
Earlier work this paper cites.
S. Sundaram and S. Narayanan, “An empirical text transformation method for spontaneous speech synthesizers,” in Eighth European Conference on Speech Communication and Technology , 2003
2003
Earlier work this paper cites.
M. Viswanathan and M. Viswanathan, “Measuring speech quality for text-to-speech systems: development and assessment of a modified mean opinion score (mos) scale,” Computer Speech & Language , vol. 19, no. 1, pp. 55–83, 2005
2005
Earlier work this paper cites.
J. Yamagishi and T. Kobayashi, “Average-voice-based speech synthesis using hsmm-based speaker adaptation and adaptive training,” IEICE TRANSACTIONS on Information and Systems , vol. 90, no. 2, pp. 533–543, 2007
2007
Earlier work this paper cites.
S. E. Yuksel, J. N. Wilson, and P. D. Gader, “Twenty years of mixture of experts,” IEEE transactions on neural networks and learning systems , vol. 23, no. 8, pp. 1177–1193, 2012
2012
Earlier work this paper cites.
A. Rochet-Capellan and S. Fuchs, “Take a breath and take the turn: how breathing meets turns in spontaneous dialogue,” Philosophical Transactions of the Royal Society B: Biological Sciences , vol. 369, no. 1658, p. 20130399, 2014
2014
Earlier work this paper cites.
S. Masoudnia and R. Ebrahimpour, “Mixture of experts: a literature survey,” Artificial Intelligence Review , vol. 42, no. 2, pp. 275–293, 2014
2014
Earlier work this paper cites.
M. Wester, O. Watts, and G. E. Henter, “Evaluating comprehension of natural and synthetic conversational speech,” in Proc. Speech Prosody , vol. 8, 2016, pp. 736–740
2016
Earlier work this paper cites.
C. Veaux, J. Yamagishi, K. MacDonald et al. , “Superseded-cstr vctk corpus: English multi-speaker corpus for cstr voice cloning toolkit,” 2016
2016
Earlier work this paper cites.
2017
Cited alongside, same era.
E. Székely, J. Mendelson, and J. Gustafson, “Synthesising uncertainty: The interplay of vocal effort and hesitation disfluencies.” in INTERSPEECH , 2017, pp. 804–808
2017
Cited alongside, same era.
T. Nagata, H. Mori, and T. Nose, “Dimensional paralinguistic information control based on multiple-regression hsmm for spontaneous dialogue speech synthesis with robust parameter estimation,” Speech Communication , vol. 88, pp. 137–148, 2017
2017
Cited alongside, same era.
R. Dall, “Statistical parametric speech synthesis using conversational data and phenomena,” 2017
2017
Cited alongside, same era.
É. Székely, G. E. Henter, and J. Gustafson, “Casting to corpus: Segmenting and selecting spontaneous dialogue for tts with a cnn-lstm speaker-dependent breath detector,” in ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2019, pp. 6925–6929
2019
Later among the works it cites.
H. Sun, X. Tan, J.-W. Gan, H. Liu, S. Zhao, T. Qin, and T.-Y. Liu, “Token-level ensemble distillation for grapheme-to-phoneme conversion,” in INTERSPEECH , 2019
2019
Later among the works it cites.
2019
Later among the works it cites.
K. Kumar, R. Kumar, T. de Boissiere, L. Gestin, W. Z. Teoh, J. Sotelo, A. de Brébisson, Y. Bengio, and A. C. Courville, “Melgan: Generative adversarial networks for conditional waveform synthesis,” in NIPS , 2019, pp. 14 910–14 921
2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M. McAuliffe, M. Socolof, S. Mihuc, M. Wagner, and M. Sonderegger, “Montreal forced aligner: Trainable text-speech alignment using kaldi.” in Interspeech , 2017, pp. 498–502
2017
Cited alongside, same era.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in NIPS , 2017, pp. 5998–6008
2017
Cited alongside, same era.
W. Ping, K. Peng, A. Gibiansky, S. O. Arik, A. Kannan, S. Narang, J. Raiman, and J. Miller, “Deep voice 3: 2000-speaker neural text-to-speech,” in International Conference on Learning Representations , 2018
2018
Cited alongside, same era.
Y. Ren, Y. Ruan, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T.-Y. Liu, “Fastspeech: Fast, robust and controllable text to speech,” in NIPS , 2019, pp. 3165–3174
2019
Cited alongside, same era.
É. Székely, G. E. Henter, J. Beskow, and J. Gustafson, “Spontaneous conversational speech synthesis from found data,” in Interspeech , 2019
2019
Cited alongside, same era.
Later among the works it cites.
É. Székely, G. E. Henter, J. Beskow, and J. Gustafson, “Breathing and speech planning in spontaneous speech synthesis,” in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 7649–7653
2020
Later among the works it cites.
2021
Closest in time.
Y. Ren, C. Hu, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T.-Y. Liu, “Fastspeech 2: Fast and high-quality end-to-end text to speech,” in International Conference on Learning Representations , 2021. [Online]. Available: https://openreview.net/forum?id=piLPYqxtWuA
2021
Closest in time.
M. Chen, X. Tan, B. Li, Y. Liu, T. Qin, S. Zhao, and T.-Y. Liu, “Adaspeech: Adaptive text to speech for custom voice,” in International Conference on Learning Representations , 2021. [Online]. Available: https://openreview.net/forum?id=Drynvt7gg4L
2021
Closest in time.
Y. Yan, X. Tan, B. Li, T. Qin, S. Zhao, Y. Shen, and T.-Y. Liu, “Adaspeech 2: Adaptive text to speech with untranscribed data,” in ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2021, pp. 6613–6617
2021
Closest in time.