Fetching the paper…
Reading the bibliography…
Speech provides a natural way for human-computer interaction.
D. H. Klatt, “Software for a cascade/parallel formant synthesizer,”
1980
Earlier work this paper cites.
D. Griffin and J. Lim, “Signal estimation from modified short-time fourier transform,”
1984
Earlier work this paper cites.
I. Seara, “Estudo estatístico dos fonemas do português brasileiro falado na capital de santa catarina para elaboração de frases foneticamente balanceadas,” Ph.D. dissertation, Dissertação de Mestrado, Universidade Federal de Santa Catarina …, 1994
1994
Earlier work this paper cites.
K. Tokuda, T. Yoshimura, T. Masuko, T. Kobayashi, and T. Kitamura, “Speech parameter generation algorithms for hmm-based speech synthesis,” in
2000
Earlier work this paper cites.
J. P. Teixeira, D. Freitas, D. Braga, M. J. Barros, and V. Latsch, “Phonetic events from the labeling the european portuguese database for speech synthesis, feup/ipbdb,” in
2001
Earlier work this paper cites.
J. P. Teixeira, D. Freitas, and H. Fujisaki, “Prediction of fujisaki model’s phrase commands,” in
2003
Earlier work this paper cites.
X. Zhu, G. T. Beauregard, and L. L. Wyse, “Real-time signal estimation from modified short-time fourier transform magnitude spectra,”
2007
Earlier work this paper cites.
V. Alencar and A. Alcaim, “Lsf and lpc-derived features for large vocabulary distributed continuous speech recognition in brazilian portuguese,” in
2008
Earlier work this paper cites.
T. R. Gruber, “Siri, a virtual personal assistant-bringing intelligence to the interface,” in
2009
Earlier work this paper cites.
W. Y. Wang and K. Georgila, “Automatic detection of unnatural word-level segments in unit-selection speech synthesis,” in
2011
Earlier work this paper cites.
J. Benesty, J. Chen, and E. A. Habets,
2011
Earlier work this paper cites.
F. Ribeiro, D. Florêncio, C. Zhang, and M. Seltzer, “Crowdmos: An approach for crowdsourcing mean opinion score studies,” in
2011
Earlier work this paper cites.
H. Ze, A. Senior, and M. Schuster, “Statistical parametric speech synthesis using deep neural networks,” in
2013
Earlier work this paper cites.
D. A. Braude, H. Shimodaira, and A. B. Youssef, “Template-warping based speech driven head motion synthesis.” in
2013
Earlier work this paper cites.
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
A. Aroon and S. Dhonde, “Statistical parametric speech synthesis: A review,” in
2015
Earlier work this paper cites.
F. Yu and V. Koltun, “Multi-scale context aggregation by dilated convolutions,”
2015
Cited alongside, same era.
R. K. Srivastava, K. Greff, and J. Schmidhuber, “Training very deep networks,” in
2015
Cited alongside, same era.
I. Goodfellow, Y. Bengio, A. Courville, and Y. Bengio,
2016
Cited alongside, same era.
M. Morise, F. Yokomori, and K. Ozawa, “World: a vocoder-based high-quality speech synthesis system for real-time applications,”
2016
Cited alongside, same era.
2016
Cited alongside, same era.
A. Tamamori, T. Hayashi, K. Kobayashi, K. Takeda, and T. Toda, “Speaker-dependent wavenet vocoder,” in
2017
Later among the works it cites.
J.-M. Valin, “A hybrid dsp/deep learning approach to real-time full-band speech enhancement,”
2017
Later among the works it cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in
2017
Later among the works it cites.
F. Charpentier and M. Stella, “Diphone synthesis using an overlap-add technique for speech waveforms concatenation,” in
2018
Later among the works it cites.
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Skerrv-Ryan
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2016
Cited alongside, same era.
J. L. Ba, J. R. Kiros, and G. E. Hinton, “Layer normalization,”
2016
Cited alongside, same era.
D. P. Kingma, T. Salimans, R. Jozefowicz, X. Chen, I. Sutskever, and M. Welling, “Improved variational inference with inverse autoregressive flow,” in
2016
Cited alongside, same era.
A. Purington, J. G. Taft, S. Sannon, N. N. Bazarova, and S. H. Taylor, “” alexa is my new bff” social roles, user satisfaction, and personification of the amazon echo,” in
2017
Cited alongside, same era.
P. Dempsey, “The teardown: Google home personal assistant,”
2017
Cited alongside, same era.
2017
Cited alongside, same era.
D. Siddhi, J. M. Verghese, and D. Bhavik, “Survey on various methods of text to speech synthesis,”
2017
Cited alongside, same era.
K. Park, “A tensorflow implementation of dc-tts,” https://github.com/kyubyong/dc_tts, 2018
2018
Later among the works it cites.
E. Gölge, “Deep learning for text to speech,” https://github.com/mozilla/TTS, 2019
2019
Later among the works it cites.
2019
Later among the works it cites.
C. Durkan, A. Bekasov, I. Murray, and G. Papamakarios, “Neural spline flows,” in
2019
Later among the works it cites.
2020
Closest in time.
2020
Closest in time.
I. M. Quintanilha, S. L. Netto, and L. W. P. Biscainho, “An open-source end-to-end asr system for brazilian portuguese using dnns built from newly assembled corpora,”
2020
Closest in time.
V. Pratap, Q. Xu, A. Sriram, G. Synnaeve, and R. Collobert, “Mls: A large-scale multilingual dataset for speech research,”
2020
Closest in time.
K. Ito, “The lj speech dataset,” https://keithito.com/LJ-Speech-Dataset/, 2017, accessed: 2020-04-29
2020
Closest in time.
S. Quintas and I. Trancoso, “Evaluation of deep learning approaches to text-to-speech systems for european portuguese,” in
2020
Closest in time.
2020
Closest in time.
C. Miao, S. Liang, M. Chen, J. Ma, S. Wang, and J. Xiao, “Flow-tts: A non-autoregressive network for text to speech based on flow,” in
2020
Closest in time.