Fetching the paper…
Reading the bibliography…
Neural end-to-end text-to-speech (TTS) , which adopts either a recurrent model, e.g.
S. Hochreiter, Y. Bengio, P. Frasconi, and J. Schmidhuber, “Gradient flow in recurrent nets: the difficulty of learning long-term dependencies,” in
2001
Earlier work this paper cites.
P. C. Loizou, “Speech quality assessment,”
2011
Earlier work this paper cites.
A. Graves, “Sequence transduction with recurrent neural networks,” in
2012
Earlier work this paper cites.
2016
Earlier work this paper cites.
N. Jaitly, D. Sussillo, Q. V. Le, O. Vinyals, I. Sutskever, and S. Bengio, “An online sequence-to-sequence model using partial conditioning,” in
2016
Earlier work this paper cites.
W. Chan, N. Jaitly, Q. Le, and O. Vinyals, “Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,” in
2016
Earlier work this paper cites.
Y. Wang, R. Skerry-Ryan, D. Stanton, Y. Wu, R. J. Weiss, N. Jaitly, Z. Yang, Y. Xiao, Z. Chen, and S. B. et al., “Tacotron: Towards end-to-end speech synthesis,” in
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” in
2017
Earlier work this paper cites.
C.-C. Chiu and C. Raffel, “Monotonic chunkwise attention,” in
2017
Cited alongside, same era.
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, and R. S.-R. et al., “Natural tts synhtesis by conditioning wavenet on mel spectrogram predictions,” in
2018
Cited alongside, same era.
J.-X. Zhang, Z.-H. Ling, and L.-R. Dai, “Forward attention in sequenceto-sequence acoustic modeling for speech synthesis,” in
2018
Cited alongside, same era.
P. Shaw, J. Uszkoreit, and A. Vaswani, “Self-attention with relative position representations,” in
2018
Cited alongside, same era.
C.-Z. A. Huang, A. Vaswani, J. Uszkoreit, I. Simon, C. Hawthorne, N. Shazeer, A. M. Dai, M. D. Hoffman, M. Dinculescu, and D. Eck, “Music transformer: Generating music with long-term structure,” in
K. Kumar, R. Kumar, T. de Boissiere, L. Gestin, W. Z. Teoh, A. d. B. J. Sotelo, Y. Bengio, and A. Courville, “Melgan: Generative adversarial networks for conditional waveform synthesis,” in
2019
Later among the works it cites.
M. He, Y. Deng, and L. He, “Robust sequence-to-sequence acoustic modeling with stepwise monotonic attention for neural tts,” in
2019
Later among the works it cites.
2019
Later among the works it cites.
N. Moritz, T. Hori, and J. L. Roux, “Triggered attention for end-to-end speech recognition,” in
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2018
Cited alongside, same era.
N. Li, S. Liu, Y. Liu, S. Zhao, M. Liu, and M. Zhou, “Neural speech synthesis with transformer network,” in
2019
Cited alongside, same era.
Y. Ren, Y. Ruan, X. Tan, Q. Tao, S. Zhao, Z. Zhao, and T. Liu, “Fastspeech: Fast, robust and controllable text to speech,” in
2019
Cited alongside, same era.
J. Valin and J. Skoglund, “Lpcnet: Improving neural speech synthesis through linear prediction,” in
2019
Cited alongside, same era.
2019
Later among the works it cites.
——, “Streaming automatic speech recognition with the transformer model,” in
2020
Closest in time.
T. N. Sainath, Y. He, B. Li, A. Narayanan, R. Pang, A. Bruguier, S. Chang, and et al., “A streaming on-device end-to-end model surpassing server-side conventional model quality and latency,” in
2020
Closest in time.