Fetching the paper…
Reading the bibliography…
This paper introduces Parallel Tacotron 2, a non-autoregressive neural text-to-speech model with a fully differentiable duration model which does not require supervised duration signals.
R. J. Williams and D. Zipser, “A Learning Algorithm for Continually Running Fully Recurrent Neural Networks,” Neural Computation , vol. 1, no. 2, pp. 270–280, 1989
1989
Earlier work this paper cites.
D. Bahdanau, K. Cho, and Y. Bengio, “Neural Machine Translation by Jointly Learning to Align and Translate,” in Proc. ICLR , 2015
2015
Earlier work this paper cites.
2016
Earlier work this paper cites.
J. Sotelo, S. Mehri, K. Kumar, J. F. Santos, K. Kastner, A. C. Courville, and Y. Bengio, “Char2Wav: End-to-End Speech Synthesis,” in Proc. ICLR , 2017
2017
Earlier work this paper cites.
Y. Wang, R. Skerry-Ryan, D. Stanton, Y. Wu, R. J. Weiss, N. Jaitly, Z. Yang, Y. Xiao, Z. Chen, S. Bengio, Q. Le, Y. Agiomyrgiannakis, R. Clark, and R. A. Saurous, “Tacotron: Towards End-to-End Speech Synthesis,” in Proc. Interspeech , 2017, pp. 4006–4010
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention Is All You Need,” in Proc. NeurIPS , 2017
2017
Earlier work this paper cites.
A. Oord, Y. Li, I. Babuschkin, K. Simonyan, O. Vinyals, K. Kavukcuoglu, G. Driessche, E. Lockhart, L. Cobo, F. Stimberg et al. , “Parallel WaveNet: Fast high-fidelity speech synthesis,” in Prof. ICML , 2018, pp. 3918–3926
2018
Earlier work this paper cites.
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Skerrv-Ryan, R. A. Saurous, Y. Agiomyrgiannakis, and Y. Wu, “Natural TTS Synthesis by Conditioning WaveNet on Mel Spectrogram Predictions,” in Proc. ICASSP , 2018
2018
Earlier work this paper cites.
N. Kalchbrenner, E. Elsen, K. Simonyan, S. Noury, N. Casagrande, E. Lockhart, F. Stimberg, A. van den Oord, S. Dieleman, and K. Kavukcuoglu, “Efficient Neural Audio Synthesis,” in Proc. ICML , 2018, pp. 2410–2419
2018
Earlier work this paper cites.
M. Cuturi and M. Blondel, “Soft-DTW: a Differentiable Loss Function for Time-Series,” 2018
2018
Cited alongside, same era.
2018
Cited alongside, same era.
N. Li, S. Liu, Y. Liu, S. Zhao, and M. Liu, “Neural Speech Synthesis with Transformer Network,” in Proc. AAAI , vol. 33, 2019, pp. 6706–6713
2019
Cited alongside, same era.
M. He, Y. Deng, and L. He, “Robust sequence-to-sequence acoustic modeling with stepwise monotonic attention for neural TTS,” in Proc. Interspeech , 2019, pp. 1293–1297
2019
Cited alongside, same era.
Y. Zheng, J. Tao, W. Zhengqi, and J. Yi, “Forward–backward decoding sequence for regularizing end-to-end TTS,” IEEE/ACM Trans. Audio Speech & Lang. Process. , vol. 27, no. 12, pp. 2067–2079, 2019
2020
Later among the works it cites.
I. Elias, H. Zen, J. Shen, Y. Zhang, Y. Jia, R. Weiss, and Y. Wu, “Parallel Tacotron: Non-autoregressive and controllable TTS,” 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
Z. Zeng, J. Wang, N. Cheng, T. Xia, and J. Xiao, “AlignTTS: Efficient Feed-Forward Text-to-Speech System without Explicit Alignment,” in Proc. ICASSP , 2020, pp. 6714–6718
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2019
Cited alongside, same era.
H. Guo, F. K. Soong, L. He, and L. Xie, “A new GAN-based end-to-end TTS training algorithm,” in Proc. Interspeech , 2019, pp. 1288–1292
2019
Cited alongside, same era.
W.-N. Hsu, Y. Zhang, R. J. Weiss, H. Zen, Y. Wu, Y. Wang, Y. Cao, Y. Jia, Z. Chen, J. Shen, P. Nguyen, and R. Pang, “Hierarchical Generative Modeling for Controllable Speech Synthesis,” in Proc. ICLR , 2019
2019
Cited alongside, same era.
2019
Cited alongside, same era.
F. Wu, A. Fan, A. Baevski, Y. N. Dauphin, and M. Auli, “Pay Less Attention with Lightweight and Dynamic Convolutions,” in Proc. ICLR , 2019
2019
Cited alongside, same era.
2020
Cited alongside, same era.
2020
Later among the works it cites.
C. Miao, S. Liang, Z. Liu, M. Chen, J. Ma, S. Wang, and J. Xiao, “EfficientTTS: An Efficient and High-Quality Text-to-Speech Architecture,” 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
C. Miao, S. Liang, M. Chen, J. Ma, S. Wang, and J. Xiao, “Flow-TTS: A Non-Autoregressive Network for Text to Speech Based on Flow,” in Proc. ICASSP , 2020, pp. 7209–7213
2020
Later among the works it cites.
A. Tjandra, C. Liu, F. Zhang, X. Zhang, Y. Wang, G. Synnaeve, S. Nakamura, and G. Zweig, “DEJA-VU: Double Feature Presentation and Iterated Loss in Deep Transformer Networks,” in Proc. ICASSP , 2020, pp. 6899–6903
2020
Later among the works it cites.
T. Kenter, M. K. Sharma, and R. Clark, “Improving Prosody of RNN-based English Text-To-Speech Synthesis by Incorporating a BERT model,” in Proc. Interspeech , 2020
2020
Later among the works it cites.