Fetching the paper…
Reading the bibliography…
Speech-to-text alignment is a critical component of neural textto-speech (TTS) models.
1904
Earlier work this paper cites.
L. R. Rabiner, “A tutorial on hidden Markov models and selected applications in speech recognition,” Proceedings of the IEEE , vol. 77, no. 2, pp. 257–286, Feb. 1989
1989
Earlier work this paper cites.
R. Kubichek, “Mel-cepstral distance measure for objective speech quality assessment,” in Proceedings of IEEE Pacific Rim Conference on Communications Computers and Signal Processing , vol. 1, 1993, pp. 125–128 vol.1
1993
Earlier work this paper cites.
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, “Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks,” ser. ICML ’06. New York, NY, USA: Association for Computing Machinery, 2006, p. 369–376
2006
Earlier work this paper cites.
2013
Earlier work this paper cites.
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
2017
Earlier work this paper cites.
2017
Cited alongside, same era.
M. McAuliffe, M. Socolof, S. Mihuc, M. Wagner, and M. Sonderegger, “Montreal forced aligner: Trainable text-speech alignment using kaldi,” in INTERSPEECH , 2017
2017
Cited alongside, same era.
2017
Cited alongside, same era.
K. Ito and L. Johnson, “The lj speech dataset,” https://keithito.com/LJ-Speech-Dataset/ , 2017
2017
Cited alongside, same era.
R. Valle, K. Shih, R. Prenger, and B. Catanzaro, “Flowtron: an autoregressive flow-based generative network for text-to-speech synthesis,” 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
J. Kim, S. Kim, J. Kong, and S. Yoon, “Glow-tts: A generative flow for text-to-speech via monotonic alignment search,” 2020
2020
Later among the works it cites.
A. Łańcucki, “Fastpitch: Parallel text-to-speech with pitch prediction,” 2020
2020
Later among the works it cites.
K. Peng, W. Ping, Z. Song, and K. Zhao, “Non-autoregressive neural text-to-speech,” in International Conference on Machine Learning . PMLR, 2020, pp. 7586–7598
2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
H. Tachibana, K. Uenoyama, and S. Aihara, “Efficiently trainable text-to-speech system based on deep convolutional networks with guided attention,” in 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2018, pp. 4784–4788
2018
Cited alongside, same era.
Y. Ren, Y. Ruan, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T.-Y. Liu, “Fastspeech: Fast, robust and controllable text to speech,” in Advances in Neural Information Processing Systems , H. Wallach, H. Larochelle, A. Beygelzimer, F. d'Alché-Buc, E. Fox, and R. Garnett, Eds., vol. 32. Curran Associates, Inc., 2019, pp. 3171–3180
2019
Cited alongside, same era.
2019
Cited alongside, same era.
Later among the works it cites.
E. Battenberg, R. J. Skerry-Ryan, S. Mariooryad, D. Stanton, D. Kao, M. Shannon, and T. Bagby, “Location-relative attention mechanisms for robust long-form speech synthesis,” in ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2020, pp. 6194–6198
2020
Later among the works it cites.
K. J. Shih, R. Valle, R. Badlani, A. Lancucki, W. Ping, and B. Catanzaro, “RAD-TTS: Parallel flow-based TTS with robust alignment learning and diverse synthesis,” in ICML Workshop on Invertible Neural Networks, Normalizing Flows, and Explicit Likelihood Models , 2021. [Online]. Available: https://openreview.net/forum?id=0NQwnnwAORi
2021
Closest in time.