Fetching the paper…
Reading the bibliography…
While modern Text-to-Speech (TTS) systems can produce natural-sounding speech, they remain unable to reproduce the full diversity found in natural speech data.
L. N. Vaserstein, “Markov processes over denumerable products of spaces, describing large systems of automata,” Problemy Peredachi Informatsii , 1969
1969
Earlier work this paper cites.
K. N. Stevens, Sources of Inter- and Intra-Speaker Variability in the Acoustic Properties of Speech Sounds , 1972
1972
Earlier work this paper cites.
D. Dowson and B. Landau, “The fréchet distance between multivariate normal distributions,” Journal of multivariate analysis , 1982
1982
Earlier work this paper cites.
K.-F. Lee, “On large-vocabulary speaker-independent continuous speech recognition,” Speech communication , 1988
1988
Earlier work this paper cites.
J. Makhoul and R. Schwartz, “State of the art in continuous speech recognition,” PNAS , 1995
1995
Earlier work this paper cites.
C. Kim and R. M. Stern, “Robust signal-to-noise ratio estimation based on waveform amplitude distribution analysis,” in ISCA , 2008
2008
Earlier work this paper cites.
T. H. Falk, C. Zheng, and W.-Y. Chan, “A non-intrusive quality and intelligibility measure of reverberant and dereverberated speech,” TASLP , 2010
2010
Earlier work this paper cites.
2014
Earlier work this paper cites.
V. Peddinti, D. Povey, and S. Khudanpur, “A time delay neural network architecture for efficient modeling of long temporal contexts,” in ISCA , 2015
2015
Earlier work this paper cites.
D. Povey, V. Peddinti, D. Galvez, P. Ghahremani, V. Manohar, X. Na, Y. Wang, and S. Khudanpur, “Purely sequence-trained neural networks for ASR based on lattice-free MMI,” in Interspeech , 2016
2016
Earlier work this paper cites.
M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter, “Gans trained by a two time-scale update rule converge to a local nash equilibrium,” Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
A. Nagrani, J. S. Chung, and A. Zisserman, “Voxceleb: a large-scale speaker identification dataset,” in Interspeech , 2017
2017
Cited alongside, same era.
M. Cuturi and M. Blondel, “Soft-dtw: a differentiable loss function for time-series,” in ICML , 2017
2017
Cited alongside, same era.
L. Wan, Q. Wang, A. Papir, and I. L. Moreno, “Generalized end-to-end loss for speaker verification,” in ICASSP , 2018
2018
Cited alongside, same era.
W.-N. Hsu, Y. Zhang, R. J. Weiss, H. Zen, Y. Wu, Y. Wang, Y. Cao, Y. Jia, Z. Chen, J. Shen et al. , “Hierarchical generative modeling for controllable speech synthesis,” in ICLR , 2018
2018
Cited alongside, same era.
2018
Z. Chen, A. Rosenberg, Y. Zhang, H. Zen, M. Ghodsi, Y. Huang, J. Emond, G. Wang, B. Ramabhadran, and P. J. M. Mengibar, “Semi-supervision in ASR: Sequential mixmatch and factorized TTS-based augmentation,” 2021
2021
Later among the works it cites.
R. Luo, X. Tan, R. Wang, T. Qin, J. Li, S. Zhao, E. Chen, and T.-Y. Liu, “Lightspeech: Lightweight and fast text to speech with neural architecture search,” in ICASSP , 2021
2021
Later among the works it cites.
C.-M. Chien, J.-H. Lin, C.-y. Huang, P.-c. Hsu, and H.-y. Lee, “Investigating on incorporating pretrained and learnable speaker representations for multi-speaker multi-style text-to-speech,” in ICASSP , 2021
2021
Later among the works it cites.
H. Liu, Q. Kong, Q. Tian, Y. Zhao, D. Wang, C. Huang, and Y. Wang, “Voicefixer: Toward general speech restoration with neural vocoder,” 2021
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
S. Kolouri, P. E. Pope, C. E. Martin, and G. K. Rohde, “Sliced wasserstein auto-encoders,” in ICLR , 2018
2018
Cited alongside, same era.
A. Rosenberg, Y. Zhang, B. Ramabhadran, Y. Jia, P. Moreno, Y. Wu, and Z. Wu, “Speech recognition with augmented synthesized speech,” in ASRU , 2019
2019
Cited alongside, same era.
M. Bińkowski, J. Donahue, S. Dieleman, A. Clark, E. Elsen, N. Casagrande, L. C. Cobo, and K. Simonyan, “High fidelity speech synthesis with adversarial networks,” in ICLR , 2019
2019
Cited alongside, same era.
Y. Ren, C. Hu, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T.-Y. Liu, “Fastspeech 2: Fast and high-quality end-to-end text to speech,” in ICLR , 2020
2020
Cited alongside, same era.
J. Kim, J. Kong, and J. Son, “Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,” in ICML , 2021
2021
Cited alongside, same era.
A. Łańcucki, “Fastpitch: Parallel text-to-speech with pitch prediction,” in ICASSP , 2021
2021
Cited alongside, same era.
J. M. Eargle, Music, sound, and technology . Springer
Cited in the paper.
2022
Closest in time.
E. Casanova, J. Weber, C. D. Shulby, A. C. Junior, E. Gölge, and M. A. Ponti, “YourTTS: Towards zero-shot multi-speaker tts and zero-shot voice conversion for everyone,” in ICML , 2022
2022
Closest in time.
2022
Closest in time.
A. Wali, Z. Alamgir, S. Karim, A. Fawaz, M. B. Ali, M. Adan, and M. Mujtaba, “Generative adversarial networks for speech processing: A review,” Computer Speech & Language , 2022
2022
Closest in time.
D. Stanton, M. Shannon, S. Mariooryad, R. Skerry-Ryan, E. Battenberg, T. Bagby, and D. Kao, “Speaker generation,” in ICASSP , 2022
2022
Closest in time.
T.-Y. Hu, M. Armandpour, A. Shrivastava, J.-H. R. Chang, H. Koppula, and O. Tuzel, “Synt++: Utilizing imperfect synthetic data to improve speech recognition,” in ICASSP , 2022
2022
Closest in time.