Fetching the paper…
Reading the bibliography…
An unsupervised text-to-speech synthesis (TTS) system learns to generate speech waveforms corresponding to any written sentence in a language by observing: 1) a collection of untranscribed speech waveforms in that language; 2) a collection of texts written in that language without access to any transcribed speech.
A. Graves, S. Fernández, F. J. Gomez, and J. Schmidhuber, “Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,” in ICML , 2006, pp. 369–376
2006
Earlier work this paper cites.
D. Povey, A. Ghoshal, G. Boulianne, L. Burget, O. Glembek, N. Goel, M. Hannemann, P. Motlicek, Y. Qian, P. Schwarz, J. Silovsky, G. Stemmer, and K. Vesely, “The kaldi speech recognition toolkit,” in ASRU , 2011
2011
Earlier work this paper cites.
P. K. Muthukumar and A. W. Black, “Automatic discovery of a phonetic inventory for unwritten languages for statistical speech synthesis,” in ICASSP , 2014, pp. 2594–2598
2014
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: An ASR corpus based on public domain audio books,” in ICASSP , 2015, pp. 5206–5210
2015
Earlier work this paper cites.
2017
Earlier work this paper cites.
M. Hasegawa-Johnson, A. Black, L. Ondel, O. Scharenborg, and F. Ciannella, “Image2speech: Automatically generating audio descriptions of images,” in ICNLSSP , 2017, p. 1–5
2017
Earlier work this paper cites.
K. Ito and L. Johnson, “The lj speech dataset,” https://keithito.com/LJ-Speech-Dataset/ , 2017
2017
Earlier work this paper cites.
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Skerry-Ryan, R. A. Saurous, Y. Agiomyrgiannakis, and Y. Wu1, “Natural tts synthesis by conditioning wavenet on mel spectrogram predictions,” in ICASSP , 2018
2018
Earlier work this paper cites.
W. Ping, K. Peng, A. Gibiansky, S. O. Arik, A. Kannan, S. Narang, J. Raiman, and J. Miller, “Deep voice 3: 2000-speaker neural text-to-speech,” in ICLR , 2018
2018
Earlier work this paper cites.
H. Tachibana, K. Uenoyama, and S. Aihara, “Efficiently trainable text-to-speech system based on deep convolutional networks with guided attention,” in ICASSP , 2018, pp. 4784–4788
2018
Earlier work this paper cites.
Y. Ren, Y. Ruan, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T.-Y. Liu, “Fastspeech: Fast, robust and controllable text to speech,” in Advances in Neural Information Processing Systems , 2019
2019
Earlier work this paper cites.
N.Li, S.Liu, Y.Liu, S.Zhao, and M.Liu, “Neural speech synthesis with transformer network,” in AAAI , vol. 33, 2019, p. 6706–6713
2019
Earlier work this paper cites.
K. Park and T. Mulc, “CSS10: A collection of single speaker speech datasets for 10 languages,” Interspeech , 2019
2019
Cited alongside, same era.
Y. Ren, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T.-Y. Liu, “Almost unsupervised text to speech and automatic speech recognition,” in ICML , 2019, pp. 5410–5419
2019
Cited alongside, same era.
C.-K. Yeh, J. Chen, C. Yu, and D. Yu, “Unsupervised speech recognition via segmental empirical output distribution matching,” in ICLR , 2019
2019
Cited alongside, same era.
K.-Y. Chen, C.-P. Tsai, D.-R. Liu, H.-Y. Lee, and L. shan Lee, “Completely unsupervised speech recognition by a generative adversarial network harmonized with iteratively refined hidden Markov models,” in Interspeech , 2019
2019
Cited alongside, same era.
K. Park and J. Kim, “g2pe,” https://github.com/Kyubyong/g2p , 2019
2019
J. Kahn, M. Rivière, W. Zheng, E. Kharitonov, Q. Xu, P. E. Mazaré, J. Karadayi, V. Liptchinsky, R. Collobert, C. Fuegen, T. Likhomanenko, G. Synnaeve, A. Joulin, A. Mohamed, and E. Dupoux, “Libri-light: A benchmark for asr with limited or no supervision,” in ICASSP , 2020, pp. 7669–7673
2020
Later among the works it cites.
T. Hayashi, R. Yamamoto, K. Inoue, T. Yoshimura, S. Watanabe, T. Toda, K. Takeda, Y. Zhang, and X. Tan, “ESPnet-TTS: Unified, reproducible, and integratable open source end-to-end text-to-speech toolkit,” in ICASSP , 2020, pp. 7654–7658
2020
Later among the works it cites.
J. Kong, J. Kim, and J. Bae, “HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis,” in Neural Information Processing Systems , 2020
2020
Later among the works it cites.
2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
M. Ott, S. Edunov, A. Baevski, A. Fan, S. Gross, N. Ng, D. Grangier, and M. Auli, “fairseq: A fast, extensible toolkit for sequence modeling,” in Proceedings of NAACL-HLT 2019: Demonstrations , 2019
2019
Cited alongside, same era.
K. Park and T. Mulc, “Css10: A collection of single speaker speech datasets for 10 languages,” in Interspeech , 2019
2019
Cited alongside, same era.
J. Xu, X. Tan, Y. Ren, T. Qin, J. Li, S. Zhao, and T. Liu, “LRSpeech: Extremely low-resource speech synthesis and recognition,” in KDD , 2020, pp. 2802–2812
2020
Cited alongside, same era.
A. H. Liu, T. Tu, H. Lee, and L. Lee, “Towards unsupervised speech recognition and synthesis with quantized speech representation learning,” in ICASSP , 2020, pp. 7259–7263
2020
Cited alongside, same era.
H. Zhang and Y. Lin, “Unsupervised learning for sequence-to-sequence text-to-speech for low-resource languages,” in Interspeech , 2020, pp. 3161–3165
2020
Cited alongside, same era.
A. Baevski, H. Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learning of speech representations,” in Neural Information Processing Systems , 2020
2020
Cited alongside, same era.
Later among the works it cites.
A. Baevski, W.-N. Hsu, A. Conneau, and M. Auli, “Unsupervised speech recognition,” in Neural Information Processing Systems , 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
X. Wang, S. Feng, J. Zhu, M. Hasegawa-Johnson, and O. Scharenborg, “Show and speak: directly synthesize spoken description of images,” in icassp , 2021
2021
Later among the works it cites.
W.-N. Hsu, D. Harwath, T. Miller, C. Song, and J. Glass, “Text-free image-to-speech synthesis using learned segmental units,” in ACL-IJCNLP , 2021, pp. 5284–5300
2021
Later among the works it cites.
J. Effendi, S. Sakti, and S. Nakamura, “End-to-end image-to-speech generation for untranscribed unknown languages,” IEEE Access , vol. 9, pp. 55 144–55 154, 2021
2021
Later among the works it cites.