Fetching the paper…
Reading the bibliography…
This paper introduces a new speech dataset called ``LibriTTS-R'' designed for text-to-speech (TTS) use.
J. B. Allen and D. A. Berkley, “Image method for efficiently simulating small-room acoustics,” J. Acoust. Soc. Am. , 1979
1979
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: An ASR corpus based on public domain audio books,” in Proc. ICASSP , 2015
2015
Earlier work this paper cites.
D. P. Kingma and J. L. Ba, “Adam: A method for stochastic optimization,” in Proc. ICLR , 2015
2015
Earlier work this paper cites.
2016
Earlier work this paper cites.
K. Ito and L. Johnson, “The lj speech dataset,” https://keithito.com/LJ-Speech-Dataset/ , 2017
2017
Earlier work this paper cites.
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Skerrv-Ryan, R. A. Saurous, Y. Agiomvrgiannakis, and Y. Wu, “Natural TTS synthesis by conditioning WaveNet on mel spectrogram predictions,” in Proc. ICASSP , 2018
2018
Earlier work this paper cites.
Y. Jia, Y. Zhang, R. Weiss, Q. Wang, J. Shen, F. Ren, z. Chen, P. Nguyen, R. Pang, I. Lopez Moreno, and Y. Wu, “Transfer learning from speaker verification to multispeaker text-to-speech synthesis,” in Proc. NeurIPS , 2018
2018
Earlier work this paper cites.
N. Kalchbrenner, W. Elsen, K. Simonyan, S. Noury, N. Casagrande, W. Lockhart, F. Stimberg, A. van den Oord, S. Dieleman, and K. Kavukcuoglu, “Efficient neural audio synthesis,” in Proc. ICML , 2018
2018
Earlier work this paper cites.
Y. Ren, Y. Ruan, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T.-Y. Liu, “FastSpeech: Fast, robust and controllable text to speech,” in Proc. NeurIPS , 2019
2019
Earlier work this paper cites.
S. Maiti and M. I. Mandel, “Parametric resynthesis with neural vocoders,” in Proc. IEEE WASPAA , 2019
2019
Earlier work this paper cites.
H. Zen, R. Clark, R. J. Weiss, V. Dang, Y. Jia, Y. Wu, Y. Zhang, and Z. Chen, “LibriTTS: A corpus derived from LibriSpeech for text-to-speech,” in Proc. Interspeech , 2019
2019
Earlier work this paper cites.
Y. Chen, Y. Assael, B. Shillingford, D. Budden, S. Reed, H. Zen, Q. Wang, L. C. Cobo, A. Trask, B. Laurie, C. Gulcehre, A. van den Oord, O. Vinyals, and N. de Freitas, “Sample efficient adaptive text-to-speech,” in Proc. ICLR , 2019
2019
Earlier work this paper cites.
J. Kong, J. Kim, and J. Bae, “HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis,” in Proc. NeurIPS , 2020
2020
Earlier work this paper cites.
T. Hayashi, R. Yamamoto, K. Inoue, T. Yoshimura, S. Watanabe, T. Toda, K. Takeda, Y. Zhang, and X. Tan, “ESPnet-TTS: Unified, reproducible, and integratable open source end-to-end text-to-speech toolkit,” in Proc. ICASSP , 2020
2020
Cited alongside, same era.
2020
Cited alongside, same era.
——, “Speaker independence of neural vocoders and their effect on parametric resynthesis speech enhancement,” in Proc. ICASSP , 2020
2020
Cited alongside, same era.
J. Su, Z. Jin, and A. Finkelstein, “HiFi-GAN: High-fidelity denoising and dereverberation based on speech deep features in adversarial networks,” in Proc. Interspeech , 2020
2020
Cited alongside, same era.
R. Valle, K. J. Shih, R. Prenger, and B. Catanzaro, “Flowtron: an autoregressive flow-based generative network for text-to-speech synthesis,” in Proc. Int. Conf. Learn. Represent. (ICLR) , 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
Y. Koizumi, S. Karita, S. Wisdom, H. Erdogan, J. R. Hershey, L. Jones, and M. Bacchiani, “DF-Conformer: Integrated architecture of Conv-TasNet and Conformer using linear complexity self-attention for speech enhancement,” in Proc. IEEE WASPAA , 2021
2021
Later among the works it cites.
S. Wang, A. Mesaros, T. Heittola, and T. Virtanen, “A curated dataset of urban scenes for audio-visual scene analysis,” in Proc. ICASSP , 2021
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Kim, S. Kim, J. Kong, and S. Yoon, “Glow-tts: A generative flow for text-to-speech via monotonic alignment search,” in Proc. Adv. Neural Inf. Process. Syst. (NeurIPS) , 2020
2020
Cited alongside, same era.
Y. Zhang, J. Qin, D. S. Park, W. Han, C.-C. Chiu, R. Pang, Q. V. Le, and Y. Wu, “Pushing the limits of semi-supervised learning for automatic speech recognition,” in Proc. NeurIPS SAS 2020 Workshop , 2020
2020
Cited alongside, same era.
Y. Jia, H. Zen, J. Shen, Y. Zhang, and Y. Wu, “PnG BERT: Augmented BERT on phonemes and graphemes for neural TTS,” in Proc. Interspeech , 2021
2021
Cited alongside, same era.
I. Elias, H. Zen, J. Shen, Y. Zhang, Y. Jia, R. J. Weiss, and Y. Wu, “Parallel Tacotron: Non-autoregressive and controllable TTS,” in Proc. ICASSP , 2021
2021
Cited alongside, same era.
Y. Ren, C. Hu, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T.-Y. Liu, “FastSpeech 2: Fast and high-quality end-to-end text to speech,” in Proc. Int. Conf. Learn. Represent. (ICLR) , 2021
2021
Cited alongside, same era.
2021
Cited alongside, same era.
——, “HiFi-GAN-2: Studio-quality speech enhancement via generative adversarial networks conditioned on acoustic features,” in Proc. IEEE WASPAA , 2021
2021
Cited alongside, same era.
Y.-A. Chung, Y. Zhang, W. Han, C.-C. Chiu, J. Qin, R. Pang, and Y. Wu, “w2v-BERT: Combining contrastive learning and masked language modeling for self-supervised speech pre-training,” in Proc. IEEE ASRU , 2021
2021
Cited alongside, same era.
T. Saeki, S. Takamichi, T. Nakamura, N. Tanji, and H. Saruwatari, “SelfRemaster: Self-supervised speech restoration with analysis-by-synthesis approach using channel modeling,” in Proc Interspeech , 2022
2022
Later among the works it cites.
H. Liu, X. Liu, Q. Kong, Q. Tian, Y. Zhao, D. Wang, C. Huang, and Y. Wang, “VoiceFixer: A unified framework for high-fidelity speech restoration,” in Proc Interspeech , 2022
2022
Later among the works it cites.
2022
Later among the works it cites.
Z. Borsos, M. Sharifi, and M. Tagliasacchi, “SpeechPainter: Text-conditioned speech inpainting,” in Proc. Interspeech , 2022
2022
Later among the works it cites.
Y. Koizumi, K. Yatabe, H. Zen, and M. Bacchiani, “WaveFit: An iterative and non-autoregressive neural vocoder based on fixed-point iteration,” in Proc. IEEE SLT , 2023
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.