Fetching the paper…
Reading the bibliography…
We explore cross-lingual multi-speaker speech synthesis and cross-lingual voice conversion applied to data augmentation for automatic speech recognition (ASR) systems in low/medium-resource scenarios.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: an asr corpus based on public domain audio books,” in Acoustics, Speech and Signal Processing (ICASSP), 2015 IEEE International Conference on . IEEE, 2015, pp. 5206–5210
2015
Earlier work this paper cites.
C. Veaux, J. Yamagishi, K. MacDonald et al. , “Superseded-cstr vctk corpus: English multi-speaker corpus for cstr voice cloning toolkit,” University of Edinburgh. The Centre for Speech Technology Research (CSTR) , 2016
2016
Earlier work this paper cites.
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Skerrv-Ryan et al. , “Natural tts synthesis by conditioning wavenet on mel spectrogram predictions,” in 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2018, pp. 4779–4783
2018
Earlier work this paper cites.
Y. Jia, Y. Zhang, R. Weiss, Q. Wang, J. Shen, F. Ren, P. Nguyen, R. Pang, I. L. Moreno, Y. Wu et al. , “Transfer learning from speaker verification to multispeaker text-to-speech synthesis,” in Advances in neural information processing systems , 2018, pp. 4480–4490
2018
Earlier work this paper cites.
A. Tjandra, S. Sakti, and S. Nakamura, “Machine speech chain with one-shot speaker adaptation,” Proc. Interspeech 2018 , pp. 887–891, 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
A. Rosenberg, Y. Zhang, B. Ramabhadran, Y. Jia, P. Moreno, Y. Wu, and Z. Wu, “Speech recognition with augmented synthesized speech,” in 2019 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) . IEEE, 2019, pp. 996–1002
2019
Earlier work this paper cites.
I. Solak, “The m-ailabs speech dataset,” 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
K. Park and T. Mulc, “Css10: A collection of single speaker speech datasets for 10 languages,” Proc. Interspeech 2019 , pp. 1566–1570, 2019
2019
Earlier work this paper cites.
R. Valle, K. J. Shih, R. Prenger, and B. Catanzaro, “Flowtron: an autoregressive flow-based generative network for text-to-speech synthesis,” in International Conference on Learning Representations , 2020
2020
Earlier work this paper cites.
2020
Cited alongside, same era.
2020
Cited alongside, same era.
——, “Machine speech chain,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 28, pp. 976–989, 2020
2020
Cited alongside, same era.
A. Laptev, R. Korostik, A. Svischev, A. Andrusenko, I. Medennikov, and S. Rybin, “You do not need more data: improving end-to-end speech recognition by text-to-speech data augmentation,” in 2020 13th International Congress on Image and Signal Processing, BioMedical Engineering and Informatics (CISP-BMEI) . IEEE, 2020, pp. 439–444
2020
Cited alongside, same era.
J. Kim, J. Kong, and J. Son, “Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,” in International Conference on Machine Learning . PMLR, 2021, pp. 5530–5540
2021
Later among the works it cites.
S. Elizabeth, W. Matthew, B. Jacob, R. Cattoni, M. Negri, M. Turchi, D. W. Oard, and P. Matt, “The multilingual tedx corpus for speech recognition and translation,” in Proceedings of Interspeech 2021 , 2021, pp. 3655–3659
2021
Later among the works it cites.
X. Hao, X. Su, R. Horaud, and X. Li, “Fullsubnet: A full-band and sub-band fusion model for real-time single-channel speech enhancement,” ICASSP 2021 - 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , Jun 2021. [Online]. Available: http://dx.doi.org/10.1109/ICASSP39728.2021.9414177
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
V. Pratap, Q. Xu, A. Sriram, G. Synnaeve, and R. Collobert, “Mls: A large-scale multilingual dataset for speech research,” Proc. Interspeech 2020 , pp. 2757–2761, 2020
2020
Cited alongside, same era.
R. Ardila, M. Branson, K. Davis, M. Kohler, J. Meyer, M. Henretty, R. Morais, L. Saunders, F. Tyers, and G. Weber, “Common voice: A massively-multilingual speech corpus,” in Proceedings of the 12th Language Resources and Evaluation Conference , 2020, pp. 4218–4222
2020
Cited alongside, same era.
2020
Cited alongside, same era.
A. Baevski, Y. Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learning of speech representations,” Advances in Neural Information Processing Systems , vol. 33, 2020
2020
Cited alongside, same era.
I. M. Quintanilha, S. L. Netto, and L. W. P. Biscainho, “An open-source end-to-end asr system for brazilian portuguese using dnns built from newly assembled corpora,” Journal of Communication and Information Systems , vol. 35, no. 1, pp. 230–242, 2020
2020
Cited alongside, same era.
2020
Cited alongside, same era.
S. Kriman, S. Beliaev, B. Ginsburg, J. Huang, O. Kuchaiev, V. Lavrukhin, R. Leary, J. Li, and Y. Zhang, “Quartznet: Deep automatic speech recognition with 1d time-channel separable convolutions,” in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 6124–6128
2020
Cited alongside, same era.
2021
Later among the works it cites.
L. R. Stefanel Gris, E. Casanova, F. Santos de Oliveira, A. da Silva Soares, and A. C. Junior, “Brazilian portuguese speech recognition using wav2vec 2.0,” arXiv e-prints , pp. arXiv–2107, 2021
2021
Later among the works it cites.
E. Casanova, J. Weber, C. D. Shulby, A. C. Junior, E. Gölge, and M. A. Ponti, “Yourtts: Towards zero-shot multi-speaker tts and zero-shot voice conversion for everyone,” in International Conference on Machine Learning . PMLR, 2022, pp. 2709–2720
2022
Closest in time.
M. Baas and H. Kamper, “Voice Conversion Can Improve ASR in Very Low-Resource Settings,” in Proc. Interspeech 2022 , 2022, pp. 3513–3517
2022
Closest in time.
E. Casanova, A. C. Junior, C. Shulby, F. S. d. Oliveira, J. P. Teixeira, M. A. Ponti, and S. Aluísio, “Tts-portuguese corpus: a corpus for speech synthesis in brazilian portuguese,” Language Resources and Evaluation , pp. 1–13, 2022
2022
Closest in time.
Y. Zhang, D. S. Park, W. Han, J. Qin, A. Gulati, J. Shor, A. Jansen, Y. Xu, Y. Huang, S. Wang et al. , “Bigssl: Exploring the frontier of large-scale semi-supervised learning for automatic speech recognition,” IEEE Journal of Selected Topics in Signal Processing , 2022
2022
Closest in time.
2022
Closest in time.
S. Chen, C. Wang, Z. Chen, Y. Wu, S. Liu, Z. Chen, J. Li, N. Kanda, T. Yoshioka, X. Xiao et al. , “Wavlm: Large-scale self-supervised pre-training for full stack speech processing,” IEEE Journal of Selected Topics in Signal Processing , vol. 16, no. 6, pp. 1505–1518, 2022
2022
Closest in time.