Fetching the paper…
Reading the bibliography…
Various applications of voice synthesis have been developed independently despite the fact that they generate "voice" as output in common.
Yet another algorithm for pitch tracking:(yaapt)
Kasi, K · 2002
Earlier work this paper cites.
A review of vector quantization techniques
Vasuki, A. and Vanathi, P · 2006
Earlier work this paper cites.
Librispeech: an asr corpus based on public domain audio books
Panayotov, V., Chen, G., Povey, D., and Khudanpur, S · 2015
Earlier work this paper cites.
Musan: A music, speech, and noise corpus
Snyder, D., Chen, G., and Povey, D · 2015
Earlier work this paper cites.
Tacotron: Towards end-to-end speech synthesis
Wang, Y., Skerry-Ryan, R., Stanton, D., Wu, Y., Weiss, R. J., Jaitly, N., Yang, Z., Xiao, Y., Chen, Z., Bengio, S., et al · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Earlier work this paper cites.
fairseq: A fast, extensible toolkit for sequence modeling
Ott, M., Edunov, S., Baevski, A., Fan, A., Gross, S., Ng, N., Grangier, D., and Auli, M · 2019
Earlier work this paper cites.
Fastspeech: Fast, robust and controllable text to speech
Ren, Y., Ruan, Y., Tan, X., Qin, T., Zhao, S., Zhao, Z., and Liu, T.-Y · 2019
Earlier work this paper cites.
Token-level ensemble distillation for grapheme-to-phoneme conversion
Sun, H., Tan, X., Gan, J.-W., Liu, H., Zhao, S., Qin, T., and Liu, T.-Y · 2019
Earlier work this paper cites.
wav2vec 2.0: A framework for self-supervised learning of speech representations
Baevski, A., Zhou, Y., Mohamed, A., and Auli, M · 2020
Earlier work this paper cites.
End-to-end adversarial text-to-speech
Donahue, J., Dieleman, S., Bińkowski, M., Elsen, E., and Simonyan, K · 2020
Earlier work this paper cites.
Libri-light: A benchmark for asr with limited or no supervision
Kahn, J., Riviere, M., Zheng, W., Kharitonov, E., Xu, Q., Mazaré, P.-E., Karadayi, J., Liptchinsky, V., Collobert, R., Fuegen, C., et al · 2020
Earlier work this paper cites.
Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis
Kong, J., Kim, J., and Bae, J · 2020
Earlier work this paper cites.
Unsupervised speech decomposition via triple information bottleneck
Qian, K., Zhang, Y., Chang, S., Hasegawa-Johnson, M., and Cox, D · 2020
Earlier work this paper cites.
Neural analysis and synthesis: Reconstructing speech from self-supervised representations
Choi, H.-S., Lee, J., Kim, W., Lee, J., Heo, H., and Lee, K · 2021
Earlier work this paper cites.
Hubert: Self-supervised speech representation learning by masked prediction of hidden units
Hsu, W.-N., Bolte, B., Tsai, Y.-H. H., Lakhotia, K., Salakhutdinov, R., and Mohamed, A · 2021
Cited alongside, same era.
Multi-singer: Fast multi-singer singing voice vocoder with a large-scale corpus
Huang, R., Chen, F., Ren, Y., Liu, J., Cui, C., and Zhao, Z · 2021
Cited alongside, same era.
Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech
Kim, J., Kong, J., and Son, J · 2021
Cited alongside, same era.
Textless speech-to-speech translation on real data
Lee, A., Gong, H., Duquenne, P.-A., Schwenk, H., Chen, P.-J., Wang, C., Popuri, S., Pino, J., Gu, J., and Hsu, W.-N · 2021
Cited alongside, same era.
Any-to-many voice conversion with location-relative sequence-to-sequence modeling
Liu, S., Cao, Y., Wang, D., Wu, X., Liu, X., and Meng, H · 2021
Cited alongside, same era.
High fidelity neural audio compression
Défossez, A., Copet, J., Synnaeve, G., and Adi, Y · 2022
Later among the works it cites.
textless-lib: A library for textless spoken language processing
Kharitonov, E., Copet, J., Lakhotia, K., Nguyen, T. A., Tomasello, P., Lee, A., Elkahky, A., Hsu, W.-N., Mohamed, A., Dupoux, E., et al · 2022
Later among the works it cites.
Bigvgan: A universal neural vocoder with large-scale training
Lee, S.-g., Ping, W., Ginsburg, B., Catanzaro, B., and Yoon, S · 2022
Later among the works it cites.
Utts: Unsupervised tts with conditional disentangled sequential variational auto-encoder
Lian, J., Zhang, C., Anumanchipalli, G. K., and Yu, D · 2022
Later among the works it cites.
Diffsinger: Singing voice synthesis via shallow diffusion mechanism
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Meta-stylespeech: Multi-speaker adaptive text-to-speech generation
Min, D., Lee, D. B., Yang, E., and Hwang, S. J · 2021
Cited alongside, same era.
Speech resynthesis from discrete disentangled self-supervised representations
Polyak, A., Adi, Y., Copet, J., Kharitonov, E., Lakhotia, K., Hsu, W.-N., Mohamed, A., and Dupoux, E · 2021
Cited alongside, same era.
Shah, J., Singla, Y. K., Chen, C., and Shah, R. R · 2021
Cited alongside, same era.
Soundstream: An end-to-end neural audio codec
Zeghidour, N., Luebs, A., Omran, A., Skoglund, J., and Tagliasacchi, M · 2021
Cited alongside, same era.
Seen and unseen emotional style transfer for voice conversion with a new emotional speech dataset
Zhou, K., Sisman, B., Liu, R., and Li, H · 2021
Cited alongside, same era.
Data2vec: A general framework for self-supervised learning in speech, vision and language
Baevski, A., Hsu, W.-N., Xu, Q., Babu, A., Gu, J., and Auli, M · 2022
Cited alongside, same era.
Audiolm: a language modeling approach to audio generation
Borsos, Z., Marinier, R., Vincent, D., Kharitonov, E., Pietquin, O., Sharifi, M., Teboul, O., Grangier, D., Tagliasacchi, M., and Zeghidour, N · 2022
Cited alongside, same era.
Liu, J., Li, C., Ren, Y., Chen, F., and Zhao, Z · 2022
Later among the works it cites.
Opencpop: A high-quality open source chinese popular song corpus for singing voice synthesis
Wang, Y., Wang, X., Zhu, P., Wu, J., Li, H., Xue, H., Zhang, Y., Xie, L., and Bi, M · 2022
Later among the works it cites.
M4singer: A multi-style, multi-singer and musical score provided mandarin singing corpus
Zhang, L., Li, R., Wang, S., Deng, L., Liu, J., Ren, Y., He, J., Huang, R., Zhu, J., Chen, X., et al · 2022
Later among the works it cites.
Musiclm: Generating music from text
Agostinelli, A., Denk, T. I., Borsos, Z., Engel, J., Verzetti, M., Caillon, A., Huang, Q., Jansen, A., Roberts, A., Tagliasacchi, M., et al · 2023
Closest in time.
Speak, read and prompt: High-fidelity text-to-speech with minimal supervision
Kharitonov, E., Vincent, D., Borsos, Z., Marinier, R., Girgin, S., Pietquin, O., Sharifi, M., Tagliasacchi, M., and Zeghidour, N · 2023
Closest in time.
Unifyspeech: A unified framework for zero-shot text-to-speech and voice conversion
Liu, H., Wang, T., Fu, R., Yi, J., Wen, Z., and Tao, J · 2023
Closest in time.
Generative spoken dialogue language modeling
Nguyen, T. A., Kharitonov, E., Copet, J., Adi, Y., Hsu, W.-N., Elkahky, A., Tomasello, P., Algayres, R., Sagot, B., Mohamed, A., et al · 2023
Closest in time.
Naturalspeech 2: Latent diffusion models are natural and zero-shot speech and singing synthesizers
Shen, K., Ju, Z., Tan, X., Liu, Y., Leng, Y., He, L., Qin, T., Zhao, S., and Bian, J · 2023
Closest in time.
Neural codec language models are zero-shot text to speech synthesizers
Wang, C., Chen, S., Wu, Y., Zhang, Z., Zhou, L., Liu, S., Chen, Z., Liu, Y., Wang, H., Li, J., et al · 2023
Closest in time.
Speak foreign languages with your own voice: Cross-lingual neural codec language modeling
Zhang, Z., Zhou, L., Wang, C., Chen, S., Wu, Y., Liu, S., Chen, Z., Liu, Y., Wang, H., Li, J., et al · 2023
Closest in time.