Fetching the paper…
Reading the bibliography…
Language model (LM) based audio generation frameworks, e.g., AudioLM, have recently achieved new state-of-the-art performance in zero-shot audio generation.
J. Kominek and A. W. Black, “The CMU Arctic speech databases,” in 5th ISCA Workshop on Speech Synthesis (SSW 5) , 2004, pp. 223–224
2004
Earlier work this paper cites.
M. Wester, “The EMIME bilingual database,” The University of Edinburgh, Tech. Rep., 2010
2010
Earlier work this paper cites.
M. Steinschneider, K. V. Nourski, and Y. I. Fishman, “Representation of speech in human auditory cortex: is it special?” Hearing research , vol. 305, pp. 57–73, 2013
2013
Earlier work this paper cites.
C. Gulcehre, O. Firat, K. Xu, K. Cho, L. Barrault, H.-C. Lin, F. Bougares, H. Schwenk, and Y. Bengio, “On using monolingual corpora in neural machine translation,” Arxiv , 2015
2015
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “LibriSpeech: an ASR corpus based on public domain audio books,” in International conference on acoustics, speech and signal processing (ICASSP) , 2015, pp. 5206–5210
2015
Earlier work this paper cites.
L. Sun, K. Li, H. Wang, S. Kang, and H. Meng, “Phonetic Posteriorgrams for Many-to-One Voice Conversion without Parallel Data Training,” in International Conference on Multimedia and Expo (ICME) , 2016, pp. 1–6
2016
Earlier work this paper cites.
C. Veaux, J. Yamagishi, and K. MacDonald, “CSTR VCTK corpus: English multi-speaker corpus for CSTR voice cloning toolkit.” University of Edinburgh. The Centre for Speech Technology Research (CSTR), 2016
2016
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. u. Kaiser, and I. Polosukhin, “Attention is all you need,” in Neural Information Processing Systems (NerurIPS) , vol. 30, 2017
2017
Earlier work this paper cites.
A. v. d. Oord, Y. Li, and O. Vinyals, “Representation learning with contrastive predictive coding,” Arxiv , 2018
2018
Earlier work this paper cites.
K. Qian, Y. Zhang, S. Chang, X. Yang, and M. Hasegawa-Johnson, “AutoVC: Zero-shot voice style transfer with only autoencoder loss,” in International Conference on Machine Learning (ICML) , 2019, pp. 5210–5219
2019
Earlier work this paper cites.
J. chieh Chou and H.-Y. Lee, “One-shot voice conversion by separating speaker and content representations with instance normalization,” in International Speech Communication Association (Interspeech) , 2019, pp. 664–668
2019
Earlier work this paper cites.
L. Dong, N. Yang, W. Wang, F. Wei, X. Liu, Y. Wang, J. Gao, M. Zhou, and H.-W. Hon, “Unified language model pre-training for natural language understanding and generation,” in Neural Information Processing Systems (NerurIPS) , vol. 32, 2019
2019
Earlier work this paper cites.
H. Zen, V. Dang, R. Clark, Y. Zhang, R. J. Weiss, Y. Jia, Z. Chen, and Y. Wu, “LibriTTS: A corpus derived from LibriSpeech for text-to-speech,” in International Speech Communication Association (Interspeech) , 2019, pp. 1526–1530
2019
Earlier work this paper cites.
K. Qian, Y. Zhang, S. Chang, M. Hasegawa-Johnson, and D. Cox, “Unsupervised speech decomposition via triple information bottleneck,” in International Conference on Machine Learning (ICML) , 2020, pp. 7836–7846
2020
Earlier work this paper cites.
A. Baevski, Y. Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learning of speech representations,” in Neural Information Processing Systems (NerurIPS) , vol. 33, 2020, pp. 12 449–12 460
2020
Cited alongside, same era.
B. Desplanques, J. Thienpondt, and K. Demuynck, “ECAPA-TDNN: Emphasized channel attention, propagation and aggregation in TDNN based speaker verification,” in International Speech Communication Association (Interspeech) , 2020, pp. 3830–3834
2020
Cited alongside, same era.
Y. Gu, Z. Zhang, X. Yi, and X. Zhao, “MediumVC: Any-to-any voice conversion using synthetic specific-speaker speeches as intermedium features,” Arxiv , 2021
2021
Cited alongside, same era.
H.-S. Choi, J. Lee, W. Kim, J. Lee, H. Heo, and K. Lee, “Neural analysis and synthesis: Reconstructing speech from self-supervised representations,” in Neural Information Processing Systems(NeurIPS) , 2021, pp. 16 251–16 265
2021
Cited alongside, same era.
2021
Later among the works it cites.
G. Maimon and Y. Adi, “Speaking style conversion with discrete self-supervised units,” Arxiv , 2022
2022
Later among the works it cites.
H. Tang, X. Zhang, J. Wang, N. Cheng, and J. Xiao, “AVQVC: One-shot voice conversion by vector quantization with applying contrastive learning,” in Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2022, pp. 4613–4617
2022
Later among the works it cites.
Z. Borsos, R. Marinier, D. Vincent, E. Kharitonov, O. Pietquin, M. Sharifi, O. Teboul, D. Grangier, M. Tagliasacchi, and N. Zeghidour, “AudioLM: a language modeling approach to audio generation,” Arxiv , 2022
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Y. Lin, C. M. Chien, J. hao Lin, H. yi Lee, and L.-S. Lee, “FragmentVC: Any-to-any voice conversion by end-to-end extracting and fusing fine-grained voice fragments with attention,” in International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2021, pp. 5939–5943
2021
Cited alongside, same era.
D. Yin, X. Ren, C. Luo, Y. Wang, Z. Xiong, and W. Zeng, “Retriever: Learning content-style representation as a token-level bipartite graph,” in International Conference on Learning Representations (ICLR) , 2021
2021
Cited alongside, same era.
D. Wang, L. Deng, Y. T. Yeung, X. Chen, X. Liu, and H. Meng, “VQMIVC: Vector quantization and mutual information-based unsupervised speech representation disentanglement for one-shot voice conversion,” in International Speech Communication Association (Interspeech) , 2021, pp. 1344–1348
2021
Cited alongside, same era.
J. Wang, J. Li, X. Zhao, Z. Wu, S. Kang, and H. Meng, “Adversarially learning disentangled speech representations for robust multi-factor voice conversion,” Arxiv , 2021
2021
Cited alongside, same era.
J. Ebbers, M. Kuhlmann, T. Cord-Landwehr, and R. Haeb-Umbach, “Contrastive predictive coding supported factorized variational autoencoder for unsupervised learning of disentangled speech representations,” in International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2021, pp. 3860–3864
2021
Cited alongside, same era.
W.-N. Hsu, B. Bolte, Y.-H. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed, “HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,” Transactions on Audio, Speech, and Language Processing , vol. 29, pp. 3451–3460, 2021
2021
Cited alongside, same era.
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, “Swin Transformer: Hierarchical vision transformer using shifted windows,” in International conference on computer vision (ICCV) , 2021, pp. 10 012–10 022
2021
Cited alongside, same era.
N. Zeghidour, A. Luebs, A. Omran, J. Skoglund, and M. Tagliasacchi, “SoundStream: An end-to-end neural audio codec,” Transactions on Audio, Speech, and Language Processing , vol. 30, pp. 495–507, 2021
2021
Cited alongside, same era.
E. Casanova, J. Weber, C. D. Shulby, A. C. Junior, E. Gölge, and M. A. Ponti, “YourTTS: Towards zero-shot multi-speaker TTS and zero-shot voice conversion for everyone,” in International Conference on Machine Learning (ICML) , 2022, pp. 2709–2720
2022
Later among the works it cites.
Z. Du, Y. Qian, X. Liu, M. Ding, J. Qiu, Z. Yang, and J. Tang, “GLM: General language model pretraining with autoregressive blank infilling,” in Association for Computational Linguistics (ACL) , 2022, pp. 320–335
2022
Later among the works it cites.
X. Shu, Y. Chen, C. Shang, Y. Zhao, C. Zhao, Y. Zhu, C. Huang, and Y. Wang, “Non-intrusive speech quality assessment with a multi-task learning based subband adaptive attention temporal convolutional neural network,” in International Speech Communication Association (Interspeech) , 2022, pp. 3298–3302
2022
Later among the works it cites.
Z. Wang, L. Xue, Q. Kong, L. Xie, Y. Chen, Q. Tian, and Y. Wang, “Multi-level temporal-channel speaker retrieval for robust zero-shot voice conversion,” Arxiv , 2023
2023
Closest in time.
A. Agostinelli, T. I. Denk, Z. Borsos, J. Engel, M. Verzetti, A. Caillon, Q. Huang, A. Jansen, A. Roberts, M. Tagliasacchi, M. Sharifi, N. Zeghidour, and C. Frank, “MusicLM: Generating music from text,” Arxiv , 2023
2023
Closest in time.
C. Wang, S. Chen, Y. Wu, Z. Zhang, L. Zhou, S. Liu, Z. Chen, Y. Liu, H. Wang, J. Li, L. He, S. Zhao, and F. Wei, “Neural codec language models are zero-shot text to speech synthesizers,” Arxiv , 2023
2023
Closest in time.
E. Kharitonov, D. Vincent, Z. Borsos, R. Marinier, S. Girgin, O. Pietquin, M. Sharifi, M. Tagliasacchi, and N. Zeghidour, “Speak, Read and Prompt: High-fidelity text-to-speech with minimal supervision,” ArXiv , 2023
2023
Closest in time.
T. Dao, “Flashattention-2: Faster attention with better parallelism and work partitioning,” Arxiv , 2023
2023
Closest in time.
L. Yu, D. Simig, C. Flaherty, A. Aghajanyan, L. Zettlemoyer, and M. Lewis, “Megabyte: Predicting million-byte sequences with multiscale transformers,” Arxiv , 2023
2023
Closest in time.