Fetching the paper…
Reading the bibliography…
Amphion is an open-source toolkit for Audio, Music, and Speech Generation, targeting to ease the way for junior researchers and engineers into these fields.
“WaveNet: A generative model for raw audio,”
Aäron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew W. Senior, and Koray Kavukcuoglu, · 2016
Earlier work this paper cites.
“Efficient neural audio synthesis,”
Nal Kalchbrenner, Erich Elsen, Karen Simonyan, Seb Noury, Norman Casagrande, Edward Lockhart, Florian Stimberg, Aäron van den Oord, Sander Dieleman, and Koray Kavukcuoglu, · 2018
Earlier work this paper cites.
“Waveglow: A flow-based generative network for speech synthesis,”
Ryan Prenger, Rafael Valle, and Bryan Catanzaro, · 2019
Earlier work this paper cites.
“MelGAN: Generative adversarial networks for conditional waveform synthesis,”
Kundan Kumar, Rithesh Kumar, Thibault de Boissiere, Lucas Gestin, Wei Zhen Teoh, Jose Sotelo, Alexandre de Brébisson, Yoshua Bengio, and Aaron C. Courville, · 2019
Earlier work this paper cites.
“AudioCaps: Generating captions for audios in the wild,”
Chris Dongjoo Kim, Byeongchang Kim, Hyunmin Lee, and Gunhee Kim, · 2019
Earlier work this paper cites.
“LibriTTS: A corpus derived from librispeech for text-to-speech,”
Heiga Zen, Viet Dang, Rob Clark, Yu Zhang, Ron J. Weiss, Ye Jia, Zhifeng Chen, and Yonghui Wu, · 2019
Earlier work this paper cites.
“FastSpeech 2: Fast and high-quality end-to-end text to speech,”
Yi Ren, Chenxu Hu, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu, · 2020
Earlier work this paper cites.
“Hifi-gan: High-fidelity denoising and dereverberation based on speech deep features in adversarial networks,”
Jiaqi Su, Zeyu Jin, and Adam Finkelstein, · 2020
Earlier work this paper cites.
Won Jang, Dan Lim, and Jaesam Yoon, · 2020
Earlier work this paper cites.
“MLS: A large-scale multilingual dataset for speech research,”
Vineel Pratap, Qiantong Xu, Anuroop Sriram, Gabriel Synnaeve, and Ronan Collobert, · 2020
Earlier work this paper cites.
“Libri-light: A benchmark for ASR with limited or no supervision,”
Jacob Kahn, Morgane Rivière, Weiyi Zheng, Evgeny Kharitonov, Qiantong Xu, Pierre-Emmanuel Mazaré, Julien Karadayi, Vitaliy Liptchinsky, Ronan Collobert, Christian Fuegen, Tatiana Likhomanenko, Gabriel Synnaeve, Armand Joulin, Abdelrahman Mohamed, and Emmanuel Dupoux, · 2020
Earlier work this paper cites.
“PANNs: Large-scale pretrained audio neural networks for audio pattern recognition,”
Qiuqiang Kong, Yin Cao, Turab Iqbal, Yuxuan Wang, Wenwu Wang, and Mark D Plumbley, · 2020
Earlier work this paper cites.
“Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,”
Jaehyeon Kim, Jungil Kong, and Juhee Son, · 2021
Cited alongside, same era.
“DiffWave: A versatile diffusion model for audio synthesis,”
Zhifeng Kong, Wei Ping, Jiaji Huang, Kexin Zhao, and Bryan Catanzaro, · 2021
Cited alongside, same era.
“Hi-Fi Multi-Speaker English TTS Dataset,”
Evelina Bakhturina, Vitaly Lavrukhin, Boris Ginsburg, and Yang Zhang, · 2021
Cited alongside, same era.
“Hubert: Self-supervised speech representation learning by masked prediction of hidden units,”
Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed, · 2021
Cited alongside, same era.
“Multi-singer: Fast multi-singer singing voice vocoder with A large-scale corpus,”
Rongjie Huang, Feiyang Chen, Yi Ren, Jinglin Liu, Chenye Cui, and Zhou Zhao, · 2021
Cited alongside, same era.
“Bigvgan: A universal neural vocoder with large-scale training,”
Sang-gil Lee, Wei Ping, Boris Ginsburg, Bryan Catanzaro, and Sungroh Yoon, · 2023
Closest in time.
“APNet: An all-frame-level neural vocoder incorporating direct prediction of amplitude and phase spectra,”
Yang Ai and Zhen-Hua Ling, · 2023
Closest in time.
“Diffsound: Discrete diffusion model for text-to-sound generation,”
Dongchao Yang, Jianwei Yu, Helin Wang, Wen Wang, Chao Weng, Yuexian Zou, and Dong Yu, · 2023
Closest in time.
“Naturalspeech 2: Latent diffusion models are natural and zero-shot speech and singing synthesizers,”
Kai Shen, Zeqian Ju, Xu Tan, Yanqing Liu, Yichong Leng, Lei He, Tao Qin, Sheng Zhao, and Jiang Bian, · 2024
Closest in time.
“SingVisio: Visual Analytics of Diffusion Model for Singing Voice Conversion,”
Liumeng Xue and Chaoren Wang and Mingxuan Wang and Xueyao Zhang and Jun Han and Zhizheng Wu, · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Alexandre Défossez, Jade Copet, Gabriel Synnaeve, and Yossi Adi, · 2022
Cited alongside, same era.
“Opencpop: A high-quality open source chinese popular song corpus for singing voice synthesis,”
Yu Wang, Xinsheng Wang, Pengcheng Zhu, Jie Wu, Hanzhao Li, Heyang Xue, Yongmao Zhang, Lei Xie, and Mengxiao Bi, · 2022
Cited alongside, same era.
“M4singer: A multi-style, multi-singer and musical score provided mandarin singing corpus,”
Lichao Zhang, Ruiqi Li, Shoutong Wang, Liqun Deng, Jinglin Liu, Yi Ren, Jinzheng He, Rongjie Huang, Jieming Zhu, Xiao Chen, and Zhou Zhao, · 2022
Cited alongside, same era.
“Neural codec language models are zero-shot text to speech synthesizers,”
Chengyi Wang, Sanyuan Chen, Yu Wu, Ziqiang Zhang, Long Zhou, Shujie Liu, Zhuo Chen, Yanqing Liu, Huaming Wang, Jinyu Li, et al., · 2023
Cited alongside, same era.
“Audioldm: Text-to-audio generation with latent diffusion models,”
Haohe Liu, Zehua Chen, Yi Yuan, Xinhao Mei, Xubo Liu, Danilo P. Mandic, Wenwu Wang, and Mark D. Plumbley, · 2023
Cited alongside, same era.
“AUDIT: Audio editing by following instructions with latent diffusion models,”
Yuancheng Wang, Zeqian Ju, Xu Tan, Lei He, Zhizheng Wu, Jiang Bian, and Sheng Zhao, · 2023
Cited alongside, same era.
“Source-filter hifi-gan: Fast and pitch controllable high-fidelity neural vocoder,”
Reo Yoneyama, Yi-Chiao Wu, and Tomoki Toda, · 2023
Cited alongside, same era.
“Leveraging diverse semantic-based audio pretrained models for singing voice conversion,”
Xueyao Zhang, Zihao Fang, Yicheng Gu, Haopeng Chen, Lexiao Zou, Junan Zhang, Liumeng Xue, and Zhizheng Wu, · 2024
Closest in time.
“Comosvc: Consistency model-based singing voice conversion,”
Yiwen Lu, Zhen Ye, Wei Xue, Xu Tan, Qifeng Liu, and Yike Guo, · 2024
Closest in time.
Zeyu Xie and Xuenan Xu and Zhizheng Wu and Mengyue Wu, · 2024
Closest in time.
“Multi-scale sub-band constant-q transform discriminator for high-fidelity vocoder,”
Yicheng Gu, Xueyao Zhang, Liumeng Xue, and Zhizheng Wu, · 2024
Closest in time.
“An Investigation of Time-Frequency Representation Discriminators for High-Fidelity Vocoder,”
Yicheng Gu and Xueyao Zhang and Liumeng Xue and Haizhou Li and Zhizheng Wu, · 2024
Closest in time.
“Naturalspeech 3: Zero-shot speech synthesis with factorized codec and diffusion models,”
Zeqian Ju, Yuancheng Wang, Kai Shen, Xu Tan, Detai Xin, Dongchao Yang, Yanqing Liu, Yichong Leng, Kaitao Song, Siliang Tang, Zhizheng Wu, Tao Qin, Xiang-Yang Li, Wei Ye, Shikun Zhang, Jiang Bian, Lei He, Jinyu Li, and Sheng Zhao, · 2024
Closest in time.
“Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation,”
He, Haorui and Shang, Zengqiang and Wang, Chaoren and Li, Xuyuan and Gu, Yicheng and Hua, Hua and Liu, Liwei and Yang, Chen and Li, Jiaqi and Shi, Peiyang and Wang, Yuancheng and Chen, Kai and Zhang, Pengyuan and Wu, Zhizheng, · 2024
Closest in time.