Fetching the paper…
Reading the bibliography…
The advent of Large Models marks a new era in machine learning, significantly outperforming smaller models by leveraging vast datasets to capture and synthesize complex patterns.
Perceptual evaluation of speech quality (pesq): A new method for speech quality assessment of telephone networks and codecs
Rix, A., Beerends, J., Hollier, M., and Hekstra, A · 2001
Earlier work this paper cites.
Training deep nets with sublinear memory cost, 2016
Chen, T., Xu, B., Zhang, C., and Guestrin, C · 2016
Earlier work this paper cites.
Wavenet: A generative model for raw audio
van den Oord, A., Dieleman, S., Zen, H., Simonyan, K., Vinyals, O., Graves, A., Kalchbrenner, N., Senior, A., and Kavukcuoglu, K · 2016
Earlier work this paper cites.
Sigmoid-weighted linear units for neural network function approximation in reinforcement learning, 2017
Elfwing, S., Uchibe, E., and Doya, K · 2017
Earlier work this paper cites.
The lj speech dataset
Ito, K. and Johnson, L · 2017
Earlier work this paper cites.
The 2016 signal separation evaluation campaign
Liutkus, A., Stöter, F.-R., Rafii, Z., Kitamura, D., Rivet, B., Ito, N., Ono, N., and Fontecave, J · 2017
Earlier work this paper cites.
Least squares generative adversarial networks, 2017
Mao, X., Li, Q., Xie, H., Lau, R. Y. K., Wang, Z., and Smolley, S. P · 2017
Earlier work this paper cites.
Cstr vctk corpus: English multi-speaker corpus for cstr voice cloning toolkit
Veaux, C., Yamagishi, J., and MacDonald, K · 2017
Earlier work this paper cites.
Efficient neural audio synthesis
Kalchbrenner, N., Elsen, E., Simonyan, K., Noury, S., Casagrande, N., Lockhart, E., Stimberg, F., van den Oord, A., Dieleman, S., and Kavukcuoglu, K · 2018
Earlier work this paper cites.
Waveglow: A flow-based generative network for speech synthesis
Prenger, R., Valle, R., and Catanzaro, B · 2019
Earlier work this paper cites.
Libritts: A corpus derived from librispeech for text-to-speech, 2019
Zen, H., Dang, V., Clark, R., Zhang, Y., Weiss, R. J., Jia, Y., Chen, Z., and Wu, Y · 2019
Cited alongside, same era.
Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis, 2020
Kong, J., Kim, J., and Bae, J · 2020
Cited alongside, same era.
auraloss: Audio focused loss functions in PyTorch
Steinmetz, C. J. and Reiss, J. D · 2020
Cited alongside, same era.
Parallel wavegan: A fast waveform generation model based on generative adversarial networks with multi-resolution spectrogram
Yamamoto, R., Song, E., and Kim, J.-M · 2020
Cited alongside, same era.
Multi-band melgan: Faster waveform generation for high-quality text-to-speech
Yang, Y. R., Chen, Y.-A., and Tsao, Y · 2020
Cited alongside, same era.
Transferring neural speech waveform synthesizers to musical instrument sounds generation
Singgan: Generative adversarial network for high-fidelity singing voice generation, 2022
Huang, R., Cui, C., Chen, F., Ren, Y., Liu, J., Zhao, Z., Huai, B., and Wang, Z · 2022
Later among the works it cites.
istftnet: Fast and lightweight mel-spectrogram vocoder incorporating inverse short-time fourier transform, 2022
Kaneko, T., Tanaka, K., Kameoka, H., and Seki, S · 2022
Later among the works it cites.
A convnet for the 2020s, 2022
Liu, Z., Mao, H., Wu, C.-Y., Feichtenhofer, C., Darrell, T., and Xie, S · 2022
Later among the works it cites.
Chunked autoregressive gan for conditional waveform synthesis, 2022
Morrison, M., Kumar, R., Kumar, K., Seetharaman, P., Courville, A., and Bengio, Y · 2022
Later among the works it cites.
Release nsf-hifigan with 44.1 khz sampling rate · openvpi/vocoders, Dec 2022
Openvpi · 2022
Later among the works it cites.
M4singer: A multi-style, multi-singer and musical score provided mandarin singing corpus
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zhao, Y., Wang, X., Juvela, L., and Yamagishi, J · 2020
Cited alongside, same era.
Univnet: A neural vocoder with multi-resolution spectrogram discriminators for high-fidelity waveform generation, 2021
Jang, W., Lim, D., Yoon, J., Kim, B., and Kim, J · 2021
Cited alongside, same era.
Xu, S., Zhao, W., and Guo, J · 2021
Cited alongside, same era.
High fidelity neural audio compression, 2022
Défossez, A., Copet, J., Synnaeve, G., and Adi, Y · 2022
Cited alongside, same era.
Zhang, L., Li, R., Wang, S., Deng, L., Liu, J., Ren, Y., He, J., Huang, R., Zhu, J., Chen, X., and Zhao, Z · 2022
Later among the works it cites.
Bigvgan: A universal neural vocoder with large-scale training, 2023
gil Lee, S., Ping, W., Ginsburg, B., Catanzaro, B., and Yoon, S · 2023
Later among the works it cites.
Multi-scale sub-band constant-q transform discriminator for high-fidelity vocoder, 2023
Gu, Y., Zhang, X., Xue, L., and Wu, Z · 2023
Later among the works it cites.
Vocos: Closing the gap between time-domain and fourier-based neural vocoders for high-quality audio synthesis, 2023
Siuzdak, H · 2023
Later among the works it cites.