Visual to sound: Generating natural sound for videos in the wild
Zhou, Y.; Wang, Z.; Fang, C.; Bui, T.; and Berg, T. L. 2018 · 2018
Cited alongside, same era.
Large scale adversarial representation learning
Donahue, J.; and Simonyan, K. 2019 · 2019
Cited alongside, same era.
A style-based generator architecture for generative adversarial networks
Karras, T.; Laine, S.; and Aila, T. 2019 · 2019
Cited alongside, same era.
Melgan: Generative adversarial networks for conditional waveform synthesis
Kumar, K.; Kumar, R.; de Boissiere, T.; Gestin, L.; Teoh, W. Z.; Sotelo, J.; de Brébisson, A.; Bengio, Y.; and Courville, A. C. 2019 · 2019
Cited alongside, same era.
Disentangling disentanglement in variational autoencoders
Mathieu, E.; Rainforth, T.; Siddharth, N.; and Teh, Y. W. 2019 · 2019
Cited alongside, same era.
Autovc: Zero-shot voice style transfer with only autoencoder loss
Qian, K.; Zhang, Y.; Chang, S.; Yang, X.; and Hasegawa-Johnson, M. 2019 · 2019
Cited alongside, same era.
Generating visually aligned sound from videos
Chen, P.; Zhang, Y.; Tan, M.; Xiao, H.; Huang, D.; and Gan, C. 2020 · 2020
Cited alongside, same era.
HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis
Kong, J.; Kim, J.; and Bae, J. 2020 · 2020
Cited alongside, same era.
Unsupervised speech decomposition via triple information bottleneck
Qian, K.; Zhang, Y.; Chang, S.; Hasegawa-Johnson, M.; and Cox, D. 2020 · 2020
Cited alongside, same era.
FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Ren, Y.; Hu, C.; Tan, X.; Qin, T.; Zhao, S.; Zhao, Z.; and Liu, T.-Y. 2020 · 2020
Cited alongside, same era.
Parallel WaveGAN: A fast waveform generation model based on generative adversarial networks with multi-resolution spectrogram
Yamamoto, R.; Song, E.; and Kim, J.-M. 2020 · 2020
Cited alongside, same era.
Singgan: Generative adversarial network for high-fidelity singing voice generation
Huang, R.; Cui, C.; Chen, F.; Ren, Y.; Liu, J.; Zhao, Z.; Huai, B.; and Wang, Z. 2022a
Cited in the paper.