Fetching the paper…
Reading the bibliography…
Generating speech from a face image is crucial for developing virtual humans capable of interacting using their unique voices, without relying on pre-recorded human speech.
J. S. Chung and A. Zisserman, “Out of time: Automated lip sync in the wild,” in ACCV 2016 Workshops , ser. Lecture Notes in Computer Science, 2016
2016
Earlier work this paper cites.
A. Nagrani, J. S. Chung, and A. Zisserman, “Voxceleb: A large-scale speaker identification dataset,” in Interspeech , 2017
2017
Earlier work this paper cites.
A. van den Oord, O. Vinyals, and K. Kavukcuoglu, “Neural discrete representation learning,” in NeurIPS , 2017
2017
Earlier work this paper cites.
J. S. Chung, A. W. Senior, O. Vinyals, and A. Zisserman, “Lip reading sentences in the wild,” in CVPR , 2017
2017
Earlier work this paper cites.
S. Arik, J. Chen, K. Peng, W. Ping, and Y. Zhou, “Neural voice cloning with a few samples,” NeurIPS , 2018
2018
Earlier work this paper cites.
J. S. Chung, A. Nagrani, and A. Zisserman, “Voxceleb2: Deep speaker recognition,” in Interspeech , 2018
2018
Earlier work this paper cites.
S. Goto, K. Onishi, Y. Saito, K. Tachibana, and K. Mori, “Face2speech: Towards multi-speaker text-to-speech synthesis using an embedding vector predicted from a face image,” in Interspeech , 2020. [Online]. Available: https://doi.org/10.21437/Interspeech.2020-2136
2020
Earlier work this paper cites.
J. Kong, J. Kim, and J. Bae, “Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,” in NeurIPS , 2020
2020
Earlier work this paper cites.
D. Min, D. B. Lee, E. Yang, and S. J. Hwang, “Meta-stylespeech : Multi-speaker adaptive text-to-speech generation,” in ICML , 2021
2021
Cited alongside, same era.
Y. Ren, C. Hu, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T. Liu, “Fastspeech 2: Fast and high-quality end-to-end text to speech,” in ICLR , 2021
2021
Cited alongside, same era.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever, “Learning transferable visual models from natural language supervision,” in ICML , 2021
2021
Cited alongside, same era.
V. Popov, I. Vovk, V. Gogoryan, T. Sadekova, and M. A. Kudinov, “Grad-tts: A diffusion probabilistic model for text-to-speech,” in ICML , 2021
2021
Cited alongside, same era.
E. Casanova, J. Weber, C. D. Shulby, A. C. Júnior, E. Gölge, and M. A. Ponti, “YourTTS: Towards Zero-Shot Multi-Speaker TTS and Zero-Shot Voice Conversion for Everyone,” in ICML , 2022
Y. Ren, M. Lei, Z. Huang, S. Zhang, Q. Chen, Z. Yan, and Z. Zhao, “Prosospeech: Enhancing prosody with quantized vector pre-training in text-to-speech,” in ICASSP , 2022
2022
Later among the works it cites.
M. Kang, D. Min, and S. J. Hwang, “Grad-stylespeech: Any-speaker adaptive text-to-speech synthesis with diffusion models,” in ICASSP 2023 , 2023
2023
Closest in time.
J. Lee, J. Son Chung, and S.-W. Chung, “Imaginary voice: Face-styled diffusion model for text-to-speech,” in ICASSP , 2023
2023
Closest in time.
2023
Closest in time.
Y. Koizumi, H. Zen, S. Karita, Y. Ding, K. Yatabe, N. Morioka, M. Bacchiani, Y. Zhang, W. Han, and A. Bapna, “Libritts-r: A restored multi-speaker text-to-speech corpus,” in Interspeech , 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2022
Cited alongside, same era.
J. Wang, Z. Wang, X. Hu, X. Li, Q. Fang, and L. Liu, “Residual-guided personalized speech synthesis based on face image,” in ICASSP , 2022
2022
Cited alongside, same era.
Y. Wu, X. Tan, B. Li, L. He, S. Zhao, R. Song, T. Qin, and T. Liu, “Adaspeech 4: Adaptive text to speech in zero-shot scenarios,” in Interspeech , 2022. [Online]. Available: https://doi.org/10.21437/Interspeech.2022-901
2022
Cited alongside, same era.
R. Badlani, A. Lancucki, K. J. Shih, R. Valle, W. Ping, and B. Catanzaro, “One TTS alignment to rule them all,” in ICASSP , 2022
2022
Cited alongside, same era.
Cited in the paper.
2023
Closest in time.
Z. Ye, R. Huang, Y. Ren, Z. Jiang, J. Liu, J. He, X. Yin, and Z. Zhao, “Clapspeech: Learning prosody from text context with contrastive language-audio pre-training,” in ACL , 2023. [Online]. Available: https://doi.org/10.18653/v1/2023.acl-long.518
2023
Closest in time.
A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, “Robust speech recognition via large-scale weak supervision,” in ICML , 2023
2023
Closest in time.