Fetching the paper…
Reading the bibliography…
In this paper, we propose a multi-speaker face-to-speech waveform generation model that also works for unseen speaker conditions.
“Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,”
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber, · 2006
Earlier work this paper cites.
“An audio-visual corpus for speech perception and automatic speech recognition,”
Martin Cooke, Jon Barker, Stuart Cunningham, and Xu Shao, · 2006
Earlier work this paper cites.
“Facenet: A unified embedding for face recognition and clustering,”
Florian Schroff, Dmitry Kalenichenko, and James Philbin, · 2015
Earlier work this paper cites.
“Lipnet,”
Yannis Assael, Brendan Shillingford, Shimon Whiteson, and Nando de Freitas, · 2016
Earlier work this paper cites.
“WORLD: a vocoder-based high-quality speech synthesis system for real-time applications,”
Masanori Morise, Fumiya Yokomori, and Kenji Ozawa, · 2016
Earlier work this paper cites.
“Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,”
William Chan, Navdeep Jaitly, Quoc Le, and Oriol Vinyals, · 2016
Earlier work this paper cites.
“Lip reading sentences in the wild,”
Joon Son Chung, Andrew Senior, Oriol Vinyals, and Andrew Zisserman, · 2017
Earlier work this paper cites.
“Vid2speech: speech reconstruction from silent video,”
Ariel Ephrat and Shmuel Peleg, · 2017
Cited alongside, same era.
“Lip2audspec: Speech reconstruction from silent lip movements video,”
Hassan Akbari, Himani Arora, Liangliang Cao, and Nima Mesgarani, · 2018
Cited alongside, same era.
“Video-Driven Speech Reconstruction Using Generative Adversarial Networks,”
Konstantinos Vougioukas, Pingchuan Ma, Stavros Petridis, and Maja Pantic, · 2019
Cited alongside, same era.
“Lipper: Synthesizing thy speech using multi-view lipreading,”
Yaman Kumar, Rohit Jain, Khwaja Mohd Salik, Rajiv Ratn Shah, Yifang Yin, and Roger Zimmermann, · 2019
Cited alongside, same era.
“Speech2face: Learning the face behind a voice,”
Tae-Hyun Oh, Tali Dekel, Changil Kim, Inbar Mosseri, William T Freeman, Michael Rubinstein, and Wojciech Matusik, · 2019
Cited alongside, same era.
“Learning individual speaking styles for accurate lip to speech synthesis,”
“Mellotron: Multispeaker expressive voice synthesis by conditioning on rhythm, pitch and global style tokens,”
Rafael Valle, Jason Li, Ryan Prenger, and Bryan Catanzaro, · 2020
Later among the works it cites.
“Vocoder-Based Speech Synthesis from Silent Videos,”
Daniel Michelsanti, Olga Slizovskaia, Gloria Haro, Emilia Gómez, Zheng-Hua Tan, and Jesper Jensen, · 2020
Later among the works it cites.
“Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,”
Jungil Kong, Jaehyeon Kim, and Jaekyoung Bae, · 2020
Later among the works it cites.
“Emotional speech synthesis with rich and granularized control,”
Se-Yun Um, Sangshin Oh, Kyungguen Byun, Inseon Jang, ChungHyun Ahn, and Hong-Goo Kang, · 2020
Later among the works it cites.
“Deep learning based assessment of synthetic speech naturalness,”
Gabriel Mittag and Sebastian Möller, · 2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
KR Prajwal, Rudrabha Mukhopadhyay, Vinay P Namboodiri, and CV Jawahar, · 2020
Cited alongside, same era.
“Face2speech: Towards multi-speaker text-to-speech synthesis using an embedding vector predicted from a face image,”
Shunsuke Goto, Kotaro Onishi, Yuki Saito, Kentaro Tachibana, and Koichiro Mori, · 2020
Cited alongside, same era.
“Vcvts: Multi-speaker video-to-speech synthesis via cross-modal knowledge transfer from voice conversion,”
Disong Wang, Shan Yang, Dan Su, Xunying Liu, Dong Yu, and Helen Meng, · 2022
Closest in time.