Fetching the paper…
Reading the bibliography…
Video-to-speech synthesis (also known as lip-to-speech) refers to the translation of silent lip movements into the corresponding audio.
“Signal estimation from modified short-time Fourier transform”
D. Griffin and Jae Lim · 1984
Earlier work this paper cites.
“Perceptual evaluation of speech quality (PESQ)-a new method for speech quality assessment of telephone networks and codecs”
Antony. Rix et al · 2001
Earlier work this paper cites.
“An audio-visual corpus for speech perception and automatic speech recognition (L)”
Martin Cooke et al · 2006
Earlier work this paper cites.
“An Algorithm for Intelligibility Prediction of Time-Frequency Weighted Noisy Speech”
Cees. Taal et al · 2011
Earlier work this paper cites.
“A fast Griffin-Lim algorithm”
Nathana“”el Perraudin, P“’eter Bal“’azs and Peter. Sndergaard · 2013
Earlier work this paper cites.
“Reconstructing intelligible audio speech from visual speech features”
Thomas Cornu and Ben Milner · 2015
Earlier work this paper cites.
“TCD-TIMIT: An Audio-Visual Corpus of Continuous Speech”
N. Harte and E. Gillen · 2015
Earlier work this paper cites.
“Librispeech: An ASR corpus based on public domain audio books”
Vassil Panayotov et al · 2015
Earlier work this paper cites.
“LipNet: Sentence-level Lipreading”
Yannis. Assael et al · 2016
Earlier work this paper cites.
“Deep Residual Learning for Image Recognition”
Kaiming He et al · 2016
Earlier work this paper cites.
“Sgdr: Stochastic gradient descent with warm restarts”
Ilya Loshchilov and Frank Hutter · 2016
Earlier work this paper cites.
“WaveNet: A Generative Model for Raw Audio”
A“”aron van Oord et al · 2016
Earlier work this paper cites.
“How far are we from solving the 2D & 3D Face Alignment problem? (and a dataset of 230,000 3D facial landmarks)”
Adrian Bulat and Georgios Tzimiropoulos · 2017
Earlier work this paper cites.
“Generating Intelligible Audio Speech From Visual Speech”
Thomas Cornu and Ben Milner · 2017
Earlier work this paper cites.
“Improved Speech Reconstruction from Silent Video”
Ariel Ephrat, Tavi Halperin and Shmuel Peleg · 2017
Cited alongside, same era.
“Vid2speech: Speech reconstruction from silent video”
Ariel Ephrat and Shmuel Peleg · 2017
Cited alongside, same era.
“Decoupled weight decay regularization”
Ilya Loshchilov and Frank Hutter · 2017
Cited alongside, same era.
“VoxCeleb: A Large-Scale Speaker Identification Dataset”
Arsha Nagrani, Joon Chung and Andrew Zisserman · 2017
Cited alongside, same era.
“LRS3-TED: a large-scale dataset for visual speech recognition”
T. Afouras, J.. Chung and A. Zisserman · 2018
Cited alongside, same era.
“Lip2Audspec: Speech Reconstruction from Silent Lip Movements Video”
“Vocoder-Based Speech Synthesis from Silent Videos”
Daniel Michelsanti et al · 2020
Later among the works it cites.
“Learning Individual Speaking Styles for Accurate Lip to Speech Synthesis”
K.. Prajwal et al · 2020
Later among the works it cites.
“Parallel Wavegan: A Fast Waveform Generation Model Based on Generative Adversarial Networks with Multi-Resolution Spectrogram”
Ryuichi Yamamoto, Eunwoo Song and Jae-Min Kim · 2020
Later among the works it cites.
“Speech Reconstruction With Reminiscent Sound Via Visual Voice Memory”
Joanna Hong et al · 2021
Later among the works it cites.
“Lip to Speech Synthesis with Visual Context Attentional GAN”
Minsu Kim, Joanna Hong and Yong Ro · 2021
Later among the works it cites.
“End-To-End Audio-Visual Speech Recognition with Conformers”
Pingchuan Ma, Stavros Petridis and Maja Pantic · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hassan Akbari et al · 2018
Cited alongside, same era.
“VoxCeleb2: Deep Speaker Recognition”
J.. Chung, A. Nagrani and A. Zisserman · 2018
Cited alongside, same era.
“Transfer Learning from Speaker Verification to Multispeaker Text-To-Speech Synthesis”
Ye Jia et al · 2018
Cited alongside, same era.
“End-to-End Audiovisual Speech Recognition”
Stavros Petridis et al · 2018
Cited alongside, same era.
“Natural TTS Synthesis by Conditioning Wavenet on MEL Spectrogram Predictions”
Jonathan Shen et al · 2018
Cited alongside, same era.
“Video-Driven Speech Reconstruction Using Generative Adversarial Networks”
Konstantinos Vougioukas et al · 2019
Cited alongside, same era.
“LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech”
Heiga Zen et al · 2019
Cited alongside, same era.
Later among the works it cites.
“StyleMelGAN: An Efficient High-Fidelity Adversarial Vocoder with Temporal Adaptive Normalization”
Ahmed Mustafa, Nicola Pia and Guillaume Fuchs · 2021
Later among the works it cites.
“Speaker disentanglement in video-to-speech conversion”
Dan Oneata, Adriana Stan and Horia Cucu · 2021
Later among the works it cites.
“Facetron: Multi-speaker Face-to-Speech Model based on Cross-modal Latent Representations”
Se-Yun Um et al · 2021
Later among the works it cites.
“Speech Prediction in Silent Videos Using Variational Autoencoders”
Ravindra Yadav et al · 2021
Later among the works it cites.
“Multi-Band Melgan: Faster Waveform Generation For High-Quality Text-To-Speech”
Geng Yang et al · 2021
Later among the works it cites.
“Visual Speech Recognition for Multiple Languages in the Wild”
Pingchuan Ma, Stavros Petridis and Maja Pantic · 2022
Closest in time.
“End-to-End Video-to-Speech Synthesis Using Generative Adversarial Networks”
Rodrigo Mira et al · 2022
Closest in time.
“Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction”
Bowen Shi et al · 2022
Closest in time.