Fetching the paper…
Reading the bibliography…
Video-to-speech is the process of reconstructing the audio speech from a video of a spoken utterance.
“Mel-cepstral distance measure for objective speech quality assessment”
R. Kubichek · 1993
Earlier work this paper cites.
“Perceptual evaluation of speech quality (PESQ)-a new method for speech quality assessment of telephone networks and codecs”
Antony. Rix, John. Beerends, Michael. Hollier and Andries. Hekstra · 2001
Earlier work this paper cites.
“An audio-visual corpus for speech perception and automatic speech recognition (L)”
Martin Cooke, Jon Barker, Stuart Cunningham and Xu Shao · 2006
Earlier work this paper cites.
“End-to-End Adversarial Text-to-Speech”
Jeff Donahue et al · 2006
Earlier work this paper cites.
“An efficient MFCC extraction method in speech recognition”
Wei Han, Cheong-fat Chan, Oliver-sing Choy and Kong-Pang Pun · 2006
Earlier work this paper cites.
“Profile View Lip Reading”
Kshitiz Kumar, Tsuhan Chen and Richard. Stern · 2007
Earlier work this paper cites.
“Real-Time Signal Estimation From Modified Short-Time Fourier Transform Magnitude Spectra”
Xinglei Zhu, Gerald Beauregard and Lonce. Wyse · 2007
Earlier work this paper cites.
“Information theoretic feature extraction for audio-visual speech recognition”
Mihai Gurban and Jean-Philippe Thiran · 2009
Earlier work this paper cites.
“Dlib-ml: A Machine Learning Toolkit”
Davis. King · 2009
Earlier work this paper cites.
“Lipreading With Local Spatiotemporal Descriptors”
Guoying Zhao, Mark Barnard and Matti Pietik“”ainen · 2009
Earlier work this paper cites.
“Speech Quality Assessment”
Philipos. Loizou · 2011
Earlier work this paper cites.
“An Algorithm for Intelligibility Prediction of Time-Frequency Weighted Noisy Speech”
Cees. Taal, Richard. Hendriks, Richard Heusdens and Jesper Jensen · 2011
Earlier work this paper cites.
“Speech synthesis techniques. A survey”
Y. Tabet and M. Boughazi · 2011
Earlier work this paper cites.
“Recurrent Neural Networks for Noise Reduction in Robust ASR”
Andrew. Maas et al · 2012
Earlier work this paper cites.
“Emotion recognition in speech using MFCC and wavelet features”
K.. Krishna Kishore and P. Krishna Satish · 2013
Earlier work this paper cites.
“Robust Audio-Visual Speech Recognition Under Noisy Audio-Video Conditions”
Darryl Stewart, Rowan Seymour, Adrian Pass and Ji Ming · 2013
Earlier work this paper cites.
“Generative Adversarial Networks”
Ian. Goodfellow et al · 2014
Earlier work this paper cites.
“OuluVS2: A multi-view audiovisual database for non-rigid mouth motion analysis”
Iryna Anina, Ziheng Zhou, Guoying Zhao and Matti Pietik“”ainen · 2015
Earlier work this paper cites.
“Reconstructing intelligible audio speech from visual speech features”
Thomas Cornu and Ben Milner · 2015
Earlier work this paper cites.
“TCD-TIMIT: An Audio-Visual Corpus of Continuous Speech”
N. Harte and E. Gillen · 2015
Earlier work this paper cites.
“LipNet: Sentence-level Lipreading”
Yannis. Assael, Brendan Shillingford, Shimon Whiteson and Nando de Freitas · 2016
Cited alongside, same era.
“Lip Reading in the Wild”
J.. Chung and A. Zisserman · 2016
Cited alongside, same era.
“Lip Reading in the Wild”
Joon Chung and Andrew Zisserman · 2016
Cited alongside, same era.
“Deep Residual Learning for Image Recognition”
Kaiming He, Xiangyu Zhang, Shaoqing Ren and Jian Sun · 2016
Cited alongside, same era.
“WORLD: A Vocoder-Based High-Quality Speech Synthesis System for Real-Time Applications”
Masanori Morise, Fumiya Yokomori and Kenji Ozawa · 2016
Cited alongside, same era.
“WaveNet: A Generative Model for Raw Audio”
A“”aron van Oord et al · 2016
Cited alongside, same era.
“Parallel WaveNet: Fast High-Fidelity Speech Synthesis”
A“”aron van Oord et al · 2018
Later among the works it cites.
“End-to-End Audiovisual Speech Recognition”
Stavros Petridis et al · 2018
Later among the works it cites.
“Natural TTS Synthesis by Conditioning Wavenet on MEL Spectrogram Predictions”
Jonathan Shen et al · 2018
Later among the works it cites.
“Visual to Sound: Generating Natural Sound for Videos in the Wild”
Yipin Zhou et al · 2018
Later among the works it cites.
“Adversarial Audio Synthesis”
Chris Donahue, Julian. McAuley and Miller. Puckette · 2019
Later among the works it cites.
“MelGAN: Generative Adversarial Networks for Conditional Waveform Synthesis”
Kundan Kumar et al · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Visually Indicated Sounds”
Andrew Owens et al · 2016
Cited alongside, same era.
Mart“’n Arjovsky, Soumith Chintala and L“’eon Bottou · 2017
Cited alongside, same era.
“Deep Cross-Modal Audio-Visual Generation”
Lele Chen, Sudhanshu Srivastava, Zhiyao Duan and Chenliang Xu · 2017
Cited alongside, same era.
“Generating Intelligible Audio Speech From Visual Speech”
Thomas Cornu and Ben Milner · 2017
Cited alongside, same era.
“Improved Speech Reconstruction from Silent Video”
Ariel Ephrat, Tavi Halperin and Shmuel Peleg · 2017
Cited alongside, same era.
“Vid2speech: Speech reconstruction from silent video”
Ariel Ephrat and Shmuel Peleg · 2017
Cited alongside, same era.
“Lipper: Synthesizing Thy Speech Using Multi-View Lipreading”
Yaman Kumar et al · 2019
Later among the works it cites.
“Jasper: An End-to-End Convolutional Neural Acoustic Model”
Jason Li et al · 2019
Later among the works it cites.
“Investigating the lombard effect influence on end-to-end audio-visual speech recognition”
Pingchuan Ma, Stavros Petridis and Maja Pantic · 2019
Later among the works it cites.
“Learning Problem-Agnostic Speech Representations from Multiple Self-Supervised Tasks”
Santiago Pascual et al · 2019
Later among the works it cites.
“Large-Scale Visual Speech Recognition”
Brendan Shillingford et al · 2019
Later among the works it cites.
“Hush-Hush Speak: Speech Reconstruction Using Silent Videos”
Shashwat Uttam et al · 2019
Later among the works it cites.
“Video-Driven Speech Reconstruction Using Generative Adversarial Networks”
Konstantinos Vougioukas, Pingchuan Ma, Stavros Petridis and Maja Pantic · 2019
Later among the works it cites.
“Vision-Infused Deep Audio Inpainting”
Hang Zhou et al · 2019
Later among the works it cites.
“Vocoder-Based Speech Synthesis from Silent Videos”
Daniel Michelsanti et al · 2020
Later among the works it cites.
“End-to-end visual speech recognition for small-scale datasets”
Stavros Petridis et al · 2020
Later among the works it cites.
“Learning Individual Speaking Styles for Accurate Lip to Speech Synthesis”
K.. Prajwal, Rudrabha Mukhopadhyay, Vinay. Namboodiri and C.. Jawahar · 2020
Later among the works it cites.
“Multi-Task Self-Supervised Learning for Robust Speech Recognition”
Mirco Ravanelli et al · 2020
Later among the works it cites.
“Realistic Speech-Driven Facial Animation with GANs”
Konstantinos Vougioukas, Stavros Petridis and Maja Pantic · 2020
Later among the works it cites.
“Parallel Wavegan: A Fast Waveform Generation Model Based on Generative Adversarial Networks with Multi-Resolution Spectrogram”
Ryuichi Yamamoto, Eunwoo Song and Jae-Min Kim · 2020
Later among the works it cites.