Fetching the paper…
Reading the bibliography…
Prior works on improving speech quality with visual input typically study each type of auditory distortion separately (e.g., separation, inpainting, video-to-speech) and present tailored algorithms.
Mel-cepstral distance measure for objective speech quality assessment
Robert Kubichek · 1993
Earlier work this paper cites.
Speech perception by humans and machines
Richard Lippmann · 1996
Earlier work this paper cites.
A glimpsing model of speech perception in noise
Martin Cooke · 2006
Earlier work this paper cites.
An audio-visual corpus for speech perception and automatic speech recognition
Martin Cooke, Jon Barker, Stuart Cunningham, and Xu Shao · 2006
Earlier work this paper cites.
Audio inpainting
Amir Adler, Valentin Emiya, Maria G Jafari, Michael Elad, Rémi Gribonval, and Mark D Plumbley · 2011
Earlier work this paper cites.
Crowdmos: An approach for crowdsourcing mean opinion score studies
Flávio Ribeiro, Dinei Florêncio, Cha Zhang, and Michael Seltzer · 2011
Earlier work this paper cites.
An algorithm for intelligibility prediction of time–frequency weighted noisy speech
Cees H Taal, Richard C Hendriks, Richard Heusdens, and Jesper Jensen · 2011
Earlier work this paper cites.
Ideal ratio mask estimation using deep neural networks for robust speech recognition
Arun Narayanan and DeLiang Wang · 2013
Earlier work this paper cites.
Phase-sensitive and recognition-boosted speech separation using deep recurrent neural networks
Hakan Erdogan, John R Hershey, Shinji Watanabe, and Jonathan Le Roux · 2015
Earlier work this paper cites.
Lip reading in the wild
Joon Son Chung and Andrew Zisserman · 2016
Earlier work this paper cites.
Out of time: automated lip sync in the wild
Joon Son Chung and Andrew Zisserman · 2016
Earlier work this paper cites.
Gaussian error linear units (gelus)
Dan Hendrycks and Kevin Gimpel · 2016
Earlier work this paper cites.
Vid2speech: speech reconstruction from silent video
Ariel Ephrat and Shmuel Peleg · 2017
Earlier work this paper cites.
Aviv Gabbay, Asaph Shamir, and Shmuel Peleg · 2017
Earlier work this paper cites.
A consolidated perspective on multimicrophone speech enhancement and source separation
Sharon Gannot, Emmanuel Vincent, Shmulik Markovich-Golan, and Alexey Ozerov · 2017
Earlier work this paper cites.
Audio set: An ontology and human-labeled dataset for audio events
Jort F Gemmeke, Daniel PW Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R Channing Moore, Manoj Plakal, and Marvin Ritter · 2017
Earlier work this paper cites.
The lj speech dataset
Keith Ito and Linda Johnson · 2017
Earlier work this paper cites.
Time-frequency masking in the complex domain for speech dereverberation and denoising
Donald S Williamson and DeLiang Wang · 2017
Earlier work this paper cites.
Lrs3-ted: a large-scale dataset for visual speech recognition
Triantafyllos Afouras, Joon Son Chung, and Andrew Zisserman · 2018
Earlier work this paper cites.
Ariel Ephrat, Inbar Mosseri, Oran Lang, Tali Dekel, Kevin Wilson, Avinatan Hassidim, William T Freeman, and Michael Rubinstein · 2018
Cited alongside, same era.
Tasnet: time-domain audio separation network for real-time, single-channel speech separation
Yi Luo and Nima Mesgarani · 2018
Cited alongside, same era.
Common voice: A massively-multilingual speech corpus
Rosana Ardila, Megan Branson, Kelly Davis, Michael Henretty, Michael Kohler, Josh Meyer, Reuben Morais, Lindsay Saunders, Francis M Tyers, and Gregor Weber · 2019
Cited alongside, same era.
Sonyc: A system for monitoring, analyzing, and mitigating urban noise pollution
Juan P Bello, Claudio Silva, Oded Nov, R Luke Dubois, Anish Arora, Justin Salamon, Charles Mydlarz, and Harish Doraiswamy · 2019
Cited alongside, same era.
Conv-tasnet: Surpassing ideal time–frequency magnitude masking for speech separation
On generative spoken language modeling from raw audio
Kushal Lakhotia, Eugene Kharitonov, Wei-Ning Hsu, Yossi Adi, Adam Polyak, Benjamin Bolte, Tu-Anh Nguyen, Jade Copet, Alexei Baevski, Abdelrahman Mohamed, et al · 2021
Later among the works it cites.
Direct speech-to-speech translation with discrete units
Ann Lee, Peng-Jen Chen, Changhan Wang, Jiatao Gu, Xutai Ma, Adam Polyak, Yossi Adi, Qing He, Yun Tang, Juan Pino, et al · 2021
Later among the works it cites.
Voicefixer: Toward general speech restoration with neural vocoder
Haohe Liu, Qiuqiang Kong, Qiao Tian, Yan Zhao, DeLiang Wang, Chuanzeng Huang, and Yuxuan Wang · 2021
Later among the works it cites.
End-to-end audio-visual speech recognition with conformers
Pingchuan Ma, Stavros Petridis, and Maja Pantic · 2021
Later among the works it cites.
Audio-visual speech inpainting with deep learning
Giovanni Morrone, Daniel Michelsanti, Zheng-Hua Tan, and Jesper Jensen · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yi Luo and Nima Mesgarani · 2019
Cited alongside, same era.
Specaugment: A simple data augmentation method for automatic speech recognition
Daniel S Park, William Chan, Yu Zhang, Chung-Cheng Chiu, Barret Zoph, Ekin D Cubuk, and Quoc V Le · 2019
Cited alongside, same era.
Towards generalized speech enhancement with generative adversarial networks
Santiago Pascual, Joan Serrà, and Antonio Bonafonte · 2019
Cited alongside, same era.
wav2vec 2.0: A framework for self-supervised learning of speech representations
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli · 2020
Cited alongside, same era.
Real time speech enhancement in the waveform domain
Alexandre Defossez, Gabriel Synnaeve, and Yossi Adi · 2020
Cited alongside, same era.
Libri-light: A benchmark for asr with limited or no supervision
J. Kahn, M. Rivière, W. Zheng, E. Kharitonov, Q. Xu, P. E. Mazaré, J. Karadayi, V. Liptchinsky, R. Collobert, C. Fuegen, T. Likhomanenko, G. Synnaeve, A. Joulin, A. Mohamed, and E. Dupoux · 2020
Cited alongside, same era.
Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis
Jungil Kong, Jaehyeon Kim, and Jaekyoung Bae · 2020
Cited alongside, same era.
Voice separation with an unknown number of multiple speakers
Eliya Nachmani, Yossi Adi, and Lior Wolf · 2020
Cited alongside, same era.
Cascaded time+ time-frequency unet for speech enhancement: Jointly addressing clipping, codec distortions, and gaps
Arun Asokan Nair and Kazuhito Koishida · 2021
Later among the works it cites.
Speech resynthesis from discrete disentangled self-supervised representations
Adam Polyak, Yossi Adi, Jade Copet, Eugene Kharitonov, Kushal Lakhotia, Wei-Ning Hsu, Abdelrahman Mohamed, and Emmanuel Dupoux · 2021
Later among the works it cites.
High fidelity speech regeneration with application to speech enhancement
Adam Polyak, Lior Wolf, Yossi Adi, Ori Kabeli, and Yaniv Taigman · 2021
Later among the works it cites.
Hifi-gan-2: Studio-quality speech enhancement via generative adversarial networks conditioned on acoustic features
Jiaqi Su, Zeyu Jin, and Adam Finkelstein · 2021
Later among the works it cites.
Changhan Wang, Morgane Riviere, Ann Lee, Anne Wu, Chaitanya Talnikar, Daniel Haziza, Mary Williamson, Juan Pino, and Emmanuel Dupoux · 2021
Later among the works it cites.
Self-training and pre-training are complementary for speech recognition
Qiantong Xu, Alexei Baevski, Tatiana Likhomanenko, Paden Tomasello, Alexis Conneau, Ronan Collobert, Gabriel Synnaeve, and Michael Auli · 2021
Later among the works it cites.
Audio signal processing for telepresence based on wearable array in noisy and dynamic scenes
Hanan Beit-On, Moti Lugasi, Lior Madmoni, Anjali Menon, Anurag Kumar, Jacob Donley, Vladimir Tourbabin, and Boaz Rafaely · 2022
Closest in time.
More than words: In-the-wild visually-driven prosody for text-to-speech
Michael Hassid, Michelle Tadmor Ramanovich, Brendan Shillingford, Miaosen Wang, Ye Jia, and Tal Remez · 2022
Closest in time.
u-hubert: Unified mixed-modal speech pretraining and zero-shot transfer to unlabeled modality
Wei-Ning Hsu and Bowen Shi · 2022
Closest in time.
Svts: Scalable video-to-speech synthesis
Rodrigo Mira, Alexandros Haliassos, Stavros Petridis, Björn W Schuller, and Maja Pantic · 2022
Closest in time.
End-to-end video-to-speech synthesis using generative adversarial networks
Rodrigo Mira, Konstantinos Vougioukas, Pingchuan Ma, Stavros Petridis, Björn W Schuller, and Maja Pantic · 2022
Closest in time.
Universal speech enhancement with score-based diffusion
Joan Serrà, Santiago Pascual, Jordi Pons, R Oguz Araz, and Davide Scaini · 2022
Closest in time.
Learning audio-visual speech representation by masked multimodal cluster prediction
Bowen Shi, Wei-Ning Hsu, Kushal Lakhotia, and Abdelrahman Mohamed · 2022
Closest in time.
Audio-visual speech codecs: Rethinking audio-visual speech enhancement by re-synthesis
Karren Yang, Dejan Marković, Steven Krenn, Vasu Agrawal, and Alexander Richard · 2022
Closest in time.