Fetching the paper…
Reading the bibliography…
Self-supervision has shown great potential for audio-visual speech recognition by vastly reducing the amount of labeled data required to build good systems.
“The origin of speech,”
Charles F Hockett and Charles D Hockett, · 1960
Earlier work this paper cites.
“The use of auditory and visual information during phonetic processing: implications for theories of speech perception. campbell r, dodd b, burnham d, editors. hearing by eye ii: advances in the psychology of speechreading and auditory–visual speech,” 1998
KP Green, · 1998
Earlier work this paper cites.
“Speech perception,”
Randy L Diehl, Andrew J Lotto, Lori L Holt, et al., · 2004
Earlier work this paper cites.
“Dlib-ml: A machine learning toolkit,”
Davis E King, · 2009
Earlier work this paper cites.
“Batch normalization: Accelerating deep network training by reducing internal covariate shift,”
Sergey Ioffe and Christian Szegedy, · 2015
Earlier work this paper cites.
“Delving deep into rectifiers: Surpassing human-level performance on imagenet classification,”
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun, · 2015
Earlier work this paper cites.
“Deep residual learning for image recognition,”
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun, · 2016
Earlier work this paper cites.
“End-to-end attention-based large vocabulary speech recognition,”
Dzmitry Bahdanau, Jan Chorowski, Dmitriy Serdyuk, Philemon Brakel, and Yoshua Bengio, · 2016
Earlier work this paper cites.
“Neural discrete representation learning,”
Aaron Van Den Oord, Oriol Vinyals, et al., · 2017
Earlier work this paper cites.
“Hybrid ctc/attention architecture for end-to-end speech recognition,”
Shinji Watanabe, Takaaki Hori, Suyoun Kim, John R Hershey, and Tomoki Hayashi, · 2017
Earlier work this paper cites.
“Deep audio-visual speech recognition,”
Triantafyllos Afouras, Joon Son Chung, Andrew Senior, Oriol Vinyals, and Andrew Zisserman, · 2018
Earlier work this paper cites.
“Audio-visual speech recognition with a hybrid ctc/attention architecture,”
Stavros Petridis, Themos Stafylakis, Pingchuan Ma, Georgios Tzimiropoulos, and Maja Pantic, · 2018
Earlier work this paper cites.
“Lrs3-ted: a large-scale dataset for visual speech recognition,”
Triantafyllos Afouras, Joon Son Chung, and Andrew Zisserman, · 2018
Earlier work this paper cites.
“Voxceleb2: Deep speaker recognition,”
Joon Son Chung, Arsha Nagrani, and Andrew Zisserman, · 2018
Earlier work this paper cites.
“Subword regularization: Improving neural network translation models with multiple subword candidates,”
Taku Kudo, · 2018
Earlier work this paper cites.
“Effectiveness of self-supervised pre-training for speech recognition,”
Alexei Baevski, Michael Auli, and Abdelrahman Mohamed, · 2019
Earlier work this paper cites.
“Large-scale visual speech recognition,”
Brendan Shillingford, Yannis Assael, Matthew W Hoffman, Thomas Paine, Cían Hughes, Utsav Prabhu, Hank Liao, Hasim Sak, Kanishka Rao, Lorrayne Bennett, et al., · 2019
Cited alongside, same era.
“Recurrent neural network transducer for audio-visual speech recognition,”
Takaki Makino, Hank Liao, Yannis Assael, Brendan Shillingford, Basilio Garcia, Otavio Braga, and Olivier Siohan, · 2019
Cited alongside, same era.
“fairseq: A fast, extensible toolkit for sequence modeling,”
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli, · 2019
Cited alongside, same era.
“BERT: Pre-training of deep bidirectional transformers for language understanding,”
Jacob Devlin, Ming-Wei Chang, et al., · 2019
Cited alongside, same era.
“Generative pre-training for speech with autoregressive predictive coding,”
Yu-An Chung and James Glass, · 2020
Cited alongside, same era.
“Learning audio-visual speech representation by masked multimodal cluster prediction,”
Bowen Shi, Wei-Ning Hsu, Kushal Lakhotia, and Abdelrahman Mohamed, · 2022
Later among the works it cites.
“Robust self-supervised audio-visual speech recognition,”
Bowen Shi, Wei-Ning Hsu, and Abdelrahman Mohamed, · 2022
Later among the works it cites.
“Transformer-based video front-ends for audio-visual speech recognition,”
Dmitriy Serdyuk, Otavio Braga, and Olivier Siohan, · 2022
Later among the works it cites.
“Data2vec: A general framework for self-supervised learning in speech, vision and language,”
Alexei Baevski, Wei-Ning Hsu, Qiantong Xu, Arun Babu, Jiatao Gu, and Michael Auli, · 2022
Later among the works it cites.
“u-hubert: Unified mixed-modal speech pretraining and zero-shot transfer to unlabeled modality,”
Wei-Ning Hsu and Bowen Shi, · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yu-An Chung, Hao Tang, and James Glass, · 2020
Cited alongside, same era.
“Deep contextualized acoustic representations for semi-supervised speech recognition,”
Shaoshi Ling, Yuzong Liu, Julian Salazar, and Katrin Kirchhoff, · 2020
Cited alongside, same era.
“Discriminative multi-modality speech recognition,”
Bo Xu, Cheng Lu, Yandong Guo, and Jacob Wang, · 2020
Cited alongside, same era.
“Asr is all you need: Cross-modal distillation for lip reading,”
Triantafyllos Afouras, Joon Son Chung, and Andrew Zisserman, · 2020
Cited alongside, same era.
“wav2vec 2.0: A framework for self-supervised learning of speech representations,”
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli, · 2020
Cited alongside, same era.
“Phonetically motivated self-supervised speech representation learning.,”
Xianghu Yue and Haizhou Li, · 2021
Cited alongside, same era.
“Non-autoregressive predictive coding for learning speech representations from local dependencies,”
Alexander H Liu, Yu-An Chung, and James Glass, · 2021
Cited alongside, same era.
Sanyuan Chen, Chengyi Wang, Zhengyang Chen, Yu Wu, Shujie Liu, Zhuo Chen, Jinyu Li, Naoyuki Kanda, Takuya Yoshioka, Xiong Xiao, Jian Wu, Long Zhou, Shuo Ren, Yanmin Qian, Yao Qian, Jian Wu, Michael Zeng, Xiangzhan Yu, and Furu Wei, · 2022
Later among the works it cites.
“Visual speech recognition for multiple languages in the wild,”
Pingchuan Ma, Stavros Petridis, and Maja Pantic, · 2022
Later among the works it cites.
“Sub-word level lip reading with visual attention,”
KR Prajwal, Triantafyllos Afouras, and Andrew Zisserman, · 2022
Later among the works it cites.
“Deep Neural Convolutive Matrix Factorization for Articulatory Representation Decomposition,”
Jiachen Lian, Alan W Black, Louis Goldstein, and Gopala Krishna Anumanchipalli, · 2022
Later among the works it cites.
“Jointly learning visual and auditory speech representations from raw data,”
Alexandros Haliassos, Pingchuan Ma, Rodrigo Mira, Stavros Petridis, and Maja Pantic, · 2023
Closest in time.
“Efficient self-supervised learning with contextualized target representations for vision, speech and language,”
Alexei Baevski, Arun Babu, Wei-Ning Hsu, and Michael Auli, · 2023
Closest in time.
“VatLM: Visual-audio-text pre-training with unified masked prediction for speech representation learning,”
Qiushi Zhu, Long Zhou, et al., · 2023
Closest in time.
“Articulatory representation learning via joint factor analysis and neural matrix factorization,”
Jiachen Lian, Alan W Black, Yijing Lu, Louis Goldstein, Shinji Watanabe, and Gopala K Anumanchipalli, · 2023
Closest in time.
“Speaker-independent acoustic-to-articulatory speech inversion,”
Peter Wu, Li-Wei Chen, Cheol Jun Cho, Shinji Watanabe, Louis Goldstein, Alan W Black, and Gopala K. Anumanchipalli, · 2023
Closest in time.
“Deep Speech Synthesis from MRI-Based Articulatory Representations,”
Peter Wu, Tingle Li, Yijing Lu, Yubin Zhang, Jiachen Lian, Alan W Black, Louis Goldstein, Shinji Watanabe, and Gopala K. Anumanchipalli, · 2023
Closest in time.