Fetching the paper…
Reading the bibliography…
While audio-visual speech models can yield superior performance and robustness compared to audio-only models, their development and adoption are hindered by the lack of labeled and unlabeled audio-visual data and the cost to deploy one model per modality.
Robust speech recognition using the modulation spectrogram
B. E. Kingsbury, N. Morgan, and S. Greenberg · 1998
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu · 2002
Earlier work this paper cites.
Visualizing data using t-sne
L. Van der Maaten and G. Hinton · 2008
Earlier work this paper cites.
On the phonetic information in ultrasonic microphone signals
K. Livescu, B. Zhu, and J. R. Glass · 2009
Earlier work this paper cites.
Development of a silent speech interface driven by ultrasound and optical images of the tongue and lips
T. Hueber, E.-L. Benaroya, G. Chollet, B. Denby, G. Dreyfus, and M. Stone · 2010
Earlier work this paper cites.
Announcing the electromagnetic articulography (day 1) subset of the mngu0 articulatory corpus
K. Richmond, P. Hoole, and S. King · 2011
Earlier work this paper cites.
Japanese and korean voice search
M. Schuster and K. Nakajima · 2012
Earlier work this paper cites.
Speech feature denoising and dereverberation via deep autoencoders for noisy reverberant speech recognition
X. Feng, Y. Zhang, and J. Glass · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2015
Earlier work this paper cites.
Moddrop: Adaptive multi-modal gesture recognition
N. Neverova, C. Wolf, G. Taylor, and F. Nebout · 2015
Earlier work this paper cites.
Librispeech: an asr corpus based on public domain audio books
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur · 2015
Earlier work this paper cites.
Musan: A music, speech, and noise corpus
D. Snyder, G. Chen, and D. Povey · 2015
Earlier work this paper cites.
Lip reading in the wild
J. S. Chung and A. Zisserman · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Google’s multilingual neural machine translation system: Enabling zero-shot translation
M. Johnson, M. Schuster, Q. V. Le, M. Krikun, Y. Wu, Z. Chen, N. Thorat, F. Viégas, M. Wattenberg, G. Corrado, et al · 2017
Earlier work this paper cites.
End-to-end speech recognition with auditory attention for multi-microphone distance speech recognition
S. Kim, I. R. Lane, S. Kim, and I. Lane · 2017
Earlier work this paper cites.
Voxceleb: a large-scale speaker identification dataset
A. Nagrani, J. S. Chung, and A. Zisserman · 2017
Earlier work this paper cites.
A. Vaswani, N. M. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
LRS3-TED: a large-scale dataset for visual speech recognition, 2018
T. Afouras, J. S. Chung, and A. Zisserman · 2018
Earlier work this paper cites.
Voxceleb2: Deep speaker recognition
J. S. Chung, A. Nagrani, and A. Zisserman · 2018
Earlier work this paper cites.
Investigating objective intelligibility in real-time emg-to-speech conversion
L. Diener and T. Schultz · 2018
Cited alongside, same era.
Ted-lium 3: twice as much data and corpus repartition for experiments on speaker adaptation
F. Hernandez, V. Nguyen, S. Ghannay, N. A. Tomashenko, and Y. Estève · 2018
Cited alongside, same era.
Subword regularization: Improving neural network translation models with multiple subword candidates
T. Kudo · 2018
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2019
Cited alongside, same era.
Cross-lingual language model pretraining
G. Lample and A. Conneau · 2019
Cited alongside, same era.
Mixout: Effective regularization to finetune large-scale pretrained language models
Slam: A unified encoder for speech and language modeling via speech-text joint pre-training
A. Bapna, Y.-a. Chung, N. Wu, A. Gulati, Y. Jia, J. H. Clark, M. Johnson, J. Riesa, A. Conneau, and Y. Zhang · 2021
Later among the works it cites.
Speechstew: Simply mix all available speech recognition data to train one large neural network
W. Chan, D. Park, C. Lee, Y. Zhang, Q. Le, and M. Norouzi · 2021
Later among the works it cites.
w2v-bert: Combining contrastive learning and masked language modeling for self-supervised speech pre-training
Y.-A. Chung, Y. Zhang, W. Han, C.-C. Chiu, J. Qin, R. Pang, and Y. Wu · 2021
Later among the works it cites.
Far-field automatic speech recognition
R. Haeb-Umbach, J. Heymann, L. Drude, S. Watanabe, M. Delcroix, T. N. P. University, H. Germany, A. J. Aachen, J. H. University, Baltimore., Usa, N. C. S. Laboratories, Kyoto, and Japan · 2021
Later among the works it cites.
Unit: Multimodal multitask learning with a unified transformer
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
C. Lee, K. Cho, and W. Kang · 2019
Cited alongside, same era.
Recurrent neural network transducer for audio-visual speech recognition
T. Makino, H. Liao, Y. Assael, B. Shillingford, B. Garcia, O. Braga, and O. Siohan · 2019
Cited alongside, same era.
Transformers without tears: Improving the normalization of self-attention
T. Q. Nguyen and J. Salazar · 2019
Cited alongside, same era.
How multilingual is multilingual bert?
T. Pires, E. Schlinger, and D. Garrette · 2019
Cited alongside, same era.
Eleatt-rnn: Adding attentiveness to neurons in recurrent neural networks
P. Zhang, J. Xue, C. Lan, W. Zeng, Z. Gao, and N. Zheng · 2019
Cited alongside, same era.
ASR is all you need: Cross-modal distillation for lip reading
T. Afouras, J. S. Chung, and A. Zisserman · 2020
Cited alongside, same era.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Cited alongside, same era.
R. Hu and A. Singh · 2021
Later among the works it cites.
Perceiver: General perception with iterative attention
A. Jaegle, F. Gimeno, A. Brock, O. Vinyals, A. Zisserman, and J. Carreira · 2021
Later among the works it cites.
Scaling up visual and vision-language representation learning with noisy text supervision
C. Jia, Y. Yang, Y. Xia, Y.-T. Chen, Z. Parekh, H. Pham, Q. V. Le, Y.-H. Sung, Z. Li, and T. Duerig · 2021
Later among the works it cites.
Swin transformer: Hierarchical vision transformer using shifted windows
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo · 2021
Later among the works it cites.
End-to-end audio-visual speech recognition with conformers
P. Ma, S. Petridis, and M. Pantic · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever · 2021
Later among the works it cites.
Meshtalk: 3d face animation from speech using cross-modality disentanglement
A. Richard, M. Zollhöfer, Y. Wen, F. de la Torre, and Y. Sheikh · 2021
Later among the works it cites.
A multi-view approach to audio-visual speaker verification
L. Sari, K. Singh, J. Zhou, L. Torresani, N. Singhal, and Y. Saraf · 2021
Later among the works it cites.
Audio-visual speech recognition is worth 32 × \times 32 × \times 8 voxels
D. Serdyuk, O. Braga, and O. Siohan · 2021
Later among the works it cites.
Flava: A foundational language and vision alignment model
A. Singh, R. Hu, V. Goswami, G. Couairon, W. Galuba, M. Rohrbach, and D. Kiela · 2021
Later among the works it cites.
Multilingual unsupervised neural machine translation with denoising adapters
A. Üstün, A. Bérard, L. Besacier, and M. Gallé · 2021
Later among the works it cites.
Nlp from scratch without large-scale pretraining: A simple and efficient framework
X. Yao, Y. Zheng, X. Yang, and Z. Yang · 2021
Later among the works it cites.
Cm3: A causal masked multimodal model of the internet
A. Aghajanyan, B. Huang, C. Ross, V. Karpukhin, H. Xu, N. Goyal, D. Okhonko, M. Joshi, G. Ghosh, M. Lewis, et al · 2022
Closest in time.
Data2vec: A general framework for self-supervised learning in speech, vision and language
A. Baevski, W.-N. Hsu, Q. Xu, A. Babu, J. Gu, and M. Auli · 2022
Closest in time.
Omnivore: A Single Model for Many Visual Modalities
R. Girdhar, M. Singh, N. Ravi, L. van der Maaten, A. Joulin, and I. Misra · 2022
Closest in time.
Visual speech recognition for multiple languages in the wild
P. Ma, S. Petridis, and M. Pantic · 2022
Closest in time.