Fetching the paper…
Reading the bibliography…
Audio-based automatic speech recognition (ASR) degrades significantly in noisy environments and is particularly vulnerable to interfering speech, as the model cannot determine which speaker to transcribe.
W. H. Sumby and I. Pollack, “Visual contribution to speech intelligibility in noise,”
1954
Earlier work this paper cites.
H. Mcgurk and J. MacDoald, “Hearing lips and seeing voices,”
1976
Earlier work this paper cites.
G. Hinton, L. Deng, D. Yu, G. E. Dahl, A.-r. Mohamed, N. Jaitly, A. Senior, V. Vanhoucke, P. Nguyen, T. N. Sainath, and B. Kingsbury, “Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups,”
2012
Earlier work this paper cites.
D. Snyder, G. Chen, and D. Povey, “Musan: A music, speech, and noise corpus,”
2015
Earlier work this paper cites.
D. Amodei
2016
Earlier work this paper cites.
W. Xiong, J. Droppo, X. Huang, F. Seide, M. Seltzer, A. Stolcke, D. Yu, and G. Zweig, “Achieving human parity in conversational speech recognition,”
2016
Earlier work this paper cites.
A. Biswas, P. Sahu, and M. Chandra, “Multiple cameras audio visual speech recognition using active appearance model visual features in car environment,”
2016
Earlier work this paper cites.
E. Vincent, S. Watanabe, A. A. Nugraha, J. Barker, and R. Marxer, “An analysis of environment, microphone and data simulation mismatches in robust speech recognition,”
2017
Earlier work this paper cites.
T. Afouras, J. S. Chung, and A. Zisserman, “The conversation: Deep audio-visual speech enhancement,” in
2018
Earlier work this paper cites.
Y. Koguchi, K. Oharada, Y. Takagi, Y. Sawada, B. Shizuki, and S. Takahashi, “A mobile command input through vowel lip shape recognition,” in
2018
Cited alongside, same era.
T. Afouras, J. S. Chung, A. Senior, O. Vinyals, and A. Zisserman, “Deep audio-visual speech recognition,”
2018
Cited alongside, same era.
2018
Cited alongside, same era.
M. Caron, P. Bojanowski, A. Joulin, and M. Douze, “Deep clustering for unsupervised learning of visual features,” in
2018
Cited alongside, same era.
J. S. Chung, A. Nagrani, and A. Zisserman, “Voxceleb2: Deep speaker recognition,” in
2018
Cited alongside, same era.
2020
Later among the works it cites.
B. Xu, C. Lu, Y. Guo, and J. Wang, “Discriminative multi-modality speech recognition,” in
2020
Later among the works it cites.
T. S. Nguyen, S. Stueker, and A. H. Waibel, “Super-human performance in online low-latency recognition of conversational speech,” in
2021
Later among the works it cites.
K. Kinoshita, T. Ochiai, M. Delcroix, and T. Nakatani, “Improving noise robust automatic speech recognition with single-channel time-domain enhancement network,” in
2021
Later among the works it cites.
W.-N. Hsu, B. Bolte, Y.-H. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed, “Hubert: Self-supervised speech representation learning by masked prediction of hidden units,” in
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” in
2019
Cited alongside, same era.
T. Makino, H. Liao, Y. Assael, B. Shillingford, B. Garcia, O. Braga, and O. Siohan, “Recurrent neural network transducer for audio-visual speech recognition,” in
2019
Cited alongside, same era.
Z. Tüske, G. Saon, K. Audhkhasi, and B. Kingsbury, “Single headed attention based sequence-to-sequence model for state-of-the-art results on switchboard-300,” in
2020
Cited alongside, same era.
2021
Later among the works it cites.
2021
Later among the works it cites.
P. Ma, S. Petridis, and M. Pantic, “End-to-end audio-visual speech recognition with conformers,” in
2021
Later among the works it cites.
B. Shi, W.-N. Hsu, K. Lakhotia, and A. Mohamed, “Learning audio-visual speech representation by masked multimodal cluster prediction,” 2022
2022
Closest in time.