Fetching the paper…
Reading the bibliography…
Although speech is a simple and effective way for humans to communicate with the outside world, a more realistic speech interaction contains multimodal information, e.g., vision, text.
Y.-A. Chung, C. Zhu, and M. Zeng, “SPLAT: Speech-language joint pre-training for spoken language understanding,” in Proc. Annu. Meeting Assoc. Comput. Linguistics , 2021, pp. 1897–1907
1907
Earlier work this paper cites.
W. H. Sumby and I. Pollack, “Visual contribution to speech intelligibility in noise,” J. Acoust. Soc. Amer. , vol. 26, no. 2, pp. 212–215, 1954
1954
Earlier work this paper cites.
S. Dupont and J. Luettin, “Audio-visual speech modeling for continuous speech recognition,” IEEE Trans. Multimedia , vol. 2, no. 3, pp. 141–151, 2000
2000
Earlier work this paper cites.
J.-S. Lee and C. H. Park, “Robust audio-visual speech recognition based on late integration,” IEEE Trans. Multimedia , vol. 10, no. 5, pp. 767–779, 2008
2008
Earlier work this paper cites.
D. E. King, “Dlib-ml: A machine learning toolkit,” J. Mach. Learn. Res. , vol. 10, pp. 1755–1758, 2009
2009
Earlier work this paper cites.
H. Li, Y. Wang, T. Mei, J. Wang, and S. Li, “Interactive multimodal visual search on mobile device,” IEEE Trans. Multimedia , vol. 15, no. 3, pp. 594–607, 2013
2013
Earlier work this paper cites.
W. Williams, N. Prasad, D. Mrva, T. Ash, and T. Robinson, “Scaling recurrent neural network language models,” in IEEE Int. Conf. Acoust., Speech, Signal Process. , 2015, pp. 5391–5395
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Proc. Adv. Neural Inf. Process. Syst. , vol. 30, pp. 6000–6010, 2017
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
S. Petridis, T. Stafylakis, P. Ma, F. Cai, G. Tzimiropoulos, and M. Pantic, “End-to-end audiovisual speech recognition,” in IEEE Int. Conf. Acoust., Speech, Signal Process. , 2018, pp. 6548–6552
2018
Earlier work this paper cites.
T. Afouras, J. S. Chung, A. Senior, O. Vinyals, and A. Zisserman, “Deep audio-visual speech recognition,” IEEE Trans. Pattern Anal. Mach. Intell. , pp. 1–1, 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
J. S. Chung, A. Nagrani, and A. Zisserman, “Voxceleb2: Deep speaker recognition,” in ISCA Interspeech , 2018, pp. 1086–1090
2018
Earlier work this paper cites.
F. Hernandez, V. Nguyen, S. Ghannay, N. Tomashenko, and Y. Estève, “Ted-lium 3: Twice as much data and corpus repartition for experiments on speaker adaptation,” in Speech and Computer , 2018, pp. 198–208
2018
Earlier work this paper cites.
T. Kudo and J. Richardson, “SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing,” in Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing: System Demonstrations , Nov. 2018, pp. 66–71
2018
Earlier work this paper cites.
S. Schneider, A. Baevski, R. Collobert, and M. Auli, “Wav2vec: Unsupervised Pre-Training for Speech Recognition,” in ISCA Interspeech , 2019, pp. 3465–3469
2019
Earlier work this paper cites.
Y.-A. Chung, W.-N. Hsu, H. Tang, and J. Glass, “An Unsupervised Autoregressive Model for Speech Representation Learning,” in ISCA Interspeech , 2019, pp. 146–150
2019
Earlier work this paper cites.
T. Makino, H. Liao, Y. Assael, B. Shillingford, B. Garcia, O. Braga, and O. Siohan, “Recurrent neural network transducer for audio-visual speech recognition,” in IEEE Autom. Speech Recognit. Understanding Workshop , 2019, pp. 905–912
2019
Earlier work this paper cites.
Y. Ren, Y. Ruan, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T.-Y. Liu, “Fastspeech: Fast, robust and controllable text to speech,” in Proc. Adv. Neural Inf. Process. Syst. , vol. 32, 2019
2019
Earlier work this paper cites.
X. Zhang, F. Cheng, and S. Wang, “Spatio-temporal fusion based convolutional sequence learning for lip reading,” in Proc. IEEE/CVF Int. Conf. Comput. Vis. , 2019, pp. 713–722
2019
Earlier work this paper cites.
B. Shillingford, Y. Assael, M. W. Hoffman, T. Paine, C. Hughes, U. Prabhu, H. Liao, H. Sak, K. Rao, L. Bennett, M. Mulville, M. Denil, B. Coppin, B. Laurie, A. Senior, and N. de Freitas, “Large-Scale Visual Speech Recognition,” in ISCA Interspeech , 2019, pp. 4135–4139
2019
Earlier work this paper cites.
M. Ott, S. Edunov, A. Baevski, A. Fan, S. Gross, N. Ng, D. Grangier, and M. Auli, “fairseq: A fast, extensible toolkit for sequence modeling,” in Proc. Annu. Meeting Assoc. Comput. Linguistics , 2019, pp. 48–53
2019
Earlier work this paper cites.
A. Baevski, Y. Zhou, A. Mohamed, and M. Auli, “Wav2vec 2.0: A framework for self-supervised learning of speech representations,” in Proc. Adv. Neural Inf. Process. Syst. , 2020, pp. 12 449–12 460
2020
Cited alongside, same era.
B. Xu, C. Lu, Y. Guo, and J. Wang, “Discriminative multi-modality speech recognition,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. , 2020, pp. 14 421–14 430
2020
Cited alongside, same era.
T. Afouras, J. S. Chung, and A. Zisserman, “Asr is all you need: Cross-modal distillation for lip reading,” in IEEE Int. Conf. Acoust., Speech, Signal Process. , 2020, pp. 2143–2147
2020
Cited alongside, same era.
Q. Zhang, H. Lu, H. Sak, A. Tripathi, E. McDermott, S. Koo, and S. Kumar, “Transformer transducer: A streamable speech recognition model with transformer encoders and rnn-t loss,” in IEEE Int. Conf. Acoust., Speech, Signal Process. , 2020, pp. 7829–7833
2020
Cited alongside, same era.
A. Baevski, W. Hsu, Q. Xu, A. Babu, J. Gu, and M. Auli, “data2vec: A general framework for self-supervised learning in speech, vision and language,” in Int. Conf. Machine Learning (ICML) , vol. 162, 2022, pp. 1298–1312
2022
Closest in time.
2022
Closest in time.
K. He, X. Chen, S. Xie, Y. Li, P. Dollár, and R. Girshick, “Masked autoencoders are scalable vision learners,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. , 2022, pp. 16 000–16 009
2022
Closest in time.
2022
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
B. Martinez, P. Ma, S. Petridis, and M. Pantic, “Lipreading using temporal convolutional networks,” in IEEE Int. Conf. Acoust., Speech, Signal Process. , 2020, pp. 6319–6323
2020
Cited alongside, same era.
Y. Zhao, R. Xu, X. Wang, P. Hou, H. Tang, and M. Song, “Hearing lips: Improving lip reading by distilling speech recognizers,” in Proc. AAAI Conf. Artif. Intell. , vol. 34, no. 04, 2020, pp. 6917–6924
2020
Cited alongside, same era.
2020
Cited alongside, same era.
H. Zhu, M.-D. Luo, R. Wang, A.-H. Zheng, and R. He, “Deep audio-visual learning: A survey,” Int. J. Autom. Comput. , vol. 18, no. 3, pp. 351–376, 2021
2021
Cited alongside, same era.
W.-N. Hsu, B. Bolte, Y.-H. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed, “Hubert: Self-supervised speech representation learning by masked prediction of hidden units,” IEEE/ACM Trans. Audio, Speech, Lang. Process. , vol. 29, pp. 3451–3460, 2021
2021
Cited alongside, same era.
A. T. Liu, S.-W. Li, and H.-y. Lee, “Tera: Self-supervised learning of transformer encoder representation for speech,” IEEE/ACM Trans. Audio, Speech, Lang. Process. , vol. 29, pp. 2351–2366, 2021
2021
Cited alongside, same era.
Y.-A. Chung, Y. Zhang, W. Han, C.-C. Chiu, J. Qin, R. Pang, and Y. Wu, “w2v-bert: Combining contrastive learning and masked language modeling for self-supervised speech pre-training,” in IEEE Autom. Speech Recognit. Understanding Workshop , 2021, pp. 244–250
2021
Cited alongside, same era.
P. Ma, S. Petridis, and M. Pantic, “End-to-end audio-visual speech recognition with conformers,” in IEEE Int. Conf. Acoust., Speech, Signal Process. , 2021, pp. 7613–7617
2021
Cited alongside, same era.
2022
Closest in time.
J. Ao, R. Wang, L. Zhou, C. Wang, S. Ren, Y. Wu, S. Liu, T. Ko, Q. Li, Y. Zhang, Z. Wei, Y. Qian, J. Li, and F. Wei, “SpeechT5: Unified-modal encoder-decoder pre-training for spoken language processing,” in Proc. Annu. Meeting Assoc. Comput. Linguistics , 2022, pp. 5723–5738
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
X. Pan, P. Chen, Y. Gong, H. Zhou, X. Wang, and Z. Lin, “Leveraging unimodal self-supervised learning for multimodal audio-visual speech recognition,” in Proc. Annu. Meeting Assoc. Comput. Linguistics , 2022, pp. 4491–4503
2022
Closest in time.
Z.-Q. Zhang, J. Zhang, J.-S. Zhang, M.-H. Wu, X. Fang, and L.-R. Dai, “Learning contextually fused audio-visual representations for audio-visual speech recognition,” in IEEE Int. Conf. Image Process. , 2022, pp. 1346–1350
2022
Closest in time.
B. Shi, W.-N. Hsu, and A. Mohamed, “Robust Self-Supervised Audio-Visual Speech Recognition,” in ISCA Interspeech , 2022, pp. 2118–2122
2022
Closest in time.
Y. Tang, H. Gong, N. Dong, C. Wang, W.-N. Hsu, J. Gu, A. Baevski, X. Li, A. Mohamed, M. Auli, and J. Pino, “Unified speech-text pre-training for speech translation and recognition,” in Proc. Annu. Meeting Assoc. Comput. Linguistics , 2022, pp. 1488–1499
2022
Closest in time.
W. Wang, S. Ren, Y. Qian, S. Liu, Y. Shi, Y. Qian, and M. Zeng, “Optimizing alignment of speech and language latent spaces for end-to-end speech recognition and understanding,” in IEEE Int. Conf. Acoust., Speech, Signal Process. , 2022, pp. 7802–7806
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
D. M. Chan, S. Ghosh, D. Chakrabarty, and B. Hoffmeister, “Multi-modal pre-training for automated speech recognition,” in IEEE Int. Conf. Acoust., Speech, Signal Process. , 2022, pp. 246–250
2022
Closest in time.
K. R. Prajwal, T. Afouras, and A. Zisserman, “Sub-word level lip reading with visual attention,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , June 2022, pp. 5162–5172
2022
Closest in time.
M. Kim, J. H. Yeo, and Y. M. Ro, “Distinguishing homophenes using multi-head visual-audio memory for lip reading,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 36, no. 1, 2022, pp. 1174–1182
2022
Closest in time.
P. Ma, S. Petridis, and M. Pantic, “Visual speech recognition for multiple languages in the wild,” Nature Machine Intelligence , pp. 1–10, 2022
2022
Closest in time.
Z. Chen, Y. Zhang, A. Rosenberg, B. Ramabhadran, P. J. Moreno, A. Bapna, and H. Zen, “MAESTRO: Matched Speech Text Representations through Modality Matching,” in ISCA Interspeech , 2022, pp. 4093–4097
2022
Closest in time.
2022
Closest in time.
B. Shi, W.-N. Hsu, and A. Mohamed, “Robust Self-Supervised Audio-Visual Speech Recognition,” in ISCA Interspeech , 2022, pp. 2118–2122
2022
Closest in time.