Fetching the paper…
Reading the bibliography…
Visual Speech Recognition (VSR) aims to infer speech into text depending on lip movements alone.
J. Neto, L. Almeida, M. Hochberg, C. Martins, L. Nunes, S. Renals, and T. Robinson, “Speaker-adaptation for hybrid hmm-ann continuous speech recognition system,” in 4th European Conference on Speech Communication and Technology (Eurospeech 1995) . ISCA, 1995, p. 2171
1995
Earlier work this paper cites.
T. Anastasakos, J. McDonough, and J. Makhoul, “Speaker adaptive training: A maximum likelihood approach to speaker normalization,” in 1997 IEEE International Conference on Acoustics, Speech, and Signal Processing , vol. 2. IEEE, 1997, pp. 1043–1046
1997
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural computation , vol. 9, no. 8, pp. 1735–1780, 1997
1997
Earlier work this paper cites.
R. A. Gopinath, “Maximum likelihood modeling with gaussian distributions for classification,” in Proceedings of the 1998 IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP’98 (Cat. No. 98CH36181) , vol. 2. IEEE, 1998, pp. 661–664
1998
Earlier work this paper cites.
I. Matthews, T. F. Cootes, J. A. Bangham, S. Cox, and R. Harvey, “Extraction of visual features for lipreading,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 24, no. 2, pp. 198–213, 2002
2002
Earlier work this paper cites.
K. Saenko, T. Darrell, and J. R. Glass, “Articulatory features for robust visual speech recognition,” in Proceedings of the 6th international conference on Multimodal interfaces , 2004, pp. 152–158
2004
Earlier work this paper cites.
I. Cohen, F. G. Cozman, N. Sebe, M. C. Cirelo, and T. S. Huang, “Semisupervised learning of classifiers: Theory, algorithms, and their application to human-computer interaction,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 26, no. 12, pp. 1553–1566, 2004
2004
Earlier work this paper cites.
K. Saenko, K. Livescu, M. Siracusa, K. Wilson, J. Glass, and T. Darrell, “Visual speech recognition with loosely synchronized feature streams,” in Tenth IEEE International Conference on Computer Vision (ICCV’05) Volume 1 , vol. 2. IEEE, 2005, pp. 1424–1431
2005
Earlier work this paper cites.
K. Saenko, K. Livescu, J. Glass, and T. Darrell, “Production domain modeling of pronunciation for visual speech recognition,” in Proceedings.(ICASSP’05). IEEE International Conference on Acoustics, Speech, and Signal Processing, 2005. , vol. 5. IEEE, 2005, pp. v–473
2005
Earlier work this paper cites.
X. He, S. Yan, Y. Hu, P. Niyogi, and H.-J. Zhang, “Face recognition using laplacianfaces,” IEEE transactions on pattern analysis and machine intelligence , vol. 27, no. 3, pp. 328–340, 2005
2005
Earlier work this paper cites.
M. Cooke, J. Barker, S. Cunningham, and X. Shao, “An audio-visual corpus for speech perception and automatic speech recognition,” The Journal of the Acoustical Society of America , vol. 120, no. 5, pp. 2421–2424, 2006
2006
Earlier work this paper cites.
S. Lafon, Y. Keller, and R. R. Coifman, “Data fusion and multicue data matching by diffusion maps,” IEEE Transactions on pattern analysis and machine intelligence , vol. 28, no. 11, pp. 1784–1797, 2006
2006
Earlier work this paper cites.
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, “Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,” in Proceedings of the 23rd international conference on Machine learning , 2006, pp. 369–376
2006
Earlier work this paper cites.
X. Li and J. Bilmes, “Regularized adaptation of discriminative classifiers,” in 2006 IEEE International Conference on Acoustics Speech and Signal Processing Proceedings , vol. 1. IEEE, 2006, pp. I–I
2006
Earlier work this paper cites.
K. Saenko, K. Livescu, J. Glass, and T. Darrell, “Multistream articulatory feature-based models for visual speech recognition,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 31, no. 9, pp. 1700–1707, 2009
2009
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in 2009 IEEE conference on computer vision and pattern recognition . Ieee, 2009, pp. 248–255
2009
Earlier work this paper cites.
B. Li and K. C. Sim, “Comparison of discriminative input and output transformations for speaker adaptation in the hybrid nn/hmm systems,” in Eleventh Annual Conference of the International Speech Communication Association , 2010
2010
Earlier work this paper cites.
N. Dehak, P. J. Kenny, R. Dehak, P. Dumouchel, and P. Ouellet, “Front-end factor analysis for speaker verification,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 19, no. 4, pp. 788–798, 2010
2010
Earlier work this paper cites.
F. Seide, G. Li, X. Chen, and D. Yu, “Feature engineering in context-dependent deep neural networks for conversational speech transcription,” in 2011 IEEE Workshop on Automatic Speech Recognition & Understanding . IEEE, 2011, pp. 24–29
2011
Earlier work this paper cites.
G. Hinton, L. Deng, D. Yu, G. E. Dahl, A.-r. Mohamed, N. Jaitly, A. Senior, V. Vanhoucke, P. Nguyen, T. N. Sainath et al. , “Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups,” IEEE Signal processing magazine , vol. 29, no. 6, pp. 82–97, 2012
2012
Earlier work this paper cites.
J. R. Barr, K. W. Bowyer, P. J. Flynn, and S. Biswas, “Face recognition from video: A review,” International journal of pattern recognition and artificial intelligence , vol. 26, no. 05, p. 1266002, 2012
2012
Earlier work this paper cites.
Z. Zhou, X. Hong, G. Zhao, and M. Pietikäinen, “A compact representation of visual speech data using latent variables,” IEEE transactions on pattern analysis and machine intelligence , vol. 36, no. 1, pp. 1–1, 2013
2013
Earlier work this paper cites.
O. Abdel-Hamid and H. Jiang, “Fast speaker adaptation of hybrid nn/hmm model for speech recognition based on discriminative learning of speaker code,” in 2013 IEEE International Conference on Acoustics, Speech and Signal Processing . IEEE, 2013, pp. 7942–7946
2013
Earlier work this paper cites.
O. Abdel-Hamid and H. Jiang, “Rapid and effective speaker adaptation of convolutional neural network based models for speech recognition.” in Proc. of Interspeech , 2013, pp. 1248–1252
2013
Earlier work this paper cites.
H. Liao, E. McDermott, and A. Senior, “Large scale deep neural network acoustic modeling with semi-supervised training data for youtube video transcription,” in 2013 IEEE Workshop on Automatic Speech Recognition and Understanding . IEEE, 2013, pp. 368–373
2013
Earlier work this paper cites.
D. Yu, K. Yao, H. Su, G. Li, and F. Seide, “Kl-divergence regularized deep neural network adaptation for improved large vocabulary speech recognition,” in 2013 IEEE International Conference on Acoustics, Speech and Signal Processing . IEEE, 2013, pp. 7893–7897
2013
Earlier work this paper cites.
Y. Miao, H. Zhang, and F. Metze, “Towards speaker adaptive training of deep neural network acoustic models,” in Fifteenth annual conference of the international speech communication association , 2014
2014
Earlier work this paper cites.
S. Xue, O. Abdel-Hamid, H. Jiang, L. Dai, and Q. Liu, “Fast adaptation of deep neural network based on discriminant codes for speech recognition,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 22, no. 12, pp. 1713–1725, 2014
2014
Earlier work this paper cites.
K. Simonyan and A. Zisserman, “Two-stream convolutional networks for action recognition in videos,” Advances in neural information processing systems , vol. 27, 2014
2014
Earlier work this paper cites.
I. Sutskever, O. Vinyals, and Q. V. Le, “Sequence to sequence learning with neural networks,” Advances in neural information processing systems , vol. 27, 2014
2014
Earlier work this paper cites.
P. Swietojanski and S. Renals, “Learning hidden unit contributions for unsupervised speaker adaptation of neural network acoustic models,” in 2014 IEEE Spoken Language Technology Workshop (SLT) . IEEE, 2014, pp. 171–176
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
Y. Miao, H. Zhang, and F. Metze, “Speaker adaptive training of deep neural network acoustic models using i-vectors,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 23, no. 11, pp. 1938–1949, 2015
2015
Earlier work this paper cites.
T. N. Sainath, B. Kingsbury, G. Saon, H. Soltau, A.-r. Mohamed, G. Dahl, and B. Ramabhadran, “Deep convolutional neural networks for large-scale speech tasks,” Neural networks , vol. 64, pp. 39–48, 2015
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
I. J. Goodfellow, J. Shlens, and C. Szegedy, “Explaining and harnessing adversarial examples,” in International Conference on Learning Representations , 2015
2015
Earlier work this paper cites.
K. Fragkiadaki, S. Levine, P. Felsen, and J. Malik, “Recurrent network models for human dynamics,” in Proceedings of the IEEE international conference on computer vision , 2015, pp. 4346–4354
2015
Cited alongside, same era.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in International Conference on Learning Representations , 2015
2015
Cited alongside, same era.
Y. Ganin and V. Lempitsky, “Unsupervised domain adaptation by backpropagation,” in International conference on machine learning . PMLR, 2015, pp. 1180–1189
2015
Cited alongside, same era.
2016
Cited alongside, same era.
J. S. Chung and A. Zisserman, “Lip reading in the wild,” in Asian conference on computer vision . Springer, 2016, pp. 87–103
M. Kim, J. Hong, and Y. M. Ro, “Lip to speech synthesis with visual context attentional gan,” Advances in Neural Information Processing Systems , vol. 34, pp. 2758–2770, 2021
2021
Later among the works it cites.
Q. Zhang, S. Wang, and G. Chen, “Speaker-independent lipreading by disentangled representation learning,” in 2021 IEEE International Conference on Image Processing (ICIP) . IEEE, 2021, pp. 2493–2497
2021
Later among the works it cites.
L. Chen, Y. Fan, and Y. Ye, “Adversarial reprogramming of pretrained neural networks for fraud detection,” in Proceedings of the 30th ACM International Conference on Information & Knowledge Management , 2021, pp. 2935–2939
2021
Later among the works it cites.
X. L. Li and P. Liang, “Prefix-tuning: Optimizing continuous prompts for generation,” in Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers) , 2021, pp. 4582–4597
2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2016
Cited alongside, same era.
J. S. Chung and A. Zisserman, “Out of time: automated lip sync in the wild,” in Asian conference on computer vision . Springer, 2016, pp. 251–263
2016
Cited alongside, same era.
I. Almajai, S. Cox, R. Harvey, and Y. Lan, “Improved speaker independent lip reading using speaker adaptive training and deep neural networks,” in 2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2016, pp. 2722–2726
2016
Cited alongside, same era.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 770–778
2016
Cited alongside, same era.
T. Stafylakis and G. Tzimiropoulos, “Combining residual networks with lstms for lipreading,” in Proc. of Interspeech , 2017
2017
Cited alongside, same era.
J. S. Chung, A. Senior, O. Vinyals, and A. Zisserman, “Lip reading sentences in the wild,” in 2017 IEEE conference on computer vision and pattern recognition (CVPR) . IEEE, 2017, pp. 3444–3453
2017
Cited alongside, same era.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems , vol. 30, 2017
2017
Cited alongside, same era.
S. Petridis, T. Stafylakis, P. Ma, F. Cai, G. Tzimiropoulos, and M. Pantic, “End-to-end audiovisual speech recognition,” in 2018 IEEE international conference on acoustics, speech and signal processing (ICASSP) . IEEE, 2018, pp. 6548–6552
2018
Cited alongside, same era.
Later among the works it cites.
B. Lester, R. Al-Rfou, and N. Constant, “The power of scale for parameter-efficient prompt tuning,” in Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing , 2021, pp. 3045–3059
2021
Later among the works it cites.
P. Ma, S. Petridis, and M. Pantic, “End-to-end audio-visual speech recognition with conformers,” in ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2021, pp. 7613–7617
2021
Later among the works it cites.
M. Kim, J. Hong, S. J. Park, and Y. M. Ro, “Multi-modality associative bridging through memory: Speech sound recollected from face video,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 296–306
2021
Later among the works it cites.
S. Ren, Y. Du, J. Lv, G. Han, and S. He, “Learning from the master: Distilling cross-modal advanced knowledge for lip reading,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2021, pp. 13 325–13 333
2021
Later among the works it cites.
P. Ma, B. Martinez, S. Petridis, and M. Pantic, “Towards practical lipreading with distilled and efficient models,” in ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2021, pp. 7608–7612
2021
Later among the works it cites.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al. , “Learning transferable visual models from natural language supervision,” in International Conference on Machine Learning . PMLR, 2021, pp. 8748–8763
2021
Later among the works it cites.
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly et al. , “An image is worth 16x16 words: Transformers for image recognition at scale,” in International Conference on Learning Representations , 2021
2021
Later among the works it cites.
M. Kim, J. H. Yeo, and Y. M. Ro, “Distinguishing homophenes using multi-head visual-audio memory for lip reading,” in Proceedings of the 36th AAAI Conference on Artificial Intelligence, Vancouver, BC, Canada , vol. 22, 2022
2022
Later among the works it cites.
P. Ma, S. Petridis, and M. Pantic, “Visual speech recognition for multiple languages in the wild,” Nature Machine Intelligence , vol. 4, no. 11, pp. 930–939, 2022
2022
Later among the works it cites.
B. Shi, W.-N. Hsu, K. Lakhotia, and A. Mohamed, “Learning audio-visual speech representation by masked multimodal cluster prediction,” in International Conference on Learning Representations , 2022
2022
Later among the works it cites.
M. Kim, H. Kim, and Y. M. Ro, “Speaker-adaptive lip reading with user-dependent padding,” in European Conference on Computer Vision . Springer, 2022, pp. 576–593
2022
Later among the works it cites.
2022
Later among the works it cites.
P. Neekhara, S. Hussain, J. Du, S. Dubnov, F. Koushanfar, and J. McAuley, “Cross-modal adversarial reprogramming,” in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision , 2022, pp. 2427–2435
2022
Later among the works it cites.
Y. Yu, H. J. Lee, H. Lee, and Y. M. Ro, “Defending person detection against adversarial patch attack by using universal defensive frame,” IEEE Transactions on Image Processing , vol. 31, pp. 6976–6990, 2022
2022
Later among the works it cites.
K. Zhou, J. Yang, C. C. Loy, and Z. Liu, “Conditional prompt learning for vision-language models,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 16 816–16 825
2022
Later among the works it cites.
X. Liu, K. Ji, Y. Fu, W. Tam, Z. Du, Z. Yang, and J. Tang, “P-tuning: Prompt tuning can be comparable to fine-tuning across scales and tasks,” in Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers) . Association for Computational Linguistics, May 2022, pp. 61–68
2022
Later among the works it cites.
M. Jia, L. Tang, B.-C. Chen, C. Cardie, S. Belongie, B. Hariharan, and S.-N. Lim, “Visual prompt tuning,” in European Conference on Computer Vision . Springer, 2022, pp. 709–727
2022
Later among the works it cites.
M. Xu, Y. Shen, S. Zhang, Y. Lu, D. Zhao, J. Tenenbaum, and C. Gan, “Prompting decision transformer for few-shot policy generalization,” in International Conference on Machine Learning . PMLR, 2022, pp. 24 631–24 645
2022
Later among the works it cites.
K. Zhou, J. Yang, C. C. Loy, and Z. Liu, “Learning to prompt for vision-language models,” International Journal of Computer Vision , vol. 130, no. 9, pp. 2337–2348, 2022
2022
Later among the works it cites.
K. Prajwal, T. Afouras, and A. Zisserman, “Sub-word level lip reading with visual attention,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 5162–5172
2022
Later among the works it cites.
M. Li, L. Chen, Y. Duan, Z. Hu, J. Feng, J. Zhou, and J. Lu, “Bridge-prompt: Towards ordinal action understanding in instructional videos,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 19 880–19 889
2022
Later among the works it cites.
2022
Later among the works it cites.
J. Hong, M. Kim, D. Yoo, and Y. M. Ro, “Visual context-driven audio feature enhancement for robust end-to-end audio-visual speech recognition,” in Proc. of Interspeech , 2022
2022
Later among the works it cites.
J. He, C. Zhou, X. Ma, T. Berg-Kirkpatrick, and G. Neubig, “Towards a unified view of parameter-efficient transfer learning,” in International Conference on Learning Representations , 2022
2022
Later among the works it cites.
E. J. Hu, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, W. Chen et al. , “Lora: Low-rank adaptation of large language models,” in International Conference on Learning Representations , 2022
2022
Later among the works it cites.
Y. Zheng, X. Feng, Z. Xia, X. Jiang, A. Demontis, M. Pintor, B. Biggio, and F. Roli, “Why adversarial reprogramming works, when it fails, and how to tell the difference,” Information Sciences , vol. 632, pp. 130–143, 2023
2023
Closest in time.
P. Liu, W. Yuan, J. Fu, Z. Jiang, H. Hayashi, and G. Neubig, “Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing,” ACM Computing Surveys , vol. 55, no. 9, pp. 1–35, 2023
2023
Closest in time.
J. S. Smith, L. Karlinsky, V. Gutta, P. Cascante-Bonilla, D. Kim, A. Arbelle, R. Panda, R. Feris, and Z. Kira, “Coda-prompt: Continual decomposed attention-based prompting for rehearsal-free continual learning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 11 909–11 919
2023
Closest in time.
Y. Gao, X. Shi, Y. Zhu, H. Wang, Z. Tang, X. Zhou, M. Li, and D. N. Metaxas, “Visual prompt tuning for test-time domain adaptation,” in International Conference on Learning Representations , 2023
2023
Closest in time.
J. Wu, Y. Gaur, Z. Chen, L. Zhou, Y. Zhu, T. Wang, J. Li, S. Liu, B. Ren, L. Liu et al. , “On decoder-only architecture for speech-to-text and large language model integration,” in 2023 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) . IEEE, 2023, pp. 1–8
2023
Closest in time.
2024
Closest in time.
Y. Fathullah, C. Wu, E. Lakomkin, J. Jia, Y. Shangguan, K. Li, J. Guo, W. Xiong, J. Mahadeokar, O. Kalinli et al. , “Prompting large language models with speech recognition abilities,” in ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2024, pp. 13 351–13 355
2024
Closest in time.