Fetching the paper…
Reading the bibliography…
This paper explores sentence-level multilingual Visual Speech Recognition (VSR) that can recognize different languages with a single trained model.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Empirical evaluation of gated recurrent neural networks on sequence modeling. In NIPS 2014 Workshop on Deep Learning, December 2014
Junyoung Chung, Caglar Gulcehre, Kyunghyun Cho, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. 2015 · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization. In International Conference on Learning Representations
Diederik P Kingma and Jimmy Ba. 2015 · 2015
Earlier work this paper cites.
Memory networks. In 3rd International Conference on Learning Representations, ICLR 2015
Jason Weston, Sumit Chopra, and Antoine Bordes. 2015 · 2015
Earlier work this paper cites.
Lipnet: End-to-end sentence-level lipreading
Yannis M Assael, Brendan Shillingford, Shimon Whiteson, and Nando De Freitas. 2016 · 2016
Earlier work this paper cites.
Lip reading in the wild. In Computer Vision–ACCV 2016: 13th Asian Conference on Computer Vision, Taipei, Taiwan, November 20-24, 2016, Revised Selected Papers, Part II 13 . Springer, 87–103
Joon Son Chung and Andrew Zisserman. 2017b · 2016
Earlier work this paper cites.
Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition . 770–778
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Earlier work this paper cites.
Deep complementary bottleneck features for visual speech recognition. In 2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2304–2308
Stavros Petridis and Maja Pantic. 2016 · 2016
Earlier work this paper cites.
Lip Reading Sentences in the Wild. In 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) . IEEE
Joon Son Chung, Andrew Senior, Oriol Vinyals, and Andrew Zisserman. 2017 · 2017
Earlier work this paper cites.
Lip reading in profile. In British Machine Vision Conference, 2017 . British Machine Vision Association and Society for Pattern Recognition
Joon Son Chung and Andrew Zisserman. 2017a · 2017
Earlier work this paper cites.
End-to-end visual speech recognition with LSTMs. In 2017 IEEE international conference on acoustics, speech and signal processing (ICASSP) . IEEE, 2592–2596
Stavros Petridis, Zuwei Li, and Maja Pantic. 2017 · 2017
Earlier work this paper cites.
Combining residual networks with LSTMs for lipreading. In Proc. Interspeech
Themos Stafylakis and Georgios Tzimiropoulos. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Hybrid CTC/attention architecture for end-to-end speech recognition
Shinji Watanabe, Takaaki Hori, Suyoun Kim, John R Hershey, and Tomoki Hayashi. 2017 · 2017
Earlier work this paper cites.
Deep audio-visual speech recognition
Triantafyllos Afouras, Joon Son Chung, Andrew Senior, Oriol Vinyals, and Andrew Zisserman. 2018b · 2018
Earlier work this paper cites.
LRS3-TED: a large-scale dataset for visual speech recognition
Triantafyllos Afouras, Joon Son Chung, and Andrew Zisserman. 2018a · 2018
Earlier work this paper cites.
VoxCeleb2: Deep Speaker Recognition
Joon Son Chung, Arsha Nagrani, and Andrew Zisserman. 2018 · 2018
Earlier work this paper cites.
Looking to listen at the cocktail party: a speaker-independent audio-visual model for speech separation
Ariel Ephrat, Inbar Mosseri, Oran Lang, Tali Dekel, Kevin Wilson, Avinatan Hassidim, William T Freeman, and Michael Rubinstein. 2018 · 2018
Earlier work this paper cites.
Taku Kudo and John Richardson. 2018 · 2018
Earlier work this paper cites.
End-to-end audiovisual speech recognition. In 2018 IEEE international conference on acoustics, speech and signal processing (ICASSP) . IEEE, 6548–6552
Stavros Petridis, Themos Stafylakis, Pingehuan Ma, Feipeng Cai, Georgios Tzimiropoulos, and Maja Pantic. 2018a · 2018
Earlier work this paper cites.
Audio-visual speech recognition with a hybrid ctc/attention architecture. In 2018 IEEE Spoken Language Technology Workshop (SLT) . IEEE, 513–520
Stavros Petridis, Themos Stafylakis, Pingchuan Ma, Georgios Tzimiropoulos, and Maja Pantic. 2018b · 2018
Earlier work this paper cites.
Multilingual speech recognition with a single end-to-end model. In 2018 IEEE international conference on acoustics, speech and signal processing (ICASSP) . IEEE, 4904–4908
Shubham Toshniwal, Tara N Sainath, Ron J Weiss, Bo Li, Pedro Moreno, Eugene Weinstein, and Kanishka Rao. 2018 · 2018
Earlier work this paper cites.
Massively Multilingual Adversarial Speech Recognition. In Proceedings of NAACL-HLT . 96–108
Oliver Adams, Matthew Wiesner, Shinji Watanabe, and David Yarowsky. 2019 · 2019
Earlier work this paper cites.
Cross-lingual language model pretraining
Alexis Conneau and Guillaume Lample. 2019 · 2019
Earlier work this paper cites.
Multilingual end-to-end speech translation. In 2019 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) . IEEE, 570–577
Hirofumi Inaguma, Kevin Duh, Tatsuya Kawahara, and Shinji Watanabe. 2019 · 2019
Earlier work this paper cites.
Asr is all you need: Cross-modal distillation for lip reading. In ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2143–2147
Triantafyllos Afouras, Joon Son Chung, and Andrew Zisserman. 2020 · 2020
Cited alongside, same era.
wav2vec 2.0: A framework for self-supervised learning of speech representations
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli. 2020 · 2020
Cited alongside, same era.
Retinaface: Single-shot multi-level face localisation in the wild. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 5203–5212
Jiankang Deng, Jia Guo, Evangelos Ververas, Irene Kotsia, and Stefanos Zafeiriou. 2020 · 2020
Cited alongside, same era.
ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN based speaker verification. In Proc. Interspeech . International Speech Communication Association (ISCA), 3830–3834
Brecht Desplanques, Jenthe Thienpondt, and Kris Demuynck. 2020 · 2020
Cited alongside, same era.
Visual speech recognition for multiple languages in the wild
Pingchuan Ma, Stavros Petridis, and Maja Pantic. 2022a · 2022
Later among the works it cites.
Training strategies for improved lip-reading. In ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 8472–8476
Pingchuan Ma, Yujiang Wang, Stavros Petridis, Jie Shen, and Maja Pantic. 2022b · 2022
Later among the works it cites.
Enhanced direct speech-to-speech translation using self-supervised pre-training and data augmentation. In Proc. Interspeech
Sravya Popuri, Peng-Jen Chen, Changhan Wang, Juan Pino, Yossi Adi, Jiatao Gu, Wei-Ning Hsu, and Ann Lee. 2022 · 2022
Later among the works it cites.
Sub-word level lip reading with visual attention. In Proceedings of the IEEE/CVF conference on Computer Vision and Pattern Recognition . 5162–5172
KR Prajwal, Triantafyllos Afouras, and Andrew Zisserman. 2022 · 2022
Later among the works it cites.
Adaspeech 4: Adaptive text to speech in zero-shot scenarios. In Proc. Interspeech
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Synchronous bidirectional learning for multilingual lip reading. In British Machine Vision Conference
Mingshuang Luo, Shuang Yang, Xilin Chen, Zitao Liu, and Shiguang Shan. 2020 · 2020
Cited alongside, same era.
Mls: A large-scale multilingual dataset for speech research. In Proc. Interspeech
Vineel Pratap, Qiantong Xu, Anuroop Sriram, Gabriel Synnaeve, and Ronan Collobert. 2020 · 2020
Cited alongside, same era.
Hearing lips: Improving lip reading by distilling speech recognizers. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 34. 6917–6924
Ya Zhao, Rui Xu, Xinchao Wang, Peng Hou, Haihong Tang, and Mingli Song. 2020 · 2020
Cited alongside, same era.
The Multilingual TEDx Corpus for Speech Recognition and Translation
Salesky Elizabeth, Wiesner Matthew, Bremerman Jacob, Roldano Cattoni, Matteo Negri, Marco Turchi, Douglas W Oard, and Post Matt. 2021 · 2021
Cited alongside, same era.
Recent developments on espnet toolkit boosted by conformer. In Proc. ICASSP . IEEE, 5874–5878
Pengcheng Guo, Florian Boyer, Xuankai Chang, Tomoki Hayashi, Yosuke Higuchi, Hirofumi Inaguma, Naoyuki Kamo, Chenda Li, Daniel Garcia-Romero, Jiatong Shi, et al · 2021
Cited alongside, same era.
Hubert: Self-supervised speech representation learning by masked prediction of hidden units
Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed. 2021 · 2021
Cited alongside, same era.
Cromm-vsr: Cross-modal memory augmented visual speech recognition
Minsu Kim, Joanna Hong, Se Jin Park, and Yong Man Ro. 2021 · 2021
Cited alongside, same era.
On generative spoken language modeling from raw audio
Kushal Lakhotia, Eugene Kharitonov, Wei-Ning Hsu, Yossi Adi, Adam Polyak, Benjamin Bolte, Tu-Anh Nguyen, Jade Copet, Alexei Baevski, Abdelrahman Mohamed, et al · 2021
Cited alongside, same era.
Yihan Wu, Xu Tan, Bohan Li, Lei He, Sheng Zhao, Ruihua Song, Tao Qin, and Tie-Yan Liu. 2022 · 2022
Later among the works it cites.
Conformers are All You Need for Visual Speech Recogntion
Oscar Chang, Hank Liao, Dmitriy Serdyuk, Ankit Shah, and Olivier Siohan. 2023a · 2023
Later among the works it cites.
Intelligible Lip-to-Speech Synthesis with Speech Units. In Proc. Interspeech
Jeongsoo Choi, Minsu Kim, and Yong Man Ro. 2023 · 2023
Later among the works it cites.
Yasser Abdelaziz Dahou Djilali, Sanath Narayan, Haithem Boussaid, Ebtessam Almazrouei, and Merouane Debbah. 2023 · 2023
Later among the works it cites.
Watch or Listen: Robust Audio-Visual Speech Recognition with Visual Corruption Modeling and Reliability Scoring. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 18783–18794
Joanna Hong, Minsu Kim, Jeongsoo Choi, and Yong Man Ro. 2023 · 2023
Later among the works it cites.
Mega-TTS: Zero-Shot Text-to-Speech at Scale with Intrinsic Inductive Bias
Ziyue Jiang, Yi Ren, Zhenhui Ye, Jinglin Liu, Chen Zhang, Qian Yang, Shengpeng Ji, Rongjie Huang, Chunfeng Wang, Xiang Yin, et al · 2023
Later among the works it cites.
Minsu Kim, Jeongsoo Choi, Dahun Kim, and Yong Man Ro. 2023a · 2023
Later among the works it cites.
Minsu Kim, Jeongsoo Choi, Soumi Maiti, Jeong Hun Yeo, Shinji Watanabe, and Yong Man Ro. 2023b · 2023
Later among the works it cites.
Auto-AVSR: Audio-visual speech recognition with automatic labels. In ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 1–5
Pingchuan Ma, Alexandros Haliassos, Adriana Fernandez-Lopez, Honglie Chen, Stavros Petridis, and Maja Pantic. 2023 · 2023
Later among the works it cites.
Soumi Maiti, Yifan Peng, Shukjae Choi, Jee-weon Jung, Xuankai Chang, and Shinji Watanabe. 2023 · 2023
Later among the works it cites.
Generative spoken dialogue language modeling
Tu Anh Nguyen, Eugene Kharitonov, Jade Copet, Yossi Adi, Wei-Ning Hsu, Ali Elkahky, Paden Tomasello, Robin Algayres, Benoit Sagot, Abdelrahman Mohamed, et al · 2023
Later among the works it cites.
SeiT: Storage-Efficient Vision Training with Tokens Using 1% of Pixel Storage. In Proc. International Conference on Computer Vision
Song Park, Sanghyuk Chun, Byeongho Heo, Wonjae Kim, and Sangdoo Yun. 2023 · 2023
Later among the works it cites.
End-to-End Speech Recognition: A Survey
Rohit Prabhavalkar, Takaaki Hori, Tara N Sainath, Ralf Schlüter, and Shinji Watanabe. 2023 · 2023
Later among the works it cites.
Robust speech recognition via large-scale weak supervision. In International Conference on Machine Learning . PMLR, 28492–28518
Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever. 2023 · 2023
Later among the works it cites.
Analysing discrete self supervised speech representation for spoken language modeling. In ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 1–5
Amitay Sicherman and Yossi Adi. 2023 · 2023
Later among the works it cites.
Multi-Temporal Lip-Audio Memory for Visual Speech Recognition. In ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 1–5
Jeong Hun Yeo, Minsu Kim, and Yong Man Ro. 2023a · 2023
Later among the works it cites.
Visual Speech Recognition for Low-resource Languages with Automatic Labels From Whisper Model
Jeong Hun Yeo, Minsu Kim, Shinji Watanabe, and Yong Man Ro. 2023b · 2023
Later among the works it cites.
Speak foreign languages with your own voice: Cross-lingual neural codec language modeling
Ziqiang Zhang, Long Zhou, Chengyi Wang, Sanyuan Chen, Yu Wu, Shujie Liu, Zhuo Chen, Yanqing Liu, Huaming Wang, Jinyu Li, et al · 2023
Later among the works it cites.
Vatlm: Visual-audio-text pre-training with unified masked prediction for speech representation learning
Qiushi Zhu, Long Zhou, Ziqiang Zhang, Shujie Liu, Binxing Jiao, Jie Zhang, Lirong Dai, Daxin Jiang, Jinyu Li, and Furu Wei. 2023 · 2023
Later among the works it cites.
Learning Cross-Lingual Visual Speech Representations. In ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 1–5
Andreas Zinonos, Alexandros Haliassos, Pingchuan Ma, Stavros Petridis, and Maja Pantic. 2023 · 2023
Later among the works it cites.
AKVSR: Audio Knowledge Empowered Visual Speech Recognition by Compressing Audio Knowledge of a Pretrained Model
Jeong Hun Yeo, Minsu Kim, Jeongsoo Choi, Dae Hoe Kim, and Yong Man Ro. 2024 · 2024
Closest in time.