Fetching the paper…
Reading the bibliography…
This paper proposes a powerful Visual Speech Recognition (VSR) method for multiple languages, especially for low-resource languages that have a limited number of labeled data.
“Language identification: A tutorial”
Eliathamby Ambikairajah et al · 2011
Earlier work this paper cites.
“Deep speech: Scaling up end-to-end speech recognition”
Awni Hannun et al · 2014
Earlier work this paper cites.
“Adam: A method for stochastic optimization”
Diederik Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
“Deep speech 2: End-to-end speech recognition in english and mandarin”
Dario Amodei et al · 2016
Earlier work this paper cites.
“Lipnet: End-to-end sentence-level lipreading”
Yannis Assael et al · 2016
Earlier work this paper cites.
“Deep residual learning for image recognition”
Kaiming He et al · 2016
Earlier work this paper cites.
“Hybrid CTC/attention architecture for end-to-end speech recognition”
Shinji Watanabe et al · 2017
Earlier work this paper cites.
“Lip reading sentences in the wild”
Joon Chung et al · 2017
Earlier work this paper cites.
“Lip reading in the wild”
Joon Chung and Andrew Zisserman · 2017
Earlier work this paper cites.
“Attention is all you need”
Ashish Vaswani et al · 2017
Earlier work this paper cites.
“Deep audio-visual speech recognition”
Triantafyllos Afouras et al · 2018
Earlier work this paper cites.
“End-to-end audiovisual speech recognition”
Stavros Petridis et al · 2018
Earlier work this paper cites.
“LRS3-TED: a large-scale dataset for visual speech recognition”
Triantafyllos Afouras, Joon Chung and Andrew Zisserman · 2018
Earlier work this paper cites.
“VoxCeleb2: Deep speaker recognition”
J Chung, A Nagrani and A Zisserman · 2018
Earlier work this paper cites.
“Looking to listen at the cocktail party: a speaker-independent audio-visual model for speech separation”
Ariel Ephrat et al · 2018
Earlier work this paper cites.
Taku Kudo and John Richardson · 2018
Cited alongside, same era.
“Decoupled Weight Decay Regularization”
Ilya Loshchilov and Frank Hutter · 2018
Cited alongside, same era.
“Spatio-temporal fusion based convolutional sequence learning for lip reading”
Xingxuan Zhang, Feng Cheng and Shilin Wang · 2019
Cited alongside, same era.
“Conformer: Convolution-augmented transformer for speech recognition”
Anmol Gulati et al · 2020
Cited alongside, same era.
“Hearing lips: Improving lip reading by distilling speech recognizers”
Ya Zhao et al · 2020
Cited alongside, same era.
“Asr is all you need: Cross-modal distillation for lip reading”
“Distinguishing homophenes using multi-head visual-audio memory for lip reading”
Minsu Kim, Jeong Yeo and Yong Ro · 2022
Later among the works it cites.
“Visual speech recognition in a driver assistance system”
Denis Ivanko et al · 2022
Later among the works it cites.
“Visual speech recognition for multiple languages in the wild”
Pingchuan Ma, Stavros Petridis and Maja Pantic · 2022
Later among the works it cites.
“End-to-End Speech Recognition: A Survey”
Rohit Prabhavalkar et al · 2023
Closest in time.
“Exploration of Efficient End-to-End ASR using Discretized Input from Self-Supervised Learning”
Xuankai Chang et al · 2023
Closest in time.
“Lip-to-speech synthesis in the wild with multi-task learning”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Triantafyllos Afouras, Joon Chung and Andrew Zisserman · 2020
Cited alongside, same era.
“CMU-MOSEAS: A multimodal language dataset for Spanish, Portuguese, German and French”
Amir Zadeh et al · 2020
Cited alongside, same era.
“Retinaface: Single-shot multi-level face localisation in the wild”
Jiankang Deng et al · 2020
Cited alongside, same era.
“Recent developments on espnet toolkit boosted by conformer”
Pengcheng Guo et al · 2021
Cited alongside, same era.
“End-to-end audio-visual speech recognition with conformers”
Pingchuan Ma, Stavros Petridis and Maja Pantic · 2021
Cited alongside, same era.
“Multilingual TEDx Corpus for Speech Recognition and Translation”
Elizabeth Salesky et al · 2021
Cited alongside, same era.
“Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction”
Bowen Shi et al · 2021
Cited alongside, same era.
Minsu Kim, Joanna Hong and Yong Ro · 2023
Closest in time.
“Intelligible Lip-to-Speech Synthesis with Speech Units”
Jeongsoo Choi, Minsu Kim and Yong Ro · 2023
Closest in time.
“Multi-Temporal Lip-Audio Memory for Visual Speech Recognition”
Jeong Yeo, Minsu Kim and Yong Ro · 2023
Closest in time.
Jeong Yeo et al · 2023
Closest in time.
“Conformers are All You Need for Visual Speech Recogntion”
Oscar Chang et al · 2023
Closest in time.
“Auto-AVSR: Audio-visual speech recognition with automatic labels”
Pingchuan Ma et al · 2023
Closest in time.
“Lip Reading for Low-resource Languages by Learning and Combining General Speech Knowledge and Language-specific Knowledge”
Minsu Kim et al · 2023
Closest in time.
“Learning Cross-Lingual Visual Speech Representations”
Andreas Zinonos et al · 2023
Closest in time.
“Robust speech recognition via large-scale weak supervision”
Alec Radford et al · 2023
Closest in time.