Fetching the paper…
Reading the bibliography…
In visual speech processing, context modeling capability is one of the most important requirements due to the ambiguous nature of lip movements.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020 · 1901
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Connectionist temporal classification
Alex Graves and Alex Graves. 2012 · 2012
Earlier work this paper cites.
Empirical evaluation of gated recurrent neural networks on sequence modeling
Junyoung Chung, Caglar Gulcehre, Kyunghyun Cho, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
Lipnet: End-to-end sentence-level lipreading
Yannis M Assael, Brendan Shillingford, Shimon Whiteson, and Nando De Freitas. 2016 · 2016
Earlier work this paper cites.
Lip reading in the wild
Joon Son Chung and Andrew Zisserman. 2017a · 2016
Earlier work this paper cites.
Lip reading in the wild
Joon Son Chung and Andrew Zisserman. 2017b · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Earlier work this paper cites.
Deep complementary bottleneck features for visual speech recognition
Stavros Petridis and Maja Pantic. 2016 · 2016
Earlier work this paper cites.
Lip reading sentences in the wild
Joon Son Chung, Andrew Senior, Oriol Vinyals, and Andrew Zisserman. 2017 · 2017
Earlier work this paper cites.
End-to-end visual speech recognition with lstms
Stavros Petridis, Zuwei Li, and Maja Pantic. 2017 · 2017
Earlier work this paper cites.
Combining residual networks with lstms for lipreading
Themos Stafylakis and Georgios Tzimiropoulos. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Lrs3-ted: a large-scale dataset for visual speech recognition
Triantafyllos Afouras, Joon Son Chung, and Andrew Zisserman. 2018 · 2018
Earlier work this paper cites.
Voxceleb2: Deep speaker recognition
Joon Son Chung, Arsha Nagrani, and Andrew Zisserman. 2018 · 2018
Earlier work this paper cites.
End-to-end audiovisual speech recognition
Stavros Petridis, Themos Stafylakis, Pingehuan Ma, Feipeng Cai, Georgios Tzimiropoulos, and Maja Pantic. 2018 · 2018
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
Recurrent neural network transducer for audio-visual speech recognition
Takaki Makino, Hank Liao, Yannis Assael, Brendan Shillingford, Basilio Garcia, Otavio Braga, and Olivier Siohan. 2019 · 2019
Cited alongside, same era.
Large-scale visual speech recognition
Brendan Shillingford, Yannis M. Assael, Matthew W. Hoffman, Thomas Paine, Cían Hughes, Utsav Prabhu, Hank Liao, Hasim Sak, Kanishka Rao, Lorrayne Bennett, Marie Mulville, Misha Denil, Ben Coppin, Ben Laurie, Andrew W. Senior, and Nando de Freitas. 2019 · 2019
Cited alongside, same era.
Asr is all you need: Cross-modal distillation for lip reading
Triantafyllos Afouras, Joon Son Chung, and Andrew Zisserman. 2020 · 2020
Cited alongside, same era.
Retinaface: Single-shot multi-level face localisation in the wild
Jiankang Deng, Jia Guo, Evangelos Ververas, Irene Kotsia, and Stefanos Zafeiriou. 2020 · 2020
Cited alongside, same era.
Hearing lips: Improving lip reading by distilling speech recognizers
Ya Zhao, Rui Xu, Xinchao Wang, Peng Hou, Haihong Tang, and Mingli Song. 2020 · 2020
Cited alongside, same era.
Learning audio-visual speech representation by masked multimodal cluster prediction
Bowen Shi, Wei-Ning Hsu, Kushal Lakhotia, and Abdelrahman Mohamed. 2022 · 2022
Later among the works it cites.
Bloom: A 176b-parameter open-access multilingual language model
BigScience Workshop, Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilić, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, et al. 2022 · 2022
Later among the works it cites.
Speechut: Bridging speech and text with hidden-unit for encoder-decoder based speech-text pre-training
Ziqiang Zhang, Long Zhou, Junyi Ao, Shujie Liu, Lirong Dai, Jinyu Li, and Furu Wei. 2022b · 2022
Later among the works it cites.
Mohamed Anwar, Bowen Shi, Vedanuj Goswami, Wei-Ning Hsu, Juan Pino, and Changhan Wang. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Beyond english-centric multilingual machine translation
Angela Fan, Shruti Bhosale, Holger Schwenk, Zhiyi Ma, Ahmed El-Kishky, Siddharth Goyal, Mandeep Baines, Onur Celebi, Guillaume Wenzek, Vishrav Chaudhary, et al. 2021 · 2021
Cited alongside, same era.
Hubert: Self-supervised speech representation learning by masked prediction of hidden units
Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed. 2021 · 2021
Cited alongside, same era.
Cromm-vsr: Cross-modal memory augmented visual speech recognition
Minsu Kim, Joanna Hong, Se Jin Park, and Yong Man Ro. 2021 · 2021
Cited alongside, same era.
On generative spoken language modeling from raw audio
Kushal Lakhotia, Eugene Kharitonov, Wei-Ning Hsu, Yossi Adi, Adam Polyak, Benjamin Bolte, Tu-Anh Nguyen, Jade Copet, Alexei Baevski, Abdelrahman Mohamed, and Emmanuel Dupoux. 2021 · 2021
Cited alongside, same era.
Towards practical lipreading with distilled and efficient models
Pingchuan Ma, Brais Martinez, Stavros Petridis, and Maja Pantic. 2021a · 2021
Cited alongside, same era.
End-to-end audio-visual speech recognition with conformers
Pingchuan Ma, Stavros Petridis, and Maja Pantic. 2021b · 2021
Cited alongside, same era.
Learning from the master: Distilling cross-modal advanced knowledge for lip reading
Sucheng Ren, Yong Du, Jianming Lv, Guoqiang Han, and Shengfeng He. 2021 · 2021
Cited alongside, same era.
Oscar Chang, Hank Liao, Dmitriy Serdyuk, Ankit Shah, and Olivier Siohan. 2023 · 2023
Later among the works it cites.
Mixspeech: Cross-modality self-learning with audio-visual stream mixup for visual speech translation and recognition
Xize Cheng, Tao Jin, Rongjie Huang, Linjun Li, Wang Lin, Zehan Wang, Ye Wang, Huadai Liu, Aoxiong Yin, and Zhou Zhao. 2023 · 2023
Later among the works it cites.
Qlora: Efficient finetuning of quantized llms
Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer. 2023 · 2023
Later among the works it cites.
Prompting large language models with speech recognition abilities
Yassir Fathullah, Chunyang Wu, Egor Lakomkin, Junteng Jia, Yuan Shangguan, Ke Li, Jinxi Guo, Wenhan Xiong, Jay Mahadeokar, Ozlem Kalinli, et al. 2023 · 2023
Later among the works it cites.
Imagebind-llm: Multi-modality instruction tuning
Jiaming Han, Renrui Zhang, Wenqi Shao, Peng Gao, Peng Xu, Han Xiao, Kaipeng Zhang, Chris Liu, Song Wen, Ziyu Guo, et al. 2023 · 2023
Later among the works it cites.
Auto-avsr: Audio-visual speech recognition with automatic labels
Pingchuan Ma, Alexandros Haliassos, Adriana Fernandez-Lopez, Honglie Chen, Stavros Petridis, and Maja Pantic. 2023 · 2023
Later among the works it cites.
Comparative layer-wise analysis of self-supervised speech models
Ankita Pasad, Bowen Shi, and Karen Livescu. 2023 · 2023
Later among the works it cites.
Audiopalm: A large language model that can speak and listen
Paul K Rubenstein, Chulayuth Asawaroengchai, Duc Dung Nguyen, Ankur Bapna, Zalán Borsos, Félix de Chaumont Quitry, Peter Chen, Dalia El Badawy, Wei Han, Eugene Kharitonov, et al. 2023 · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023 · 2023
Later among the works it cites.
On decoder-only architecture for speech-to-text and large language model integration
Jian Wu, Yashesh Gaur, Zhuo Chen, Long Zhou, Yimeng Zhu, Tianrui Wang, Jinyu Li, Shujie Liu, Bo Ren, Linquan Liu, et al. 2023a · 2023
Later among the works it cites.
Multi-temporal lip-audio memory for visual speech recognition
Jeong Hun Yeo, Minsu Kim, and Yong Man Ro. 2023b · 2023
Later among the works it cites.
Qiushi Zhu, Long Zhou, Ziqiang Zhang, Shujie Liu, Binxing Jiao, Jie Zhang, Lirong Dai, Daxin Jiang, Jinyu Li, and Furu Wei. 2023 · 2023
Later among the works it cites.
Lipvoicer: Generating speech from silent videos guided by lip reading
Yochai Yemini, Aviv Shamsian, Lior Bracha, Sharon Gannot, and Ethan Fetaya. 2024 · 2024
Closest in time.