Fetching the paper…
Reading the bibliography…
Audio-visual automatic speech recognition (AV-ASR) introduces the video modality into the speech recognition process, often by relying on information conveyed by the motion of the speaker's mouth.
“Gradient-Based learning applied to document recognition,”
Y L LeCun, Leon Bottou, Yoshua Bengio, and Patrick Haffner, · 1998
Earlier work this paper cites.
“Audio-visual speech recognition,”
Chalapathy Neti, Gerasimos Potamianos, Juergen Luettin, Iain Matthews, Herve Glotin, Dimitra Vergyri, June Sison, and Azad Mashari, · 2000
Earlier work this paper cites.
“Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks,”
Alex Graves, Santiago Fernandez, Faustino Gomez, and Jurgen Schmidhuber, · 2006
Earlier work this paper cites.
“Sequence transduction with recurrent neural networks,”
Alex Graves, · 2012
Earlier work this paper cites.
“Large scale deep neural network acoustic modeling with semi-supervised training data for YouTube video transcription,”
Hank Liao, Erik McDermott, and Andrew Senior, · 2013
Earlier work this paper cites.
“Neural machine translation by jointly learning to align and translate,”
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio, · 2014
Earlier work this paper cites.
“Adam: A method for stochastic optimization,”
Diederik P Kingma and Jimmy Ba, · 2014
Earlier work this paper cites.
“Very deep convolutional networks for Large-Scale image recognition,”
Karen Simonyan and Andrew Zisserman, · 2015
Earlier work this paper cites.
“Data augmentation for deep neural network acoustic modeling,”
Xiaodong Cui, Vaibhava Goel, and Brian Kingsbury, · 2015
Earlier work this paper cites.
“LIPNET: Sentence-Level lipreading,”
Yannis M Assael, Brendan Shillingford, Shimon Whiteson, and Nando de Freitas, · 2016
Earlier work this paper cites.
“Lip reading sentences in the wild,”
Joon Son Chung, Andrew Senior, Oriol Vinyals, and Andrew Zisserman, · 2016
Earlier work this paper cites.
“Visual features for context-aware speech recognition,”
Abhinav Gupta, Yajie Miao, Leonardo Neves, and Florian Metze, · 2017
Earlier work this paper cites.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin, · 2017
Cited alongside, same era.
“A closer look at spatiotemporal convolutions for action recognition,”
Du Tran, Heng Wang, Lorenzo Torresani, Jamie Ray, Yann LeCun, and Manohar Paluri, · 2017
Cited alongside, same era.
“Rethinking spatiotemporal feature learning: Speed-Accuracy trade-offs in video classification,”
Saining Xie, Chen Sun, Jonathan Huang, Zhuowen Tu, and Kevin Murphy, · 2017
Cited alongside, same era.
“Deep audio-visual speech recognition,”
Triantafyllos Afouras, Joon Son Chung, Andrew Senior, Oriol Vinyals, and Andrew Zisserman, · 2018
Cited alongside, same era.
“BERT: Pre-training of deep bidirectional transformers for language understanding,”
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova, · 2018
Cited alongside, same era.
Qian Zhang, Han Lu, Hasim Sak, Anshuman Tripathi, Erik McDermott, Stephen Koo, and Shankar Kumar, · 2020
Later among the works it cites.
“Conformer: Convolution-augmented transformer for speech recognition,”
Anmol Gulati, James Qin, Chung-Cheng Chiu, Niki Parmar, Yu Zhang, Jiahui Yu, Wei Han, Shibo Wang, Zhengdong Zhang, Yonghui Wu, and Ruoming Pang, · 2020
Later among the works it cites.
“Discriminative multi-modality speech recognition,”
Bo Xu, Cheng Lu, Yandong Guo, and Jacob Wang, · 2020
Later among the works it cites.
“End-to-End Multi-Person Audio/Visual automatic speech recognition,”
Otavio Braga, Takaki Makino, Olivier Siohan, and Hank Liao, · 2020
Later among the works it cites.
“ViViT: A video vision transformer,”
Anurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun, Mario Lučić, and Cordelia Schmid, · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Self-Attention with relative position representations,”
Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani, · 2018
Cited alongside, same era.
“LRS3-TED: a large-scale dataset for visual speech recognition,”
Triantafyllos Afouras, Joon Son Chung, and Andrew Zisserman, · 2018
Cited alongside, same era.
“Recurrent neural network transducer for Audio-Visual speech recognition,”
Takaki Makino, Hank Liao, Yannis Assael, Brendan Shillingford, Basilio Garcia, Otavio Braga, and Olivier Siohan, · 2019
Cited alongside, same era.
“ASR is all you need: cross-modal distillation for lip reading,”
Triantafyllos Afouras, Joon Son Chung, and Andrew Zisserman, · 2019
Cited alongside, same era.
“Language models are unsupervised multitask learners,”
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever, · 2019
Cited alongside, same era.
“Large-Scale visual speech recognition,”
Brendan Shillingford, Yannis Assael, Matthew W Hoffman, Thomas Paine, Cían Hughes, Utsav Prabhu, Hank Liao, Hasim Sak, Kanishka Rao, Lorrayne Bennett, Marie Mulville, Misha Denil, Ben Coppin, Ben Laurie, Andrew Senior, and Nando de Freitas, · 2019
Cited alongside, same era.
“An image is worth 16x16 words: Transformers for image recognition at scale,”
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby, · 2020
Cited alongside, same era.
Closest in time.
“Is Space-Time attention all you need for video understanding?,”
Gedas Bertasius, Heng Wang, and Lorenzo Torresani, · 2021
Closest in time.
“End-to-end audio-visual speech recognition with conformers,”
Pingchuan Ma, Stavros Petridis, and Maja Pantic, · 2021
Closest in time.
“An image is worth 16x16 words, what is a video worth?,”
Gilad Sharir, Asaf Noy, and Lihi Zelnik-Manor, · 2021
Closest in time.
Daniel Neimark, Omri Bar, Maya Zohar, and Dotan Asselmann, · 2021
Closest in time.
“A closer look at Audio-Visual Multi-Person speech recognition and active speaker selection,”
Otavio Braga and Olivier Siohan, · 2021
Closest in time.
“Our principles – google AI,” https://ai.google/principles/, · 2021
Closest in time.