Fetching the paper…
Reading the bibliography…
In this paper, we consider the task of spotting spoken keywords in silent video sequences -- also known as visual keyword spotting.
Dynamic programming algorithm optimization for spoken word recognition
Hiroaki Sakoe and Seibi Chiba · 1978
Earlier work this paper cites.
Application of hidden markov models for recognition of a limited set of words in unconstrained speech
Jay Wilpon, Chin-Hui Lee, and Lawrence Rabiner · 1989
Earlier work this paper cites.
Minimum Prediction Residual Principle Applied to Speech Recognition , page 154–158
Fumitada Itakura · 1990
Earlier work this paper cites.
A hidden markov model based keyword recognition system
Richard C Rose and Douglas B Paul · 1990
Earlier work this paper cites.
The Hands Are The Head of The Mouth. The Mouth as Articulator in Sign Languages
P Boyes Braem and RL Sutton-Spence · 2001
Earlier work this paper cites.
Gradient flow in recurrent nets: the difficulty of learning long-term dependencies, 2001
Sepp Hochreiter, Yoshua Bengio, Paolo Frasconi, Jürgen Schmidhuber, et al · 2001
Earlier work this paper cites.
Recent advances in the automatic recognition of audiovisual speech
Gerasimos Potamianos, Chalapathy Neti, Guillaume Gravier, Ashutosh Garg, and Andrew W Senior · 2003
Earlier work this paper cites.
Dbn based multi-stream models for audio-visual speech recognition
John N Gowdy, Amarnag Subramanya, Chris Bartels, and Jeff Bilmes · 2004
Earlier work this paper cites.
An application of recurrent neural networks to discriminative keyword spotting
Santiago Fernandez, Alex Graves, and Jurgen Schmidhuber · 2007
Earlier work this paper cites.
Mouthings and simultaneity in british sign language
Rachel Sutton-Spence · 2007
Earlier work this paper cites.
Adaptive multimodal fusion by uncertainty compensation with application to audiovisual speech recognition
George Papandreou, Athanassios Katsamanis, Vassilis Pitsikalis, and Petros Maragos · 2009
Earlier work this paper cites.
Unsupervised spoken keyword spotting via segmental dtw on gaussian posteriorgrams
Yaodong Zhang and James Glass · 2009
Earlier work this paper cites.
Lattice indexing for spoken term detection
Dogan Can and M. Saraçlar · 2011
Earlier work this paper cites.
Building the British Sign Language Corpus
Adam Schembri, Jordan Fenlon, Ramas Rentelis, Sally Reynolds, and Kearsy Cormier · 2013
Earlier work this paper cites.
CMU pronouncing dictionary
Speech Group at Carnegie Mellon University · 2014
Earlier work this paper cites.
A review of recent advances in visual speech decoding
Ziheng Zhou, Guoying Zhao, Xiaopeng Hong, and Matti Pietikäinen · 2014
Earlier work this paper cites.
Online keyword spotting with a character-level recurrent neural network
Kyuyeon Hwang, Minjae Lee, and Wonyong Sung · 2015
Earlier work this paper cites.
ADAM: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Convolutional neural networks for small-footprint keyword spotting
Tara N. Sainath and Carolina Parada · 2015
Earlier work this paper cites.
Lipnet: Sentence-level lipreading
Yannis M. Assael, Brendan Shillingford, Shimon Whiteson, and Nando de Freitas · 2016
Earlier work this paper cites.
An end-to-end architecture for keyword spotting and voice activity detection
Chris Lengerich and Awni Hannun · 2016
Earlier work this paper cites.
Jointly learning to locate and classify words using convolutional networks
Dimitri Palaz, Gabriel Synnaeve, and Ronan Collobert · 2016
Earlier work this paper cites.
Max-pooling loss training of long short-term memory networks for small-footprint keyword spotting
Ming Sun, Anirudh Raju, George Tucker, Sankaran Panchapagesan, Gengshen Fu, Arindam Mandal, Spyridon Matsoukas, Nikko Strom, and Shiv Vitaladevuni · 2016
Earlier work this paper cites.
A novel lip descriptor for audio-visual keyword spotting based on adaptive decision fusion
Pingping Wu, Hong Liu, Xiaofei Li, Ting Fan, and Xuewu Zhang · 2016
Earlier work this paper cites.
Unrestricted Vocabulary Keyword Spotting Using LSTM-CTC
Yimeng Zhuang, Xuankai Chang, Yanmin Qian, and Kai Yu · 2016
Earlier work this paper cites.
Convolutional recurrent neural networks for small-footprint keyword spotting
Sercan Arik, Markus Kliegl, Rewon Child, Joel Hestness, Andrew Gibiansky, Chris Fougner, Ryan Prenger, and Adam Coates · 2017
Cited alongside, same era.
End to-end asr-free keyword search from speech
Kartik Audhkhasi, Andrew Rosenberg, Abhinav Sethy, Bhuvana Ramabhadran, and Brian Kingsbury · 2017
Cited alongside, same era.
Lip reading sentences in the wild
Joon Son Chung, Andrew Senior, Oriol Vinyals, and Andrew Zisserman · 2017
Cited alongside, same era.
Tall: Temporal activity localization via language query
Jiyang Gao, Chen Sun, Zhenheng Yang, and Ram Nevatia · 2017
Cited alongside, same era.
Streaming small-footprint keyword spotting using sequence-to-sequence models
Yanzhang He, Rohit Prabhavalkar, Kanishka Rao, Wei Li, Anton Bakhtin, and Ian McGraw · 2017
Cited alongside, same era.
Localizing moments in video with natural language
Lisa Anne Hendricks, O. Wang, E. Shechtman, Josef Sivic, Trevor Darrell, and Bryan C. Russell · 2017
Excl: Extractive clip localization using natural language descriptions
Soham Ghosh, Anuva Agarwal, Zarana Parekh, and Alexander Hauptmann · 2019
Later among the works it cites.
Improving transformer-based end-to-end speech recognition with connectionist temporal classification and language model integration
Shigeki Karita, Nelson Yalta, Shinji Watanabe, M. Delcroix, A. Ogawa, and T. Nakatani · 2019
Later among the works it cites.
Temporal feedback convolutional recurrent neural networks for keyword spotting
Taejun Kim and Juhan Nam · 2019
Later among the works it cites.
Improving audio-visual speech recognition performance with cross-modal student-teacher training
Wei Li, Sicheng Wang, Ming Lei, Sabato Marco Siniscalchi, and Chin-Hui Lee · 2019
Later among the works it cites.
Recurrent neural network transducer for audio-visual speech recognition
Takaki Makino, Hank Liao, Yannis Assael, Brendan Shillingford, Basilio Garcia, Otavio Braga, and Olivier Siohan · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
End-to-end speech recognition and keyword search on low-resource languages
Andrew Rosenberg, Kartik Audhkhasi, Abhinav Sethy, Bhuvana Ramabhadran, and Michael Picheny · 2017
Cited alongside, same era.
British Sign Language Corpus Project: A corpus of digital video data and annotations of British Sign Language 2008-2017 (Third Edition), 2017
Adam Schembri, Jordan Fenlon, Ramas Rentelis, and Kearsy Cormier · 2017
Cited alongside, same era.
Combining residual networks with lstms for lipreading
Themos Stafylakis and Georgios Tzimiropoulos · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Trainable frontend for robust and far-field keyword spotting
Yuxuan Wang, Pascal Getreuer, Thad Hughes, Richard Lyon, and Rif Saurous · 2017
Cited alongside, same era.
Hello edge: Keyword spotting on microcontrollers
Yundong Zhang, Naveen Suda, Liangzhen Lai, and Vikas Chandra · 2017
Cited alongside, same era.
Later among the works it cites.
Transformers with convolutional context for asr
Abdelrahman Mohamed, Dmytro Okhonko, and Luke Zettlemoyer · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Later among the works it cites.
Large-Scale Visual Speech Recognition
Brendan Shillingford, Yannis Assael, Matthew W. Hoffman, Thomas Paine, Cían Hughes, Utsav Prabhu, Hank Liao, Hasim Sak, Kanishka Rao, Lorrayne Bennett, Marie Mulville, Ben Coppin, Ben Laurie, Andrew Senior, and Nando de Freitas · 2019
Later among the works it cites.
Multilevel language and vision integration for text-to-clip retrieval
Huijuan Xu, Kun He, Bryan A. Plummer, L. Sigal, S. Sclaroff, and Kate Saenko · 2019
Later among the works it cites.
Spotting visual keywords from temporal sliding windows
Yue Yao, Tianyu Wang, Heming Du, Liang Zheng, and Tom Gedeon · 2019
Later among the works it cites.
To find where you talk: Temporal sentence localization in video with attention based location regression
Yitian Yuan, T. Mei, and Wenwu Zhu · 2019
Later among the works it cites.
Spatio-temporal fusion based convolutional sequence learning for lip reading
Xingxuan Zhang, Feng Cheng, and Shilin Wang · 2019
Later among the works it cites.
Asr is all you need: Cross-modal distillation for lip reading
Triantafyllos Afouras, Joon Son Chung, and Andrew Zisserman · 2020
Later among the works it cites.
BSL-1K: Scaling up co-articulated sign language recognition using mouthing cues
Samuel Albanie, Gül Varol, Liliane Momeni, Triantafyllos Afouras, Joon Son Chung, Neil Fox, and Andrew Zisserman · 2020
Later among the works it cites.
Conformer: Convolution-augmented transformer for speech recognition
Anmol Gulati, James Qin, Chung-Cheng Chiu, Niki Parmar, Yu Zhang, Jiahui Yu, Wei Han, Shibo Wang, Zhengdong Zhang, Yonghui Wu, and Ruoming Pang · 2020
Later among the works it cites.
Visual transformers: Token-based image representation and processing for computer vision
Bichen Wu, Chenfeng Xu, Xiaoliang Dai, Alvin Wan, Peizhao Zhang, Masayoshi Tomizuka, Kurt Keutzer, and Peter Vajda · 2020
Later among the works it cites.
Audio-visual recognition of overlapped speech for the lrs2 dataset
Jianwei Yu, Shi-Xiong Zhang, Jian Wu, Shahram Ghorbani, Bo Wu, Shiyin Kang, Shansong Liu, Xunying Liu, Helen Meng, and Dong Yu · 2020
Later among the works it cites.
Dense regression network for video grounding
Runhao Zeng, H. Xu, W. Huang, Peihao Chen, Mingkui Tan, and Chuang Gan · 2020
Later among the works it cites.
Keyword transformer: A self-attention model for keyword spotting
Axel Berg, Mark O’Connor, and Miguel Tairum Cruz · 2021
Closest in time.
Is space-time attention all you need for video understanding?
Gedas Bertasius, Heng Wang, and Lorenzo Torresani · 2021
Closest in time.
Aligning subtitles in sign language videos
Hannah Bull, Triantafyllos Afouras, Gül Varol, Samuel Albanie, Liliane Momeni, and Andrew Zisserman · 2021
Closest in time.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Closest in time.
Sub-word level lip reading with visual attention
Prajwal K R, Triantafyllos Afouras, Andrew Zisserman, et al · 2021
Closest in time.
Read and attend: Temporal localisation in sign language videos
Gül Varol, Liliane Momeni, Samuel Albanie, Triantafyllos Afouras, and Andrew Zisserman · 2021
Closest in time.