Fetching the paper…
Reading the bibliography…
We present results that show it is possible to build a competitive, greatly simplified, large vocabulary continuous speech recognition system with whole words as acoustic units.
A time-delay neural network architecture for isolated word recognition
Kevin J. Lang, Alex H. Waibel, and Geoffrey E. Hinton · 1990
Earlier work this paper cites.
Context dependent modelling of phones in continuous speech using decision trees
L.R. Bahl, P.V. de Souza, P.S. Gopalakrishnan, D. Nahamoo, and M.A. Picheny · 1991
Earlier work this paper cites.
Tree-based state tying for high accuracy acoustic modelling
S.J. Young, J.J. Odell, and P.C. Woodland · 1994
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Bidirectional recurrent neural networks
Mike Schuster and Kuldip K. Paliwal · 1997
Earlier work this paper cites.
Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber · 2006
Earlier work this paper cites.
Large scale distributed deep networks
Jeffrey Dean, Greg Corrado, Rajat Monga, Kai Chen, Matthieu Devin, Quoc V. Le, Mark Z. Mao, Marc’Aurelio Ranzato, Andrew W. Senior, Paul A. Tucker, Ke Yang, and Andrew Y. Ng · 2012
Earlier work this paper cites.
Language model verbalization for automatic speech recognition
Haşim Sak, Françoise Beaufays, Kaisuke Nakajima, and Cyril Allauzen · 2013
Earlier work this paper cites.
Large scale deep neural network acoustic modeling with semi-supervised training data for youtube video transcription
Hank Liao, Erik McDermott, and Andrew Senior · 2013
Cited alongside, same era.
Long Short-Term Memory Recurrent Neural Network Architectures for Large Scale Acoustic Modeling
Hasim Sak, Andrew Senior, and Francoise Beaufays · 2014
Cited alongside, same era.
A big data approach to acoustic model training corpus selection
Olga Kapralova, John Alex, Eugene Weinstein, Pedro Moreno, and Olivier Siohan · 2014
Cited alongside, same era.
EESEN: End-to-end speech recognition using deep RNN models and WFST-based decoding
Yajie Miao, Mohammad Gowayyed, and Florian Metze · 2015
Cited alongside, same era.
LibriSpeech: an ASR corpus based on public domain audio books
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur · 2015
Cited alongside, same era.
Towards end-to-end speech recognition with deep convolutional neural networks
Ying Zhang, Mohammad Pezeshki, Philémon Brakel, Saizheng Zhang, César Laurent, Yoshua Bengio, and Aaron Courville · 2016
Closest in time.
Deep Speech 2: End-to-end speech recognition in English and Mandarin
Dario Amodei, et al · 2016
Closest in time.
On training the recurrent neural network encoder-decoder for large vocabulary end-to-end speech recognition
Liang Lu, Xing-Xing Zhang, and Steve Renals · 2016
Closest in time.
Listen, attend and spell: A neural network for large vocabulary conversational speech recognition
William Chan, Navdeep Jaitly, Quoc V. Le, and Oriol Vinyals · 2016
Closest in time.
http://googleblog.blogspot.com/2009/11/automatic-captions-in-youtube.html
Automatic captions in YouTube · 2016
Closest in time.
http://youtube.com/yt/lineups/
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Senior, H. Sak, and I. Shafran · 2015
Cited alongside, same era.
Towards end-to-end speech recognition with recurrent neural networks
Alex Graves and Navdeep Jaitly · 2016
Cited alongside, same era.
End-to-end attention-based large vocabulary speech recognition
Dzmitry Bahdanau, Jan Chorowski, Dmitriy Serdyuk, Philemon Brakel, and Yoshua Bengio · 2016
Cited alongside, same era.
Fast and accurate recurrent neural network acoustic models for speech recognition
Haşim Sak, Andrew Senior, Kanishka Rao, and Françoise Beaufays
Cited in the paper.
Learning acoustic frame labeling for speech recognition with recurrent neural networks
Haşim Sak, Andrew Senior, Kanishka Rao, Ozan Irsoy, Alex Graves, Françoise Beaufays, and Johan Schalkwyk
Cited in the paper.
Google preferred lineup explorer - YouTube · 2016
Closest in time.
Learning N-gram language models from uncertain data
Vitaly Kuznetsov, Hank Liao, Mehryar Mohri, Michael Riley, and Brian Roark · 2016
Closest in time.