Fetching the paper…
Reading the bibliography…
We show that an end-to-end deep learning approach can be used to recognize either English or Mandarin Chinese speech--two vastly different languages.
A fast storage allocator
K. C. Knowlton · 1965
Earlier work this paper cites.
Phoneme recognition using time-delay neural networks,” acoustics speech and signal processing
A. Waibel, T. Hanazawa, G. Hinton, K. Shikano, and K. Lang · 1989
Earlier work this paper cites.
An efficient gradient-based algorithm for online training of recurrent network trajectories
R. Williams and J. Peng · 1990
Earlier work this paper cites.
Connectionist Speech Recognition: A Hybrid Approach
H. Bourlard and N. Morgan · 1993
Earlier work this paper cites.
Connectionist probability estimators in HMM speech recognition
S. Renals, N. Morgan, H. Bourlard, M. Cohen, and H. Franco · 1994
Earlier work this paper cites.
The use of recurrent neural networks in continuous speech recognition
T. Robinson, M. Hochberg, and S. Renals · 1996
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Bidirectional recurrent neural networks
M. Schuster and K. K. Paliwal · 1997
Earlier work this paper cites.
Size matters: An empirical study of neural network training for large vocabulary continuous speech recognition
D. Ellis and N. Morgan · 1999
Earlier work this paper cites.
The Fisher corpus: a resource for the next generations of speech-to-text
C. Cieri, D. Miller, and K. Walker · 2004
Earlier work this paper cites.
Learning methods for generic object recognition with invariance to pose and lighting
Y. LeCun, F. J. Huang, and L. Bottou · 2004
Earlier work this paper cites.
Optimization of collective communication operations in mpich
R. Thakur and R. Rabenseifner · 2005
Earlier work this paper cites.
Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber · 2006
Earlier work this paper cites.
A fast data collection and augmentation procedure for object recognition
B. Sapp, A. Saxena, and A. Ng · 2008
Earlier work this paper cites.
Curriculum learning
Y. Bengio, J. Louradour, R. Collobert, and J. Weston · 2009
Earlier work this paper cites.
Support vector machines for noise robust ASR
M. J. F. Gales, A. Ragni, H. Aldamarki, and C. Gautier · 2009
Earlier work this paper cites.
Bandwidth optimal all-reduce algorithms for clusters of workstations
P. Patarasuk and X. Yuan · 2009
Earlier work this paper cites.
Large-scale deep unsupervised learning using graphics processors
R. Raina, A. Madhavan, and A. Ng · 2009
Earlier work this paper cites.
Search by voice in mandarin chinese
J. Shan, G. Wu, Z. Hu, X. Tang, M. Jansche, and P. Moreno · 2010
Earlier work this paper cites.
Text detection and character recognition in scene images with unsupervised feature learning
A. Coates, B. Carpenter, C. Case, S. Satheesh, B. Suresh, T. Wang, D. J. Wu, and A. Y. Ng · 2011
Earlier work this paper cites.
Large vocabulary continuous speech recognition with context-dependent DBN-HMMs
G. Dahl, D. Yu, and L. Deng · 2011
Earlier work this paper cites.
Context-dependent pre-trained deep neural networks for large vocabulary speech recognition
G. Dahl, D. Yu, L. Deng, and A. Acero · 2011
Earlier work this paper cites.
Acoustic modeling using deep belief networks
A. Mohamed, G. Dahl, and G. Hinton · 2011
Earlier work this paper cites.
Conversational speech transcription using context-dependent deep neural networks
F. Seide, G. Li, and D. Yu · 2011
Cited alongside, same era.
Applying convolutional neural networks concepts to hybrid nn-hmm model for speech recognition
O. Abdel-Hamid, A.-r. Mohamed, H. Jang, and G. Penn · 2012
Cited alongside, same era.
Large scale distributed deep networks
J. Dean, G. S. Corrado, R. Monga, K. Chen, M. Devin, Q. Le, M. Mao, M. Ranzato, A. Senior, P. Tucker, K. Yang, and A. Ng · 2012
Cited alongside, same era.
Deep neural networks for acoustic modeling in speech recognition
G. Hinton, L. Deng, D. Yu, G. Dahl, A. Mohamed, N. Jaitly, A. Senior, V. Vanhoucke, P. Nguyen, T. Sainath, and B. Kingsbury · 2012
Cited alongside, same era.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. Hinton · 2012
Cited alongside, same era.
Building high-level features using large scale unsupervised learning
Sequence discriminative distributed training of long shortterm memory recurrent neural networks
H. Sak, O. Vinyals, G. Heigold, A. Senior, E. McDermott, R. Monga, and M. Mao · 2014
Later among the works it cites.
Joint training of convolutional and non-convolutional neural networks
H. Soltau, G. Saon, and T. Sainath · 2014
Later among the works it cites.
Sequence to sequence learning with neural networks
I. Sutskever, O. Vinyals, and Q. V. Le · 2014
Later among the works it cites.
Going deeper with convolutions
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich · 2014
Later among the works it cites.
W. Zaremba and I. Sutskever · 2014
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Q. Le, M. Ranzato, R. Monga, M. Devin, K. Chen, G. Corrado, J. Dean, and A. Ng · 2012
Cited alongside, same era.
Application of pretrained deep neural networks to large vocabulary speech recognition
A. S. N. Jaitly, P. Nguyen and V. Vanhoucke · 2012
Cited alongside, same era.
On the difficulty of training recurrent neural networks
R. Pascanu, T. Mikolov, and Y. Bengio · 2012
Cited alongside, same era.
Deep learning with COTS HPC
A. Coates, B. Huval, T. Wang, D. J. Wu, A. Y. Ng, and B. Catanzaro · 2013
Cited alongside, same era.
Speech recognition with deep recurrent neural networks
A. Graves, A.-r. Mohamed, and G. Hinton · 2013
Cited alongside, same era.
Scalable modified Kneser-Ney language model estimation
K. Heafield, I. Pouzyrevsky, J. H. Clark, and P. Koehn · 2013
Cited alongside, same era.
Vocal tract length perturbation (VTLP) improves speech recognition
N. Jaitly and G. Hinton · 2013
Cited alongside, same era.
Neural machine translation by jointly learning to align and translate
D. Bahdanau, K. Cho, and Y. Bengio · 2015
Closest in time.
End-to-end attention-based large vocabulary speech recognition
D. Bahdanau, J. Chorowski, D. Serdyuk, P. Brakel, and Y. Bengio · 2015
Closest in time.
The third ’CHiME’ speech separation and recognition challenge: Dataset, task and baselines
J. Barker, E. Marxer, Ricard Vincent, and S. Watanabe · 2015
Closest in time.
W. Chan, N. Jaitly, Q. Le, and O. Vinyals · 2015
Closest in time.
End-to-end continuous speech recognition using attention-based recurrent nn: First results
J. Chorowski, D. Bahdanau, K. Cho, and Y. Bengio · 2015
Closest in time.
Optimizing RNN performance
E. Elsen · 2015
Closest in time.
An empirical exploration of recurrent network architectures
R. Jozefowicz, W. Zaremba, and I. Sutskever · 2015
Closest in time.
Audio augmentation for speech recognition
T. Ko, V. Peddinti, D. Povey, and S. Khudanpur · 2015
Closest in time.
Batch normalized recurrent neural networks
C. Laurent, G. Pereyra, P. Brakel, Y. Zhang, and Y. Bengio · 2015
Closest in time.
Lexicon-free conversational speech recognition with neural networks
A. Maas, Z. Xie, D. Jurafsky, and A. Ng · 2015
Closest in time.
EESEN: End-to-end speech recognition using deep rnn models and wfst-based decoding
Y. Miao, M. Gowayyed, and F. Metz · 2015
Closest in time.
Nervana GPU
Nervana Systems · 2015
Closest in time.
Librispeech: an asr corpus based on public domain audio books
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur · 2015
Closest in time.
Convolutional, long short-term memory, fully connected deep neural networks
T. Sainath, O. Vinyals, A. Senior, and H. Sak · 2015
Closest in time.
Fast and accurate recurrent neural network acoustic models for speech recognition
H. Sak, A. Senior, K. Rao, and F. Beaufays · 2015
Closest in time.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
C. Szegedy and S. Ioffe · 2015
Closest in time.
The ntt chime-3 system: Advances in speech enhancement and recognition for mobile multi-microphone devices
T. Yoshioka, N. Ito, M. Delcroix, A. Ogawa, K. Kinoshita, M. F. C. Yu, W. J. Fabian, M. Espi, T. Higuchi, S. Araki, and T. Nakatani · 2015
Closest in time.