Fetching the paper…
Reading the bibliography…
Online speech recognition is crucial for developing natural human-machine interfaces.
J. S. Garofolo, L. F. Lamel, W. M. Fisher, J. G. Fiscus, D. S. Pallett, and N. L. Dahlgren, “DARPA TIMIT Acoustic Phonetic Continuous Speech Corpus CDROM,” 1993
1993
Earlier work this paper cites.
M. Schuster and K. K. Paliwal, “Bidirectional recurrent neural networks,”
1997
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,”
1997
Earlier work this paper cites.
X. Huang, A. Acero, and H. Hon,
2001
Earlier work this paper cites.
X. Glorot and Y. Bengio, “Understanding the difficulty of training deep feedforward neural networks,” in
2010
Earlier work this paper cites.
D. Povey, A. Ghoshal, G. Boulianne, L. Burget, O. Glembek, N. Goel, M. Hannemann, P. Motlicek, Y. Qian, P. Schwarz, J. Silovsky, G. Stemmer, and K. Vesely, “The Kaldi Speech Recognition Toolkit,” in
2011
Earlier work this paper cites.
G. E. Dahl, D. Yu, L. Deng, and A. Acero, “Context-dependent pre-trained deep neural networks for large-vocabulary speech recognition,”
2012
Earlier work this paper cites.
A. Graves, N. Jaitly, and A. Mohamed, “Hybrid speech recognition with Deep Bidirectional LSTM,” in
2013
Earlier work this paper cites.
X. Lei, A. Senior, A. Gruenstein, and J. Sorensen, “Accurate and compact large vocabulary speech recognition on mobile devices.” in
2013
Earlier work this paper cites.
G. Chen, C. Parada, and G. Heigold, “Small-footprint keyword spotting using deep neural networks,” in
2014
Earlier work this paper cites.
M. Bacchiani, A. W. Senior, and G. Heigold, “Asynchronous, online, GMM-free training of a context dependent acoustic model for speech recognition,” in
2014
Earlier work this paper cites.
G. Chen, C. Parada, and G. Heigold, “Small-footprint keyword spotting using deep neural networks,” in
2014
Earlier work this paper cites.
K. Cho, B. van Merrienboer, D. Bahdanau, and Y. Bengio, “On the properties of neural machine translation: Encoder-decoder approaches,” in
2014
Earlier work this paper cites.
D. Yu and L. Deng,
2015
Earlier work this paper cites.
Y. Wang, J. Li, and Y. Gong, “Small-footprint high-performance deep neural network-based speech recognition using split-VQ,” in
2015
Earlier work this paper cites.
T. N. Sainath and C. Parada, “Convolutional neural networks for small-footprint keyword spotting,” in
2015
Earlier work this paper cites.
A. Mohamed, F. Seide, J. Yu, D.and Droppo, A. Stolcke, G. Zweig, and G. Penn, “Deep bi-directional recurrent networks over spectral windows,” in
2015
Cited alongside, same era.
V. Peddinti, D. Povey, and S. Khudanpur, “A time delay neural network architecture for efficient modeling of long temporal contexts,” in
2015
Cited alongside, same era.
M. Ravanelli and M. Omologo, “Contaminated speech training methods for robust DNN-HMM distant speech recognition,” in
2015
Cited alongside, same era.
M. Ravanelli, L. Cristoforetti, R. Gretter, M. Pellin, A. Sosi, and M. Omologo, “The DIRHA-English corpus and related tasks for distant-speech recognition in domestic environments,” in
2015
Cited alongside, same era.
J. Barker, R. Marxer, E. Vincent, and S. Watanabe, “The third CHiME Speech Separation and Recognition Challenge: Dataset, task and baselines,” in
2015
Y. Gal and Z. Ghahramani, “A theoretically grounded application of dropout in recurrent neural networks,” in
2016
Later among the works it cites.
C. Laurent, G. Pereyra, P. Brakel, Y. Zhang, and Y. Bengio, “Batch normalized recurrent neural networks,” in
2016
Later among the works it cites.
F. M. S. Watanabe, M. Delcroix and J. R. Hershey,
2017
Later among the works it cites.
L. Lu and S. Renals, “Small-footprint highway deep neural networks for speech recognition,”
2017
Later among the works it cites.
2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: An ASR corpus based on public domain audio books,” in
2015
Cited alongside, same era.
2015
Cited alongside, same era.
T. Moon, H. Choi, H. Lee, and I. Song, “RNNDROP: A novel dropout for RNNS in ASR,” in
2015
Cited alongside, same era.
I. Goodfellow, Y. Bengio, and A. Courville,
2016
Cited alongside, same era.
A. Zeyer, R. Schlüter, and H. Ney, “Towards online-recognition with deep bidirectional LSTM acoustic models,” in
2016
Cited alongside, same era.
K. Chen and Q. Huo, “Training deep bidirectional LSTM acoustic model for LVCSR by a context-sensitive-chunk BPTT approach,”
2016
Cited alongside, same era.
N. Jaitly, Q. V. Le, O. Vinyals, I. Sutskever, D. Sussillo, and S. Bengio, “An online sequence-to-sequence model using partial conditioning,” in
2016
Cited alongside, same era.
2017
Later among the works it cites.
A. Goyal, A. Sordoni, M.-A. Côté, N. Ke, and Y. Bengio, “Z-Forcing: Training stochastic recurrent networks,” in
2017
Later among the works it cites.
S. Shabanian, D. Arpit, A. Trischler, and Y. Bengio, “Variational Bi-LSTMs,”
2017
Later among the works it cites.
M. Ravanelli, P. Brakel, M. Omologo, and Y. Bengio, “Improving speech recognition by revising gated recurrent units,” in
2017
Later among the works it cites.
A. Paszke, S. Gross, S. Chintala, G. Chanan, E. Yang, Z. DeVito, Z. Lin, A. Desmaison, L. Antiga, and A. Lerer, “Automatic differentiation in PyTorch,” in
2017
Later among the works it cites.
V. Peddinti, Y. Wang, D. Povey, and S. Khudanpur, “Low Latency Acoustic Modeling Using Temporal Convolution and LSTMs,”
2018
Closest in time.
M. Ravanelli and M. Omologo, “Automatic context window composition for distant speech recognition,”
2018
Closest in time.
D. Serdyuk, N. R. Ke, A. Sordoni, A. Trischler, C. Pal, and Y. Bengio, “Twin networks: Matching the future for sequence generation,” in
2018
Closest in time.
K. Drossos, S. I. Mimilakis, D. Serdyuk, G. Schuller, T. Virtanen, and Y. Bengio, “MaD TwinNet: Masker-denoiser architecture with Twin Networks for monaural sound source separation,”
2018
Closest in time.
M. Ravanelli, P. Brakel, M. Omologo, and Y. Bengio, “Light gated recurrent units for speech recognition,”
2018
Closest in time.