Fetching the paper…
Reading the bibliography…
We describe Microsoft's conversational speech recognition system, in which we combine recent developments in neural-network-based acoustic and language modeling to advance the state of the art on the Switchboard recognition task.
“Generalization of back-propagation to recurrent neural networks”,
F. J. Pineda, · 1987
Earlier work this paper cites.
“A learning algorithm for continually running fully recurrent neural networks”,
R. J. Williams and D. Zipser, · 1989
Earlier work this paper cites.
“Phoneme recognition using time-delay neural networks”,
A. Waibel, T. Hanazawa, G. Hinton, K. Shikano, and K. J. Lang, · 1989
Earlier work this paper cites.
“Backpropagation applied to handwritten zip code recognition”,
Y. LeCun, B. Boser, J. S. Denker, D. Henderson, R. E. Howard, W. Hubbard, and L. D. Jackel, · 1989
Earlier work this paper cites.
“Consonant recognition by modular construction of large phonemic time-delay neural networks”,
A. Waibel, H. Sawai, and K. Shikano, · 1989
Earlier work this paper cites.
“A recurrent error propagation network speech recognition system”,
T. Robinson and F. Fallside, · 1991
Earlier work this paper cites.
“Convolutional networks for images, speech, and time series”,
Y. LeCun and Y. Bengio, · 1995
Earlier work this paper cites.
“Long short-term memory”,
S. Hochreiter and J. Schmidhuber, · 1997
Earlier work this paper cites.
“The SRI March 2000 Hub-5 conversational speech transcription system”,
A. Stolcke et al., · 2000
Earlier work this paper cites.
“SRILM—an extensible language modeling toolkit”,
A. Stolcke, · 2002
Earlier work this paper cites.
“Framewise phoneme classification with bidirectional LSTM and other neural network architectures”,
A. Graves and J. Schmidhuber, · 2005
Earlier work this paper cites.
“Advances in speech transcription at IBM under the DARPA EARS program”,
S. F. Chen, B. Kingsbury, L. Mangu, D. Povey, G. Saon, H. Soltau, and G. Zweig, · 2006
Earlier work this paper cites.
“Recurrent neural network based language model”,
T. Mikolov, M. Karafiát, L. Burget, J. Cernockỳ, and S. Khudanpur, · 2010
Earlier work this paper cites.
“Noise-contrastive estimation: A new estimation principle for unnormalized statistical models”,
M. Gutmann and A. Hyvärinen, · 2010
Earlier work this paper cites.
“Front-end factor analysis for speaker verification”,
N. Dehak, P. J. Kenny, R. Dehak, P. Dumouchel, and P. Ouellet, · 2011
Earlier work this paper cites.
“Sequential classification criteria for NNs in automatic speech recognition”,
G. Wang and K. Sim, · 2011
Cited alongside, same era.
“Context dependent recurrent neural network language model”,
T. Mikolov and G. Zweig, · 2012
Cited alongside, same era.
“Efficient estimation of maximum entropy language models with N-gram features: An SRILM extension”,
T. Alumäe and M. Kurimo, · 2012
Cited alongside, same era.
“Applying convolutional neural networks concepts to hybrid NN-HMM model for speech recognition”,
O. Abdel-Hamid, A.-r. Mohamed, H. Jiang, and G. Penn, · 2012
Cited alongside, same era.
“Linguistic regularities in continuous space word representations”,
T. Mikolov, W.-t. Yih, and G. Zweig, · 2013
Cited alongside, same era.
“Speaker adaptation of neural network acoustic models using i-vectors”,
“Fast and accurate recurrent neural network acoustic models for speech recognition”,
H. Sak, A. Senior, K. Rao, and F. Beaufays, · 2015
Later among the works it cites.
“The IBM 2015 English conversational telephone speech recognition system”,
G. Saon, H.-K. J. Kuo, S. Rennie, and M. Picheny, · 2015
Later among the works it cites.
“Very deep convolutional neural networks for LVCSR”,
M. Bi, Y. Qian, and K. Yu, · 2015
Later among the works it cites.
“Deep residual learning for image recognition”,
K. He, X. Zhang, S. Ren, and J. Sun, · 2015
Later among the works it cites.
R. K. Srivastava, K. Greff, and J. Schmidhuber, · 2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
G. Saon, H. Soltau, D. Nahamoo, and M. Picheny, · 2013
Cited alongside, same era.
“Sequence-discriminative training of deep neural networks”,
K. Veselỳ, A. Ghoshal, L. Burget, and D. Povey, · 2013
Cited alongside, same era.
“Deep convolutional neural networks for LVCSR”,
T. N. Sainath, A.-r. Mohamed, B. Kingsbury, and B. Ramabhadran, · 2013
Cited alongside, same era.
“Long short-term memory recurrent neural network architectures for large scale acoustic modeling”,
H. Sak, A. W. Senior, and F. Beaufays, · 2014
Cited alongside, same era.
“Sequence to sequence learning with neural networks”,
I. Sutskever, O. Vinyals, and Q. V. Le, · 2014
Cited alongside, same era.
“Deep speech: Scaling up end-to-end speech recognition”,
A. Hannun, C. Case, J. Casper, B. Catanzaro, G. Diamos, E. Elsen, R. Prenger, S. Satheesh, S. Sengupta, A. Coates, et al., · 2014
Cited alongside, same era.
“Very deep convolutional networks for large-scale image recognition”,
K. Simonyan and A. Zisserman, · 2014
Cited alongside, same era.
T. Sercu, C. Puhrsch, B. Kingsbury, and Y. LeCun, · 2016
Closest in time.
“Very deep convolutional neural networks for noise robust speech recognition”,
Y. Qian, M. Bi, T. Tan, and K. Yu, · 2016
Closest in time.
“Deep convolutional neural networks with layer-wise context expansion and attention”,
D. Yu, W. Xiong, J. Droppo, A. Stolcke, G. Ye, J. Li, and G. Zweig, · 2016
Closest in time.
“Purely sequence-trained neural networks for ASR based on lattice-free MMI”,
D. Povey, V. Peddinti, D. Galvez, P. Ghahrmani, V. Manohar, X. Na, Y. Wang, and S. Khudanpur, · 2016
Closest in time.
“Linearly augmented deep neural network”,
P. Ghahremani, J. Droppo, and M. L. Seltzer, · 2016
Closest in time.
“The IBM 2016 English conversational telephone speech recognition system”,
G. Saon, T. Sercu, S. J. Rennie, and H. J. Kuo, · 2016
Closest in time.
“Parallelizing WFST speech decoders”,
C. Mendis, J. Droppo, S. Maleki, M. Musuvathi, T. Mytkowicz, and G. Zweig, · 2016
Closest in time.
“CUED-RNNLM: An open-source toolkit for efficient training and evaluation of recurrent neural network language models”,
X. Chen, X. Liu, Y. Qian, M. J. F. Gales, and P. C. Woodland, · 2016
Closest in time.
“Achieving human parity in conversational speech recognition”,
W. Xiong, J. Droppo, X. Huang, F. Seide, M. Seltzer, A. Stolcke, D. Yu, and G. Zweig, · 2016
Closest in time.