Fetching the paper…
Reading the bibliography…
We describe the 2017 version of Microsoft's conversational speech recognition system, in which we update our 2016 system with recent developments in neural-network-based acoustic and language modeling to further advance the state of the art on the Switchboard speech recognition task.
“Sequencing in conversational openings”,
E. A. Schegloff, · 1968
Earlier work this paper cites.
“Generalization of back-propagation to recurrent neural networks”,
F. J. Pineda, · 1987
Earlier work this paper cites.
“A learning algorithm for continually running fully recurrent neural networks”,
R. J. Williams and D. Zipser, · 1989
Earlier work this paper cites.
“Phoneme recognition using time-delay neural networks”,
A. Waibel, T. Hanazawa, G. Hinton, K. Shikano, and K. J. Lang, · 1989
Earlier work this paper cites.
“Backpropagation applied to handwritten zip code recognition”,
Y. LeCun, B. Boser, J. S. Denker, D. Henderson, R. E. Howard, W. Hubbard, and L. D. Jackel, · 1989
Earlier work this paper cites.
“Consonant recognition by modular construction of large phonemic time-delay neural networks”,
A. Waibel, H. Sawai, and K. Shikano, · 1989
Earlier work this paper cites.
“A recurrent error propagation network speech recognition system”,
T. Robinson and F. Fallside, · 1991
Earlier work this paper cites.
“Switchboard: Telephone speech corpus for research and development”,
J. J. Godfrey, E. C. Holliman, and J. McDaniel, · 1992
Earlier work this paper cites.
“Convolutional networks for images, speech, and time series”,
Y. LeCun and Y. Bengio, · 1995
Earlier work this paper cites.
“Insights into spoken language gleaned from phonetic transcription of the Switchboard corpus”,
S. Greenberg, J. Hollenback, and D. Ellis, · 1996
Earlier work this paper cites.
“Conceptual pacts and lexical choice in conversation”,
S. E. Brennan and H. H. Clark, · 1996
Earlier work this paper cites.
“Long short-term memory”,
S. Hochreiter and J. Schmidhuber, · 1997
Earlier work this paper cites.
“The SRI March 2000 Hub-5 conversational speech transcription system”,
A. Stolcke et al., · 2000
Earlier work this paper cites.
“Observations on overlap: Findings and implications for automatic processing of multi-party conversation”,
E. Shriberg, A. Stolcke, and D. Baron, · 2001
Earlier work this paper cites.
“Computing consensus translation from multiple machine translation systems”,
S. Bangalore, G. Bordel, and G. Riccardi, · 2001
Earlier work this paper cites.
“SRILM—an extensible language modeling toolkit”,
A. Stolcke, · 2002
Earlier work this paper cites.
“Getting more mileage from web text sources for conversational speech language modeling using class-dependent mixtures”,
I. Bulyko, M. Ostendorf, and A. Stolcke, · 2003
Earlier work this paper cites.
“Statistical language model adaptation: review and perspectives”,
J. R. Bellegarda, · 2004
Earlier work this paper cites.
“Multi-speaker language modeling”,
G. Ji and J. Bilmes, · 2004
Earlier work this paper cites.
“Framewise phoneme classification with bidirectional LSTM and other neural network architectures”,
A. Graves and J. Schmidhuber, · 2005
Cited alongside, same era.
“Neural probabilistic language models”,
Y. Bengio, H. Schwenk, J.-S. Senécal, F. Morin, and J.-L. Gauvain, · 2006
Cited alongside, same era.
“Advances in speech transcription at IBM under the DARPA EARS program”,
S. F. Chen, B. Kingsbury, L. Mangu, D. Povey, G. Saon, H. Soltau, and G. Zweig, · 2006
Cited alongside, same era.
“Recurrent neural network based language model”,
T. Mikolov, M. Karafiát, L. Burget, J. Cernockỳ, and S. Khudanpur, · 2010
Cited alongside, same era.
“Transcription methods for consistency, volume and efficiency”,
M. L. Glenn, S. Strassel, H. Lee, K. Maeda, R. Zakhary, and X. Li, · 2010
Cited alongside, same era.
“Conversational speech transcription using context-dependent deep neural networks”,
“Batch normalization: Accelerating deep network training by reducing internal covariate shift”,
S. Ioffe and C. Szegedy, · 2015
Later among the works it cites.
“Convolutional, long short-term memory, fully connected deep neural networks”,
T. N. Sainath, O. Vinyals, A. Senior, and H. Sak, · 2015
Later among the works it cites.
“From feedforward to recurrent LSTM neural networks for language modeling”,
M. Sundermeyer, H. Ney, and R. Schlüter, · 2015
Later among the works it cites.
“Adam: A method for stochastic optimization”,
D. P. Kingma and J. Ba, · 2015
Later among the works it cites.
“Very deep multilingual convolutional neural networks for LVCSR”,
T. Sercu, C. Puhrsch, B. Kingsbury, and Y. LeCun, · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
F. Seide, G. Li, and D. Yu, · 2011
Cited alongside, same era.
“Front-end factor analysis for speaker verification”,
N. Dehak, P. J. Kenny, R. Dehak, P. Dumouchel, and P. Ouellet, · 2011
Cited alongside, same era.
“Context dependent recurrent neural network language model”,
T. Mikolov and G. Zweig, · 2012
Cited alongside, same era.
“LSTM neural networks for language modeling”,
M. Sundermeyer, R. Schlüter, and H. Ney, · 2012
Cited alongside, same era.
“Linguistic regularities in continuous space word representations”,
T. Mikolov, W.-t. Yih, and G. Zweig, · 2013
Cited alongside, same era.
“Speaker adaptation of neural network acoustic models using i-vectors”,
G. Saon, H. Soltau, D. Nahamoo, and M. Picheny, · 2013
Cited alongside, same era.
“Long short-term memory recurrent neural network architectures for large scale acoustic modeling”,
H. Sak, A. W. Senior, and F. Beaufays, · 2014
Cited alongside, same era.
Y. Qian, M. Bi, T. Tan, and K. Yu, · 2016
Later among the works it cites.
“Improving English conversational telephone speech recognition”,
I. Medennikov, A. Prudnikov, and A. Zatvornitskiy, · 2016
Later among the works it cites.
“Achieving human parity in conversational speech recognition”,
W. Xiong, J. Droppo, X. Huang, F. Seide, M. Seltzer, A. Stolcke, D. Yu, and G. Zweig, · 2016
Later among the works it cites.
“Linearly augmented deep neural network”,
P. Ghahremani, J. Droppo, and M. L. Seltzer, · 2016
Later among the works it cites.
“Deep convolutional neural networks with layer-wise context expansion and attention”,
D. Yu, W. Xiong, J. Droppo, A. Stolcke, G. Ye, J. Li, and G. Zweig, · 2016
Later among the works it cites.
“The IBM 2016 English conversational telephone speech recognition system”,
G. Saon, T. Sercu, S. J. Rennie, and H. J. Kuo, · 2016
Later among the works it cites.
“Purely sequence-trained neural networks for ASR based on lattice-free MMI”,
D. Povey, V. Peddinti, D. Galvez, P. Ghahrmani, V. Manohar, X. Na, Y. Wang, and S. Khudanpur, · 2016
Later among the works it cites.
“Self-stabilized deep neural network”,
P. Ghahremani and J. Droppo, · 2016
Later among the works it cites.
“Using the output embedding to improve language models”,
O. Press and L. Wolf, · 2016
Later among the works it cites.
“Comparing human and machine errors in conversational speech transcription”,
A. Stolcke and J. Droppo, · 2017
Closest in time.
“English conversational telephone speech recognition by humans and machines”,
G. Saon, G. Kurata, T. Sercu, K. Audhkhasi, S. Thomas, D. Dimitriadis, X. Cui, B. Ramabhadran, M. Picheny, L.-L. Lim, B. Roomi, and P. Hall, · 2017
Closest in time.
“Deep learning-based telephony speech recognition in the wild”,
K. J. Han, S. Hahm, B.-H. Kim, J. Kim, and I. Lane, · 2017
Closest in time.
“The Microsoft 2016 conversational speech recognition system”,
W. Xiong, J. Droppo, X. Huang, F. Seide, M. Seltzer, A. Stolcke, D. Yu, and G. Zweig, · 2017
Closest in time.