Fetching the paper…
Reading the bibliography…
It is well known that a speech recognition system that combines multiple acoustic models trained on the same data significantly outperforms a single-model system.
L. I. Kuncheva and C. J. Whitaker, “Measures of diversity in classifier ensembles and their relationship with the ensemble accuracy,”
2003
Earlier work this paper cites.
R. E. Banfield, L. O. Hall, K. W. Bowyer, and W. P. Kegelmeyer, “Ensemble diversity measures and their application to thinning,”
2005
Earlier work this paper cites.
A. R. Mohamed, G. E. Dahl, and G. Hinton, “Acoustic modeling using deep belief networks,”
2011
Earlier work this paper cites.
G. E. Dahl and A. Acero, “Context-dependent pre-trained deep neural networks for large-vocabulary speech recognition,”
2011
Earlier work this paper cites.
D. Povey, A. Ghoshal, G. Boulianne, L. Burget, O. Glembek, N. Goel, M. Hannemann, P. Motlicek, Y. Qian, P. Schwarz
2011
Earlier work this paper cites.
K. Kinoshita, M. Delcroix, T. Yoshioka, T. Nakatani, A. Sehr, W. Kellermann, and R. Maas, “The reverb challenge: A common evaluation framework for dereverberation and recognition of reverberant speech,” in
2013
Earlier work this paper cites.
N. Jaitly and G. E. Hinton, “Vocal tract length perturbation (vtlp) improves speech recognition,” in
2013
Earlier work this paper cites.
T. N. Sainath, B. Kingsbury, A. R. Mohamed, and G. E. Dahl, “Improvements to deep convolutional neural networks for lvcsr,” 2013
2013
Earlier work this paper cites.
A. Ragni, K. M. Knill, S. P. Rath, and M. J. Gales, “Data augmentation for low resource languages,” 2014
2014
Cited alongside, same era.
L. Deng and J. C. Platt, “Ensemble deep learning for speech recognition,” in
2014
Cited alongside, same era.
X. Lu, Y. Tsao, S. Matsuda, and C. Hori, “Ensemble modeling of denoising autoencoder for speech spectrum restoration,” in
2014
Cited alongside, same era.
H. Sak, A. Senior, and F. Beaufays, “Long short-term memory based recurrent neural network architectures for large vocabulary speech recognition,”
2014
Cited alongside, same era.
H. Soltau, G. Saon, and T. N. Sainath, “Joint training of convolutional and non-convolutional neural networks,” in
2014
Cited alongside, same era.
G. Hinton, O. Vinyals, and J. Dean, “Distilling the knowledge in a neural network,”
2015
Later among the works it cites.
X. Cui, V. Goel, and B. Kingsbury, “Data augmentation for deep neural network acoustic modeling,”
2015
Later among the works it cites.
Y. Chebotar and A. Waters, “Distilling knowledge from ensembles of neural networks for speech recognition.” in
2016
Later among the works it cites.
T. Fukuda, M. Suzuki, G. Kurata, S. Thomas, J. Cui, and B. Ramabhadran, “Efficient knowledge distillation from an ensemble of teachers.” in
2017
Later among the works it cites.
G. Cheng, V. Peddinti, D. Povey, V. Manohar, S. Khudanpur, and Y. Yan, “An exploration of dropout with lstms.” in
2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
D. Povey, X. Zhang, and S. Khudanpur, “Parallel training of deep neural networks with natural gradient and parameter averaging,”
2014
Cited alongside, same era.
T. Ko, V. Peddinti, D. Povey, and S. Khudanpur, “Audio augmentation for speech recognition,” in
2015
Cited alongside, same era.
T. N. Sainath, O. Vinyals, A. Senior, and H. Sak, “Convolutional, long short-term memory, fully connected deep neural networks,” in
2015
Cited alongside, same era.
2018
Later among the works it cites.
N. Kanda, Y. Fujita, and K. Nagamatsu, “Sequence distillation for purely sequence trained acoustic models,” in
2018
Later among the works it cites.