Fetching the paper…
Reading the bibliography…
Most phoneme recognition state-of-the-art systems rely on a classical neural network classifiers, fed with highly tuned features, such as MFCC or PLP features.
S. Davis and P. Mermelstein, “Comparison of parametric representations for monosyllabic word recognition in continuously spoken sentences,” Acoustics, Speech and Signal Processing, IEEE Transactions on , vol. 28, no. 4, pp. 357–366, 1980
1980
Earlier work this paper cites.
Y. LeCun, “Generalization and network design strategies,” in Connectionism in Perspective , R. Pfeifer, Z. Schreter, F. Fogelman, and L. Steels, Eds. Zurich, Switzerland: Elsevier, 1989
1989
Earlier work this paper cites.
A. Waibel, T. Hanazawa, G. Hinton, K. Shikano, and K. Lang, “Phoneme recognition using time-delay neural networks,” Acoustics, Speech and Signal Processing, IEEE Transactions on , vol. 37, no. 3, pp. 328 –339, mar 1989
1989
Earlier work this paper cites.
L. Bottou, F. Fogelman Soulié, P. Blanchet, and J. S. Lienard, “Experiments with time delay networks and dynamic time warping for speaker independent isolated digit recognition,” in Proceedings of EuroSpeech 89 , vol. 2, Paris, France, 1989, pp. 537–540
1989
Earlier work this paper cites.
K. F. Lee and H. W. Hon, “Speaker-independent phone recognition using hidden markov models,” IEEE Transactions on Acoustics, Speech and Signal Processing , vol. 37, no. 11, pp. 1641–1648, 1989
1989
Earlier work this paper cites.
H. Bourlard and C. Wellekens, “Links between markov models and multilayer perceptrons,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 12, no. 12, pp. 1167 –1178, Dec. 1990
1990
Earlier work this paper cites.
H. Hermansky, “Perceptual linear predictive (plp) analysis of speech,” The Journal of the Acoustical Society of America , vol. 87, p. 1738, 1990
1990
Earlier work this paper cites.
L. Bottou, “Stochastic gradient learning in neural networks,” in Proceedings of Neuro-Nîmes 91 . Nimes, France: EC2, 1991
1991
Earlier work this paper cites.
Y. Bengio, “A connectionist approach to speech recognition,” International Journal on Pattern Recognition and Artificial Intelligence , vol. 7, no. 4, pp. 647–668, 1993
1993
Earlier work this paper cites.
P. Woodland, J. Odell, V. Valtchev, and S. Young, “Large vocabulary continuous speech recognition using htk,” in Proc. of ICASSP , vol. ii, apr 1994, pp. II/125–II/128 vol.2
1994
Earlier work this paper cites.
N. Morgan and H. Bourlard, “Continuous speech recognition,” Signal Processing Magazine, IEEE , vol. 12, no. 3, pp. 24 –42, May 1995
1995
Cited alongside, same era.
L. Bottou, Y. Bengio, and Y. LeCun, “Global training of document processing systems using graph transformer networks.” in In Proc. of Computer Vision and Pattern Recognition . Puerto-Rico., 1997, pp. 490–494
1997
Cited alongside, same era.
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner, “Gradient-based learning applied to document recognition,” Proceedings of the IEEE , vol. 86, no. 11, pp. 2278–2324, 1998
1998
Cited alongside, same era.
S. Young, G. Evermann, D. Kershaw, G. Moore, J. Odell, D. Ollason, V. Valtchev, and P. Woodland, “The htk book,” Cambridge University Engineering Department , vol. 3, 2002
2002
Cited alongside, same era.
F. Seide, G. Li, and D. Yu, “Conversational speech transcription using context-dependent deep neural networks,” in Proc. Interspeech , 2011, pp. 437–440
2011
Later among the works it cites.
N. Jaitly and G. Hinton, “Learning a better representation of speech soundwaves using restricted boltzmann machines,” in Proc. of ICASSP , 2011, pp. 5884–5887
2011
Later among the works it cites.
R. Collobert, K. Kavukcuoglu, and C. Farabet, “Torch7: A matlab-like environment for machine learning,” in BigLearn, NIPS Workshop , 2011
2011
Later among the works it cites.
A. Mohamed, G. Dahl, and G. Hinton, “Acoustic modeling using deep belief networks,” Audio, Speech, and Language Processing, IEEE Transactions on , vol. 20, no. 1, pp. 14 –22, jan. 2012
2012
Later among the works it cites.
A. Krizhevsky, I. Sutskever, and G. Hinton, “Imagenet classification with deep convolutional neural networks,” in Advances in Neural Information Processing Systems 25 , 2012, pp. 1106–1114
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. LeCun, F. J. Huang, and L. Bottou, “Learning methods for generic object recognition with invariance to pose and lighting,” in Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition , vol. 2, 2004, pp. II–97
2004
Cited alongside, same era.
G. E. Hinton, S. Osindero, and Y. W. Teh, “A fast learning algorithm for deep belief nets,” Neural computation , vol. 18, no. 7, pp. 1527–1554, 2006
2006
Cited alongside, same era.
R. Collobert and J. Weston, “A unified architecture for natural language processing: deep neural networks with multitask learning,” in Proceedings of the 25th international conference on Machine learning , 2008, pp. 160–167
2008
Cited alongside, same era.
A. Mohamed, G. Dahl, and G. Hinton, “Deep belief networks for phone recognition,” in NIPS Workshop on Deep Learning for Speech Recognition and Related Applications , 2009
2009
Cited alongside, same era.
H. Lee, P. Pham, Y. Largman, and A. Y. Ng, “Unsupervised feature learning for audio classification using convolutional deep belief networks,” in Advances in Neural Information Processing Systems 22 , 2009, pp. 1096–1104
2009
Cited alongside, same era.
R. Collobert, J. Weston, L. Bottou, M. Karlen, K. Kavukcuoglu, and P. Kuksa, “Natural language processing (almost) from scratch,” The Journal of Machine Learning Research , vol. 12, pp. 2493–2537, 2011
2011
Cited alongside, same era.
2012
Later among the works it cites.
G. E. Dahl, D. Yu, L. Deng, and A. Acero, “Context-dependent pre-trained deep neural networks for large-vocabulary speech recognition,” Audio, Speech, and Language Processing, IEEE Transactions on , vol. 20, no. 1, p. 30–42, 2012
2012
Later among the works it cites.
G. Hinton, L. Deng, D. Yu, G. E. Dahl, A.-r. Mohamed, N. Jaitly, A. Senior, V. Vanhoucke, P. Nguyen, and T. N. Sainath, “Deep neural networks for acoustic modeling in speech recognition: the shared views of four research groups,” Signal Processing Magazine, IEEE , vol. 29, no. 6, p. 82–97, 2012
2012
Later among the works it cites.
O. Abdel-Hamid, A.-r. Mohamed, H. Jiang, and G. Penn, “Applying convolutional neural networks concepts to hybrid NN-HMM model for speech recognition,” in Proc. of ICASSP , 2012, pp. 4277–4280
2012
Later among the works it cites.
D. Palaz, R. Collobert, and M. Magimai.-Doss, “Estimating phoneme class conditional probabilities from raw speech signal using convolutional neural networks,” in Proceedings of Interspeech , Aug. 2013
2013
Closest in time.