Fetching the paper…
Reading the bibliography…
Traditional automatic speech recognition (ASR) systems often use an acoustic model (AM) built on handcrafted acoustic features, such as log Mel-filter bank (FBANK) values.
“Zur theorie der orthogonalen funktionensysteme”,
A. Haar, · 1910
Earlier work this paper cites.
“Experiments in hearing”,
G. Von Békésy, E.G. Wever, · 1960
Earlier work this paper cites.
“Comparison of parametric representations for monosyllabic word recognition in continuously spoken sentences”,
S. Davis, P. Mermelstein, · 1980
Earlier work this paper cites.
“Phoneme recognition using time-delay neural networks”,
A. Waibel, T. Hanazawa, G. Hinton, K. Shikano, K.J. Lang, · 1989
Earlier work this paper cites.
“Perceptual linear predictive (PLP) analysis of speech”,
H. Hermansky, · 1990
Earlier work this paper cites.
The use of Recurrent Neural Networks in Continuous Speech Recognition
T. Robinson, M. Hochberg and S. Renals · 1996
Earlier work this paper cites.
“Gradient-based learning applied to document recognition”,
Y. LeCun, L. Bottou, Y. Bengio, P. Haffner, · 1998
Earlier work this paper cites.
“The AMI meeting corpus: A pre-announcement”,
J. Carletta, S. Ashby, S. Bourban, M. Flynn, M. Guillemot, T. Hain, J. Kadlec, V. Karaiskos, W. Kraaij, M. Kronenthal, · 2005
Earlier work this paper cites.
“Rectified linear units improve restricted Boltzmann machines”,
V. Nair, G.E. Hinton, · 2010
Earlier work this paper cites.
“Understanding how deep belief networks perform acoustic modelling”,
A.-R. Mohamed, G.E. Hinton, G. Penn, · 2012
Earlier work this paper cites.
“Speech recognition using long-span temporal patterns in a deep network model”,
S.M. Siniscalchi, D. Yu, L. Deng, C.-H. Lee, · 2013
Earlier work this paper cites.
“Acoustic modeling with deep neural networks using raw time signal for LVCSR”,
Z. Tüske, P. Golik, R. Schlüter, H. Ney, · 2014
Earlier work this paper cites.
“Convolutional neural networks for speech recognition”,
O. Abdel-Hamid, A.-R. Mohamed, H. Jiang, L. Deng, G. Penn, D. Yu, · 2014
Cited alongside, same era.
“Recurrent deep neural networks for robust speech recognition”,
C. Weng, D. Yu, S. Watanabe, B.-H.F. Juang, · 2014
Cited alongside, same era.
“Neural networks for distant speech recognition”,
S. Renals, P. Swietojanski, · 2014
Cited alongside, same era.
“Learning the speech front-end with raw waveform CLDNNs”,
T.N. Sainath, R.J. Weiss, A. Senior, K.W. Wilson, O. Vinyals, · 2015
Cited alongside, same era.
“Convolutional, long short-term memory, fully connected deep neural networks”,
T.N. Sainath, O. Vinyals, A. Senior, H. Sak, · 2015
Cited alongside, same era.
“Speaker location and microphone spacing invariant acoustic modeling from raw multichannel waveforms”,
“Learning multiscale features directly from waveforms”,
Z. Zhu, J.H. Engel, A.Y. Hannun, · 2016
Later among the works it cites.
“Complex linear projection (CLP): A discriminative approach to joint feature extraction and acoustic modeling”,
E. Variani, T.N. Sainath, I. Shafran, M. Bacchiani, · 2016
Later among the works it cites.
“The RWTH/UPB/FORTH system combination for the 4th CHiME challenge evaluation”,
T. Menne, J. Heymann, A. Alexandridis, K. Irie, A. Zeyer, M. Kitza, R. Schlüter, · 2016
Later among the works it cites.
“Robust features in deep-learning-based speech recognition”,
v. Mitra, F. Horacio, R.M. Stern, · 2017
Later among the works it cites.
“End-to-end speech recognition with auditory attention for multi-microphone distance speech recognition”,
S. Kim, I. Lane, · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
T.N. Sainath, R.J. Weiss, K.W. Wilson, A. Narayanan, M. Bacchiani, · 2015
Cited alongside, same era.
“Convolutional neural networks for acoustic modeling of raw time signal in LVCSR”,
P. Golik, Z. Tüske, R. Schlüter, H. Ney, · 2015
Cited alongside, same era.
“Very deep convolutional networks for large-scale image recognition”,
K. Simonyan, A. Zisserman, · 2015
Cited alongside, same era.
“Cambridge university transcription systems for the Multi-Genre Broadcast Challenge”,
P.C. Woodland, X. Liu, Y. Qian, C. Zhang, M.J.F. Gales, P. Karanasou, P. Lanchantin, & L. Wang, · 2015
Cited alongside, same era.
“Speech acoustic modeling from raw multichannel waveforms”,
Y. Hoshen, R.J. Weiss, K.W. Wilson, · 2015
Cited alongside, same era.
The HTK Book (for HTK version 3.5)
S. Young, G. Evermann, M. Gales, T. Hain, D. Kershaw, X. Liu, G. Moore, J. Odell, D. Ollason, D. Povey, A. Ragni, V. Valtchev, P. Woodland, & C. Zhang, · 2015
Cited alongside, same era.
“Acoustic modelling from the signal domain using CNNs”,
P. Ghahremani, V. Manohar, D. Povey, S. Khudanpur, · 2016
Cited alongside, same era.
C. Zhang, · 2017
Later among the works it cites.
“An analysis of environment, microphone and data simulation mismatches in robust speech recognition”,
E. Vincent, S. Watanabe, A.A. Nugraha, J. Barker, R. Marxer, · 2017
Later among the works it cites.
“Acoustic modeling of speech waveform based on multi-resolution, neural network signal processing”,
Z. Tüske, R. Schlüter H. Ney, · 2018
Later among the works it cites.
“Acoustic modeling from frequency-domain representations of speech”,
P. Ghahremani, H. Hadian, H. Lv, D. Povey, S. Khudanpur, · 2018
Later among the works it cites.
“Learning acoustic features from the raw waveform for automatic speech recognition”,
T. Menne, Z. Tüske, R. Schlüter, H. Ney, · 2018
Later among the works it cites.
“PyHTK: Python library and ASR pipelines for HTK”,
C. Zhang, F.L. Kreyssig, Q. Li, & P.C. Woodland, · 2019
Closest in time.