Fetching the paper…
Reading the bibliography…
State-of-the-art speech recognition systems rely on fixed, hand-crafted features such as mel-filterbanks to preprocess the waveform before the training pipeline.
D. B. Paul and J. M. Baker, “The design for the wall street journal-based csr corpus,” in
1992
Earlier work this paper cites.
2013
Earlier work this paper cites.
Z. Tüske, P. Golik, R. Schlüter, and H. Ney, “Acoustic modeling with deep neural networks using raw time signal for lvcsr,” in
2014
Earlier work this paper cites.
J. Andén and S. Mallat, “Deep scattering spectrum,”
2014
Earlier work this paper cites.
N. Srivastava, G. E. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov, “Dropout: a simple way to prevent neural networks from overfitting.”
2014
Earlier work this paper cites.
A. Graves and N. Jaitly, “Towards end-to-end speech recognition with recurrent neural networks,” in
2014
Earlier work this paper cites.
Y. Hoshen, R. J. Weiss, and K. W. Wilson, “Speech acoustic modeling from raw multichannel waveforms,” in
2015
Earlier work this paper cites.
T. N. Sainath, R. J. Weiss, A. Senior, K. W. Wilson, and O. Vinyals, “Learning the speech front-end with raw waveform cldnns,” in
2015
Earlier work this paper cites.
Y. Miao, M. Gowayyed, and F. Metze, “Eesen: End-to-end speech recognition using deep RNN models and WFST-based decoding,” in
2015
Earlier work this paper cites.
Z. Zhu, J. H. Engel, and A. Hannun, “Learning multiscale features directly from waveforms,”
2016
Cited alongside, same era.
2016
Cited alongside, same era.
P. Ghahremani, V. Manohar, D. Povey, and S. Khudanpur, “Acoustic modelling from the signal domain using cnns.” 2016
2016
Cited alongside, same era.
2016
Cited alongside, same era.
2017
Later among the works it cites.
2017
Later among the works it cites.
2017
Later among the works it cites.
2017
Later among the works it cites.
S. Kim, T. Hori, and S. Watanabe, “Joint ctc-attention based end-to-end speech recognition using multi-task learning,” in
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2016
Cited alongside, same era.
2016
Cited alongside, same era.
T. Salimans and D. P. Kingma, “Weight normalization: A simple reparameterization to accelerate training of deep neural networks,” in
2016
Cited alongside, same era.
D. Amodei, S. Ananthanarayanan, R. Anubhai, J. Bai, E. Battenberg, C. Case, J. Casper, B. Catanzaro, Q. Cheng, G. Chen
2016
Cited alongside, same era.
2016
Cited alongside, same era.
2017
Later among the works it cites.
2017
Later among the works it cites.
——, “Attention-based wav2text with feature transfer learning,”
2017
Later among the works it cites.
“Gammatone-based spectrograms, using gammatone filterbanks or fourier transform weightings.” https://github.com/detly/gammatone, accessed: 2018-03-19
2018
Closest in time.