Fetching the paper…
Reading the bibliography…
The availability of open-source software is playing a remarkable role in the popularization of speech recognition and deep learning.
“Finite-state transducers in language and speech processing,”
M. Mohri, · 1997
Earlier work this paper cites.
“Maximum Likelihood Linear Transformations for HMM-Based Speech Recognition,”
M.J.F. Gales, · 1998
Earlier work this paper cites.
HTK – Hidden Markov Model Toolkit
S. Young et al., · 2006
Earlier work this paper cites.
“The lia speech recognition system: From 10xrt to 1xrt,”
G. Linarès, P. Nocera, D. Massonié, and D. Matrouf, · 2007
Earlier work this paper cites.
“Interpretation of Multiparty Meetings the AMI and Amida Projects,”
S. Renals, T. Hain, and H. Bourlard, · 2008
Earlier work this paper cites.
“Recent development of open-source speech recognition engine julius,”
A. Lee and T. Kawahara., · 2008
Earlier work this paper cites.
“Understanding the difficulty of training deep feedforward neural networks,”
X. Glorot and Y. Bengio, · 2010
Earlier work this paper cites.
“RASR - The RWTH Aachen University Open Source Speech Recognition Toolkit,”
D. Rybach, S. Hahn, P. Lehnen, D. Nolden, M. Sundermeyer, Z. Tüske, S. Wiesler, R. Schlüter, and H. Ney, · 2011
Earlier work this paper cites.
“The Kaldi Speech Recognition Toolkit,”
D. Povey et al., · 2011
Earlier work this paper cites.
“Random search for hyper-parameter optimization,”
J. Bergstra and Y. Bengio, · 2012
Earlier work this paper cites.
“The DIRHA simulated corpus,”
L. Cristoforetti, M. Ravanelli, M. Omologo, A. Sosi, A. Abad, M. Hagmueller, and P. Maragos, · 2014
Earlier work this paper cites.
“Dropout: A simple way to prevent neural networks from overfitting,”
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov, · 2014
Earlier work this paper cites.
“Joint noise adaptive training for robust automatic speech recognition,”
A. Narayanan and D. Wang, · 2014
Earlier work this paper cites.
Automatic Speech Recognition – A Deep Learning Approach
D. Yu and L. Deng, · 2015
Earlier work this paper cites.
“The third CHiME Speech Separation and Recognition Challenge: Dataset, task and baselines,”
J. Barker, R. Marxer, E. Vincent, and S. Watanabe, · 2015
Cited alongside, same era.
“Librispeech: An ASR corpus based on public domain audio books,”
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, · 2015
Cited alongside, same era.
“Batch normalization: Accelerating deep network training by reducing internal covariate shift,”
S. Ioffe and C. Szegedy, · 2015
Cited alongside, same era.
“Adam: A method for stochastic optimization,”
D.P. Kingma and J. Ba, · 2015
Cited alongside, same era.
“Contaminated speech training methods for robust DNN-HMM distant speech recognition,”
M. Ravanelli and M. Omologo, · 2015
Cited alongside, same era.
“The DIRHA-English corpus and related tasks for distant-speech recognition in domestic environments,”
“Realistic multi-microphone data simulation for distant speech recognition,”
M. Ravanelli, P. Svaizer, and M. Omologo, · 2016
Later among the works it cites.
Deep learning for Distant Speech Recognition
M. Ravanelli, · 2017
Later among the works it cites.
“Automatic differentiation in pytorch,”
A. Paszke et al., · 2017
Later among the works it cites.
“Hybrid ctc/attention architecture for end-to-end speech recognition,”
S. Watanabe, T. Hori, S. Kim, J. R. Hershey, and T. Hayashi, · 2017
Later among the works it cites.
“Multitask learning of context-dependent targets in deep neural network acoustic models,”
P. Bell, P. Swietojanski, and S. Renals, · 2017
Later among the works it cites.
“A network of deep neural networks for distant speech recognition,”
M. Ravanelli, P. Brakel, M. Omologo, and Y. Bengio, · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M. Ravanelli, L. Cristoforetti, R. Gretter, M. Pellin, A. Sosi, and M. Omologo, · 2015
Cited alongside, same era.
“A simple way to initialize recurrent networks of rectified linear units,”
Q.V. Le, N. Jaitly, and G.E. Hinton, · 2015
Cited alongside, same era.
“RNNDROP: A novel dropout for RNNS in ASR,”
T. Moon, H. Choi, H. Lee, and I. Song, · 2015
Cited alongside, same era.
Deep Learning
I. Goodfellow, Y. Bengio, and A. Courville, · 2016
Cited alongside, same era.
“Theano: A Python framework for fast computation of mathematical expressions,”
Theano Development Team, · 2016
Cited alongside, same era.
“Tensorflow: A system for large-scale machine learning,”
M. Abadi et al., · 2016
Cited alongside, same era.
“CNTK: Microsoft’s Open-Source Deep-Learning Toolkit,”
F. Seide and A. Agarwal, · 2016
Cited alongside, same era.
“Improving speech recognition by revising gated recurrent units,”
M. Ravanelli, P. Brakel, M. Omologo, and Y. Bengio, · 2017
Later among the works it cites.
“Pykaldi: A python wrapper for kaldi,”
D. Can, V. R. Martinez, P. Papadopoulos, and S. S. Narayanan, · 2018
Closest in time.
“Automatic context window composition for distant speech recognition,”
M. Ravanelli and M. Omologo, · 2018
Closest in time.
“Light gated recurrent units for speech recognition,”
M. Ravanelli, P. Brakel, M. Omologo, and Y. Bengio, · 2018
Closest in time.
“Twin regularization for online speech recognition,”
M. Ravanelli, D. Serdyuk, and Y. Bengio, · 2018
Closest in time.
“Speaker Recognition from raw waveform with SincNet,”
M. Ravanelli and Y.Bengio, · 2018
Closest in time.
“Interpretable Convolutional Filters with SincNet,”
M. Ravanelli and Y.Bengio, · 2018
Closest in time.