Fetching the paper…
Reading the bibliography…
Deep learning is currently playing a crucial role toward higher levels of artificial intelligence.
The design for the wall street journal-based csr corpus
Douglas P. and J. M. Baker · 1992
Earlier work this paper cites.
DARPA TIMIT Acoustic Phonetic Continuous Speech Corpus CDROM, 1993
J. S. Garofolo, L. F. Lamel, W. M. Fisher, J. G. Fiscus, D. S. Pallett, and N. L. Dahlgren · 1993
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Object recognition with gradient-based learning
Y. LeCun, P. Haffner, L. Bottou, and Y. Bengio · 1999
Earlier work this paper cites.
Digital Signal Processing
S. K. Mitra · 2005
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
X. Glorot and Y. Bengio · 2010
Earlier work this paper cites.
Theory and Applications of Digital Speech Processing
L. R. Rabiner and R. W. Schafer · 2011
Earlier work this paper cites.
i-vector based speaker recognition on short utterances
A. Kanagasundaram, R. Vogt, D. Dean, S. Sridharan, and M. Mason · 2011
Earlier work this paper cites.
The Kaldi Speech Recognition Toolkit
D. Povey et al · 2011
Earlier work this paper cites.
Study of the effect of i-vector modeling on short and mismatch utterance duration for speaker verification
A. K. Sarkar, D Matrouf, P.M. Bousquet, and J.F. Bonastre · 2012
Earlier work this paper cites.
Impulse response estimation for robust speech recognition in a reverberant environment
M. Ravanelli, A. Sosi, P. Svaizer, and M. Omologo · 2012
Earlier work this paper cites.
Learning filter banks within a deep neural network framework
T. N. Sainath, B. Kingsbury, A. R. Mohamed, and B. Ramabhadran · 2013
Earlier work this paper cites.
Rectifier nonlinearities improve neural network acoustic models
A. L. Maas, A. Y. Hannun, and A. Y. Ng · 2013
Earlier work this paper cites.
Towards end-to-end speech recognition with recurrent neural networks
A. Graves and N. Jaitly · 2014
Earlier work this paper cites.
Visualizing and understanding convolutional networks
M. D. Zeiler and R. Fergus · 2014
Earlier work this paper cites.
Empirical evaluation of gated recurrent neural networks on sequence modeling
J. Chung, Ç. Gülçehre, K. Cho, and Y. Bengio · 2014
Earlier work this paper cites.
Deep neural networks for small footprint text-dependent speaker verification
E. Variani, X. Lei, E. McDermott, I. L. Moreno, and J. Gonzalez-Dominguez · 2014
Earlier work this paper cites.
Acoustic modeling with deep neural networks using raw time signal for LVCSR
Z. Tüske, P. Golik, R. Schlüter, and H. Ney · 2014
Earlier work this paper cites.
Modified-prior i-Vector Estimation for Language Identification of Short Duration Utterances
R. Travadi, M. Van Segbroeck, and S. Narayanan · 2014
Earlier work this paper cites.
The DIRHA-GRID corpus: baseline and tools for multi-room distant speech recognition using distributed microphones
M. Matassoni, R. Astudillo, A. Katsamanis, and M. Ravanelli · 2014
Earlier work this paper cites.
The DIRHA simulated corpus
L. Cristoforetti, M. Ravanelli, M. Omologo, A. Sosi, A. Abad, M. Hagmueller, and P. Maragos · 2014
Earlier work this paper cites.
On the selection of the impulse responses for distant-speech recognition based on contaminated speech training
M. Ravanelli and M. Omologo · 2014
Earlier work this paper cites.
Automatic Speech Recognition - A Deep Learning Approach
D. Yu and L. Deng · 2015
Earlier work this paper cites.
Explaining and harnessing adversarial examples
I. Goodfellow, J. Shlens, and C. Szegedy · 2015
Earlier work this paper cites.
A unified deep neural network for speaker and language recognition
F. Richardson, D. A. Reynolds, and N. Dehak · 2015
Earlier work this paper cites.
Analysis of CNN-based speech recognition system using raw speech as input
D. Palaz, M. Magimai-Doss, and R. Collobert · 2015
Cited alongside, same era.
Learning the speech front-end with raw waveform CLDNNs
T. N. Sainath, R. J. Weiss, A. W. Senior, K. W. Wilson, and O. Vinyals · 2015
Cited alongside, same era.
Speech acoustic modeling from raw multichannel waveforms
Y. Hoshen, R. Weiss, and K. W. Wilson · 2015
Cited alongside, same era.
Speaker localization and microphone spacing invariant acoustic modeling from raw multichannel waveforms
T. N. Sainath, R. J. Weiss, K. W. Wilson, A. Narayanan, M. Bacchiani, and A. Senior · 2015
Cited alongside, same era.
Librispeech: An ASR corpus based on public domain audio books
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur · 2015
Cited alongside, same era.
The DIRHA-ENGLISH corpus and related tasks for distant-speech recognition in domestic environments
Improving speech recognition by revising gated recurrent units
M. Ravanelli, P. Brakel, M. Omologo, and Y. Bengio · 2017
Later among the works it cites.
Deep neural network embeddings for text-independent speaker verification
D. Snyder, D. Garcia-Romero, D. Povey, and S. Khudanpur · 2017
Later among the works it cites.
Deep speaker embeddings for short-duration speaker verification
G. Bhattacharya, J. Alam, and P. Kenny · 2017
Later among the works it cites.
Voxceleb: a large-scale speaker identification dataset
A. Nagrani, J. S. Chung, and A. Zisserman · 2017
Later among the works it cites.
End-to-end spoofing detection with raw waveform CLDNNS
H. Dinkel, N. Chen, Y. Qian, and K. Yu · 2017
Later among the works it cites.
DNN Filter Bank Cepstral Coefficients for Spoofing Detection
H. Yu, Z. H. Tan, Y. Zhang, Z. Ma, and J. Guo · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M. Ravanelli, L. Cristoforetti, R. Gretter, M. Pellin, A. Sosi, and M. Omologo · 2015
Cited alongside, same era.
Contaminated speech training methods for robust DNN-HMM distant speech recognition
M. Ravanelli and M. Omologo · 2015
Cited alongside, same era.
A multi-channel corpus for distant-speech interaction in presence of known interferences
E. Zwyssig, M. Ravanelli, P. Svaizer, and M. Omologo · 2015
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
S. Ioffe and C. Szegedy · 2015
Cited alongside, same era.
Deep Learning
I. Goodfellow, Y. Bengio, and A. Courville · 2016
Cited alongside, same era.
End-to-end attention-based large vocabulary speech recognition
D. Bahdanau, J. Chorowski, D. Serdyuk, P. Brakel, and Y. Bengio · 2016
Cited alongside, same era.
"why should i trust you?": Explaining the predictions of any classifier
M. T. Ribeiro, S. Singh, and C. Guestrin · 2016
Cited alongside, same era.
A deep neural network integrated with filterbank learning for speech recognition
H. Seki, K. Yamamoto, and S. Nakagawa · 2017
Later among the works it cites.
Convolutional neural networks analyzed via convolutional sparse coding
V. Papyan, Y. Romano, and M. Elad · 2017
Later among the works it cites.
Deep learning for Distant Speech Recognition
M. Ravanelli · 2017
Later among the works it cites.
A network of deep neural networks for distant speech recognition
M. Ravanelli, P. Brakel, M. Omologo, and Y. Bengio · 2017
Later among the works it cites.
Interpretable Machine Learning: A Guide for Making Black Box Models Explainable
C. Molnar · 2018
Closest in time.
Visual interpretability for deep learning: a survey
Q.-S. Zhang and S.-C. Zhu · 2018
Closest in time.
Interpreting CNN Knowledge via an Explanatory Graph
Q. Zhang, R. Cao, F. Shi, Y. N. Wu, and S.-C. Zhu · 2018
Closest in time.
Interpreting and explaining deep neural networks for classification of audio signals
S. Becker, M. Ackermann, S. Lapuschkin, K.-R. Müller, and W. Samek · 2018
Closest in time.
Light gated recurrent units for speech recognition
M. Ravanelli, P. Brakel, M. Omologo, and Y. Bengio · 2018
Closest in time.
Text-independent speaker verification based on triplet convolutional neural network embeddings
C. Zhang, K. Koishida, and J. Hansen · 2018
Closest in time.
Towards directly modeling raw speech signal for speaker verification using CNNs
H. Muckenhirn, M. Magimai-Doss, and S. Marcel · 2018
Closest in time.
A complete end-to-end speaker verification system using deep neural networks: From raw signals to verification result
J.-W. Jung, H.-S. Heo, I.-H. Yang, H.-J. Shim, , and H.-J. Yu · 2018
Closest in time.
Avoiding Speaker Overfitting in End-to-End DNNs using Raw Waveform for Text-Independent Speaker Verification
J.-W. Jung, H.-S. Heo, I.-H. Yang, H.-J. Shim, and H.-J. Yu · 2018
Closest in time.
Speaker Recognition from raw waveform with SincNet
M. Ravanelli and Y. Bengio · 2018
Closest in time.
Twin regularization for online speech recognition
M. Ravanelli, D. Serdyuk, and Y. Bengio · 2018
Closest in time.
Learning filterbanks from raw speech for phone recognition
N. Zeghidour, N. Usunier, I. Kokkinos, T. Schatz, G. Synnaeve, and E. Dupoux · 2018
Closest in time.
On Learning Vocal Tract System Related Speaker Discriminative Information from Raw Signal Using CNNs
H. Muckenhirn, M. Magimai-Doss, and S. Marcel · 2018
Closest in time.
The PyTorch-Kaldi Speech Recognition Toolkit
M. Ravanelli, T. Parcollet, and Y. Bengio · 2018
Closest in time.