Fetching the paper…
Reading the bibliography…
Speech recognizers trained on close-talking speech do not generalize to distant speech and the word error rate degradation can be as large as 40% absolute.
J. B. Allen and D. A. Berkley, “Image method for efficiently simulating small-room acoustics,”
1979
Earlier work this paper cites.
Y. Ephraim and D. Malah, “Speech enhancement using a minimum mean-square error short-time spectral amplitude estimator,”
1984
Earlier work this paper cites.
B. B. Paul and E. A. Martin, “Speaker stress-resistant continuous speech recognition,” in
1988
Earlier work this paper cites.
J.-L. Gauvain and C.-H. Lee, “Maximum a posteriori estimation for multivariate Gaussian mixture observations of Markov chains,”
1994
Earlier work this paper cites.
C. J. Leggetter and P. C. Woodland, “Maximum likelihood linear regression for speaker adaptation of continuous density hidden Markov models,”
1995
Earlier work this paper cites.
J. Neto, L. Almeida, M. Hochberg, C. Martins, L. Nunes, S. Renals, and T. Robinson, “Speaker-adaptation for hybrid HMM-ANN continuous speech recognition system,” in
1995
Earlier work this paper cites.
R. M. Warren, K. R. Hainsworth, B. S. Brubaker, J. A. Bashford, and E. W. Healy, “Spectral restoration of speech: intelligibility is increased by inserting noise in spectral gaps,”
1997
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,”
1997
Earlier work this paper cites.
E. A. Wan and A. T. Nelson, “Networks for speech enhancement,” in
1998
Earlier work this paper cites.
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner, “Gradient-based learning applied to document recognition,”
1998
Earlier work this paper cites.
H. Hermansky, D. P. Ellis, and S. Sharma, “Tandem connectionist feature extraction for conventional HMM systems,” in
2000
Earlier work this paper cites.
P. Y. Simard, D. Steinkraus, and J. C. Platt, “Best practices for convolutional neural networks applied to visual document analysis,” in
2003
Earlier work this paper cites.
M. Wölfel and J. McDonough,
2009
Earlier work this paper cites.
S. Den-David, J. Blitzer, K. Crammer, A. Kulesza, F. Pereira, and J. W. Vaughan, “A theory of learning from different domains,”
2010
Cited alongside, same era.
S. Ben-David, T. Luu, T. Lu, and D. Pál, “Impossibility theorems for domain adaptation,” in
2010
Cited alongside, same era.
B. Li and K. C. Sim, “Comparison of discriminative input and output transformations for speaker adaptation in the hybrid NN/HMM systems,” in
2010
Cited alongside, same era.
T. Yoshioka, A. Sehr, M. Delcroix, K. Kinoshita, R. Maas, T. Nakatani, and W. Kellermann, “Making machines understand us in reverberant rooms,”
2012
Cited alongside, same era.
A. L. Maas, Q. V. Le, T. M. O’Neil, O. Vinyals, P. Nguyen, and A. Y. Ng, “Recurrent neural networks for noise reduction in robust ASR,” in
2012
Cited alongside, same era.
I. Himawan, P. Motlicek, D. Imseng, B. Potard, N. Kim, and J. Lee, “Learning feature mapping using deep neural network bottleneck features for distant large vocabulary speech recognition,” in
2015
Later among the works it cites.
Y. Ganin, E. Ustinova, H. Ajakan, P. Germanin, H. Larochelle, F. Laviolette, M. Marchand, and V. Lempitsky, “Domain-adversarial training of neural networks,”
2016
Later among the works it cites.
J. Du, Y.-H. Tu, L. Sun, F. Ma, H.-K. Wang, J. Pan, C. Liu, J.-D. Chen, and C.-H. Lee, “The USTC-iFlytek system for CHiME-4 challenge,”
2016
Later among the works it cites.
L. D. Jahn Heymann and R. Haeb-Umbach, “Wide residual BLSTM network with discriminative speaker adaptation for robust speech recognition,” in
2016
Later among the works it cites.
H. Erdogan, T. Hayashi, J. R. Hershey, T. Hori, C. Hori, W.-N. Hsu, S. Kim, J. L. Roux, Z. Meng, and S. Watanabe, “Multi-channel speech recognition: LSTMs all the way through,” in
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” in
2012
Cited alongside, same era.
M. L. Seltzer, D. Yu, and Y. Wang, “An investigation of deep neural networks for noise robust speech recognition,” in
2013
Cited alongside, same era.
P. Swietojanski, A. Ghoshal, and S. Renals, “Hybrid acoustic models for distant and multichannel large vocabulary speech recognition,” in
2013
Cited alongside, same era.
N. Jaitly and G. E. Hinton, “Vocal tract length perturbation (VTLP) improves speech recognition,” in
2013
Cited alongside, same era.
J. Du, Q. Wang, T. Gao, Y. Xu, L. Dai, and C.-H. Lee, “Robust speech recognition with speech enhanced deep neural networks,” in
2014
Cited alongside, same era.
Y. Xu, J. Du, L.-R. Dai, and C.-H. Lee, “An experimental study on speech enhancement based on deep neural networks,”
2014
Cited alongside, same era.
F. Weninger, H. Erdogan, S. Watanabe, E. Vincent, J. L. Roux, J. R. Hershey, and B. Schuller, “Speech enhancement with LSTM recurrent neural networks and its application to noise-robust ASR,” in
2015
Cited alongside, same era.
2016
Later among the works it cites.
Y. Qian, T. Tan, and D. Yu, “An investigation into using parallel data for far-field speech recognition,” in
2016
Later among the works it cites.
Y. Qian, T. Tan, D. Yu, and Y. Zhang, “Integrated adaptation with multi-factor joint-learning for far-field speech recognition,” in
2016
Later among the works it cites.
Y. Zhang, G. Chen, D. Yu, K. Yao, S. Khudanpur, and J. Glass, “Highway long short-term memory RNNs for distant speech recognition,” in
2016
Later among the works it cites.
V. Peddinti, V. Manohar, Y. Wang, D. Povey, and S. Khudanpur, “Far-field ASR without parallel data,” in
2016
Later among the works it cites.
T. Ko, V. Peddinti, D. Povey, M. L. Seltzer, and S. Khudanpur, “A study on data augmentation of reverberant speech for robust speech recognition,” in
2017
Later among the works it cites.
W.-N. Hsu, Y. Zhang, and J. Glass, “Unsupervised learning of disentangled and interpretable representations from sequential data,” in
2017
Later among the works it cites.
V. Peddinti, Y. Wang, D. Povey, and S. Khudanpur, “Low latency acoustic modeling using temporal convolution and LSTMs,”
2018
Closest in time.
W.-N. Hsu and J. Glass, “Extracting domain invariant features by unsupervised learning for robust automatic speech recognition,” in
2018
Closest in time.