Fetching the paper…
Reading the bibliography…
Far-field speech recognition in noisy and reverberant conditions remains a challenging problem despite recent deep learning breakthroughs.
“Beamforming: A versatile approach to spatial filtering,”
Barry D Van Veen and Kevin M Buckley, · 1988
Earlier work this paper cites.
“Long short-term memory,”
S. Hochreiter and J. Schmidhuber, · 1997
Earlier work this paper cites.
“Blind acoustic beamforming based on generalized eigenvalue decomposition,”
E. Warsitz and R. Haeb-Umbach, · 2007
Earlier work this paper cites.
“Acoustic beamforming for speaker diarization of meetings,”
X. Anguera, C. Wooters, and J. Hernando, · 2007
Earlier work this paper cites.
Microphone array signal processing
Jacob Benesty, Jingdong Chen, and Yiteng Huang, · 2008
Earlier work this paper cites.
“The kaldi speech recognition toolkit,”
Daniel Povey, Arnab Ghoshal, Gilles Boulianne, Lukas Burget, Ondrej Glembek, Nagendra Goel, Mirko Hannemann, Petr Motlicek, Yanmin Qian, Petr Schwarz, Jan Silovsky, Georg Stemmer, and Karel Vesely, · 2011
Earlier work this paper cites.
“Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups,”
G. Hinton, L. Deng, D. Yu, G. E. Dahl, A. r. Mohamed, N. Jaitly, A. Senior, V. Vanhoucke, P. Nguyen, T. N. Sainath, and B. Kingsbury, · 2012
Earlier work this paper cites.
“Context-dependent pre-trained deep neural networks for large-vocabulary speech recognition,”
G. E. Dahl, D. Yu, L. Deng, and A. Acero, · 2012
Earlier work this paper cites.
“Lstm neural networks for language modeling.,”
Martin Sundermeyer, Ralf Schlüter, and Hermann Ney, · 2012
Earlier work this paper cites.
“Revisiting recurrent neural networks for robust asr,”
O. Vinyals, S. V. Ravuri, and D. Povey, · 2012
Earlier work this paper cites.
“Speech recognition with deep recurrent neural networks,”
A. Graves, A. r. Mohamed, and G. Hinton, · 2013
Cited alongside, same era.
“Hybrid speech recognition with deep bidirectional lstm,”
A. Graves, N. Jaitly, and A. r. Mohamed, · 2013
Cited alongside, same era.
“Speech recognition with deep recurrent neural networks,”
A. Graves, A. r. Mohamed, and G. Hinton, · 2013
Cited alongside, same era.
“Memory-enhanced neural networks and nmf for robust asr,”
J. T. Geiger, F. Weninger, J. F. Gemmeke, M. Wöllmer, B. Schuller, and G. Rigoll, · 2014
Cited alongside, same era.
“Feature enhancement by deep {LSTM} networks for {ASR} in reverberant multisource environments,”
Felix Weninger, Jürgen Geiger, Martin Wöllmer, Björn Schuller, and Gerhard Rigoll, · 2014
Cited alongside, same era.
“Recurrent deep neural networks for robust speech recognition,”
C. Weng, D. Yu, S. Watanabe, and B. H. F. Juang, · 2014
“Speaker location and microphone spacing invariant acoustic modeling from raw multichannel waveforms,”
T. N. Sainath, R. J. Weiss, K. W. Wilson, A. Narayanan, M. Bacchiani, and Andrew, · 2015
Later among the works it cites.
“Phase-sensitive and recognition-boosted speech separation using deep recurrent neural networks,”
H. Erdogan, J. R. Hershey, S. Watanabe, and J. Le Roux, · 2015
Later among the works it cites.
“Attention-based models for speech recognition,”
Jan K Chorowski, Dzmitry Bahdanau, Dmitriy Serdyuk, Kyunghyun Cho, and Yoshua Bengio, · 2015
Later among the works it cites.
“The third chime speech separation and recognition challenge: Dataset, task and baselines,”
J. Barker, R. Marxer, E. Vincent, and S. Watanabe, · 2015
Later among the works it cites.
“Chainer: a next-generation open source framework for deep learning,”
Seiya Tokui, Kenta Oono, Shohei Hido, and Justin Clayton, · 2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Long short-term memory recurrent neural network architectures for large scale acoustic modeling.,”
Hasim Sak, Andrew W Senior, and Françoise Beaufays, · 2014
Cited alongside, same era.
“Strategies for distant speech recognitionin reverberant environments,”
Marc Delcroix, Takuya Yoshioka, Atsunori Ogawa, Yotaro Kubo, Masakiyo Fujimoto, Nobutaka Ito, Keisuke Kinoshita, Miquel Espi, Shoko Araki, Takaaki Hori, and Tomohiro Nakatani, · 2015
Cited alongside, same era.
“The merl/sri system for the 3rd chime challenge using beamforming, robust feature extraction, and advanced speech recognition,”
T. Hori, Z. Chen, H. Erdogan, J. R. Hershey, J. Le Roux, V. Mitra, and S. Watanabe, · 2015
Cited alongside, same era.
“Improved mvdr beamforming using single-channel mask prediction networks,”
H Erdogan, JR Hershey, S Watanabe, M Mandel, and J Le Roux, · 2016
Later among the works it cites.
“Deep beamforming networks for multi-channel speech recognition,”
X. Xiao, S. Watanabe, H. Erdogan, L. Lu, J. Hershey, M. L. Seltzer, G. Chen, Y. Zhang, M. Mandel, and D. Yu, · 2016
Later among the works it cites.
“Factored spatial and spectral multichannel raw waveform cldnns,”
T. N. Sainath, R. J. Weiss, K. W. Wilson, A. Narayanan, and M. Bacchiani, · 2016
Later among the works it cites.
“Neural network adaptive beamforming for robust multichannel speech recognition,”
Bo Li, Tara N Sainath, Ron J Weiss, Kevin W Wilson, and Michiel Bacchiani, · 2016
Later among the works it cites.