Fetching the paper…
Reading the bibliography…
Integration of multiple microphone data is one of the key ways to achieve robust speech recognition in noisy environments or when the speaker is located at some distance from the input device.
Speech recognition in noisy environments with the aid of microphone arrays
Van Compernolle, Dirk, Ma, Weiye, Xie, Fei, and Van Diest, Marc · 1990
Earlier work this paper cites.
Connectionist speech recognition: a hybrid approach, 1994
Morgan, N and Bourlard, H · 1994
Earlier work this paper cites.
Long short-term memory
Hochreiter, Sepp and Schmidhuber, Jürgen · 1997
Earlier work this paper cites.
On diagonal loading for minimum variance beamformers
Mestre, Xavier, Lagunas, Miguel, et al · 2003
Earlier work this paper cites.
Likelihood-maximizing beamforming for robust hands-free speech recognition
Seltzer, Michael L, Raj, Bhiksha, and Stern, Richard M · 2004
Earlier work this paper cites.
Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks
Graves, Alex, Fernández, Santiago, Gomez, Faustino, and Schmidhuber, Jürgen · 2006
Earlier work this paper cites.
Csr-i (wsj0) complete
Garofalo, John, Graff, David, Paul, Doug, and Pallett, David · 2007
Earlier work this paper cites.
Bridging the gap: Towards a unified framework for hands-free speech recognition using microphone arrays
Seltzer, Michael L · 2008
Earlier work this paper cites.
Binaural and multiple-microphone signal processing motivated by auditory perception
Stern, Richard M, Gouvêa, Evandro, Kim, Chanwoo, Kumar, Kshitiz, and Park, Hyung-Min · 2008
Earlier work this paper cites.
Signal separation for robust speech recognition based on phase difference information obtained in the frequency domain
Kim, Chanwoo, Kumar, Kshitiz, Raj, Bhiksha, and Stern, Richard M · 2009
Earlier work this paper cites.
Spatial separation of speech signals using amplitude estimation based on interaural comparisons of zero-crossings
Park, Hyung-Min and Stern, Richard M · 2009
Earlier work this paper cites.
Adaptive segmentation and separation of determined convolutive mixtures under dynamic conditions
Loesch, Benedikt and Yang, Bin · 2010
Earlier work this paper cites.
The kaldi speech recognition toolkit
Povey, Daniel, Ghoshal, Arnab, Boulianne, Gilles, Burget, Lukas, Glembek, Ondrej, Goel, Nagendra, Hannemann, Mirko, Motlicek, Petr, Qian, Yanmin, Schwarz, Petr, Silovsky, Jan, Stemmer, Georg, and Vesely, Karel · 2011
Earlier work this paper cites.
Conversational speech transcription using context-dependent deep neural networks
Seide, Frank, Li, Gang, and Yu, Dong · 2011
Cited alongside, same era.
Multi-source tdoa estimation in reverberant audio using angular spectra and clustering
Blandin, Charles, Ozerov, Alexey, and Vincent, Emmanuel · 2012
Cited alongside, same era.
Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups
Hinton, Geoffrey, Deng, Li, Yu, Dong, Dahl, George E, Mohamed, Abdel-rahman, Jaitly, Navdeep, Senior, Andrew, Vanhoucke, Vincent, Nguyen, Patrick, Sainath, Tara N, et al · 2012
Cited alongside, same era.
Microphone array processing for distant speech recognition: From close-talking microphones to far-field sensors
Kumatani, Kenichi, McDonough, John, and Raj, Bhiksha · 2012
Cited alongside, same era.
Acoustic modeling using deep belief networks
Mohamed, Abdel-rahman, Dahl, George E, and Hinton, Geoffrey · 2012
Cited alongside, same era.
Neural networks for distant speech recognition
Renals, Steve and Swietojanski, Pawel · 2014
Later among the works it cites.
Sak, Haşim, Senior, Andrew, and Beaufays, Françoise · 2014
Later among the works it cites.
Convolutional neural networks for distant speech recognition
Swietojanski, Pawel, Ghoshal, Arnab, and Renals, Steve · 2014
Later among the works it cites.
Vinyals, Oriol, Kaiser, Lukasz, Koo, Terry, Petrov, Slav, Sutskever, Ilya, and Hinton, Geoffrey · 2014
Later among the works it cites.
End-to-end attention-based large vocabulary speech recognition
Bahdanau, Dzmitry, Chorowski, Jan, Serdyuk, Dmitriy, Brakel, Philemon, and Bengio, Yoshua · 2015
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Pascanu, Razvan, Mikolov, Tomas, and Bengio, Yoshua · 2012
Cited alongside, same era.
Hybrid speech recognition with deep bidirectional lstm
Graves, Alan, Jaitly, Navdeep, and Mohamed, Abdel-rahman · 2013
Cited alongside, same era.
Neural machine translation by jointly learning to align and translate
Bahdanau, Dzmitry, Cho, Kyunghyun, and Bengio, Yoshua · 2014
Cited alongside, same era.
End-to-end continuous speech recognition using attention-based recurrent nn: First results
Chorowski, Jan, Bahdanau, Dzmitry, Cho, Kyunghyun, and Bengio, Yoshua · 2014
Cited alongside, same era.
Towards end-to-end speech recognition with recurrent neural networks
Graves, Alex and Jaitly, Navdeep · 2014
Cited alongside, same era.
Deepspeech: Scaling up end-to-end speech recognition
Hannun, Awni, Case, Carl, Casper, Jared, Catanzaro, Bryan, Diamos, Greg, Elsen, Erich, Prenger, Ryan, Satheesh, Sanjeev, Sengupta, Shubho, Coates, Adam, et al · 2014
Cited alongside, same era.
Using neural network front-ends on far field multiple microphones based speech recognition
Liu, Yulan, Zhang, Pengyuan, and Hain, Thomas · 2014
Cited alongside, same era.
Closest in time.
Chan, William, Jaitly, Navdeep, Le, Quoc V, and Vinyals, Oriol · 2015
Closest in time.
Learning feature mapping using deep neural network bottleneck features for distant large vocabulary speech recognition
Himawan, Ivan, Motlicek, Petr, Imseng, David, Potard, Blaise, Kim, Namhoon, and Lee, Jaewon · 2015
Closest in time.
The third ’chime’ speech separation and recognition challenge: Dataset, task and baselines
Jon Barker, Ricard Marxer, Emmanuel Vincent Shinji Watanabe · 2015
Closest in time.
Distant speech separation using predicted time–frequency masks from spatial features
Pertilä, Pasi and Nikunen, Joonas · 2015
Closest in time.
Vinyals, Oriol and Le, Quoc · 2015
Closest in time.
Show, attend and tell: Neural image caption generation with visual attention
Xu, Kelvin, Ba, Jimmy, Kiros, Ryan, Courville, Aaron, Salakhutdinov, Ruslan, Zemel, Richard, and Bengio, Yoshua · 2015
Closest in time.
Far-field speech recognition using cnn-dnn-hmm with convolution in time
Yoshioka, Takuya, Karita, Shigeki, and Nakatani, Tomohiro · 2015
Closest in time.