Fetching the paper…
Reading the bibliography…
Eliminating the negative effect of non-stationary environmental noise is a long-standing research topic for automatic speech recognition that stills remains an important challenge.
S. Boll, “Suppression of acoustic noise in speech using spectral subtraction,” IEEE Transactions on Acoustics, Speech, and Signal Processing , vol. 27, no. 2, pp. 113–120, Apr. 1979
1979
Earlier work this paper cites.
B. Li, T. N. Sainath, R. J. Weiss, K. W. Wilson, and M. Bacchiani, “Neural network adaptive beamforming for robust multichannel speech recognition,” in Proc. INTERSPEECH , San Francisco, CA, 2016, pp. 1976–1980
1980
Earlier work this paper cites.
Y. Ephraim and D. Malah, “Speech enhancement using a minimum-mean square error short-time spectral amplitude estimator,” IEEE Transactions on Acoustics, Speech, and Signal Processing , vol. 32, no. 6, pp. 1109–1121, Dec. 1984
1984
Earlier work this paper cites.
——, “Speech enhancement using a minimum mean-square error log-spectral amplitude estimator,” IEEE Transactions on Acoustics, Speech, and Signal Processing , vol. 23, no. 2, pp. 443–445, Apr. 1985
1985
Earlier work this paper cites.
H. Erdogan, J. R. Hershey, S. Watanabe, M. I. Mandel, and J. L. Roux, “Improved MVDR beamforming using single-channel mask prediction networks,” in Proc. INTERSPEECH , San Francisco, CA, 2016, pp. 1981–1985
1985
Earlier work this paper cites.
H. Cox, R. M. Zeskind, and M. M. Owen, “Robust adaptive beamforming,” IEEE Transactions on Acoustics, Speech, and Signal Processing , vol. 35, no. 10, pp. 1365–1376, Oct. 1987
1987
Earlier work this paper cites.
S. R. Quackenbush, T. P. Barnwell, and M. A. Clements, Objective Measures of Speech quality . Upper Saddle River, NJ: Prentice-Hall, 1988
1988
Earlier work this paper cites.
B. D. Van Veen and K. M. Buckley, “Beamforming: A versatile approach to spatial filtering,” IEEE ASSP Magazine , vol. 5, no. 2, pp. 4–24, Apr. 1988
1988
Earlier work this paper cites.
Y. LeCun, B. Boser, J. S. Denker, D. Henderson, R. E. Howard, W. Hubbard, and L. D. Jackel, “Backpropagation applied to handwritten zip code recognition,” Neural computation , vol. 1, no. 4, pp. 541–551, 1989
1989
Earlier work this paper cites.
J. H. Hansen and M. A. Clements, “Constrained iterative speech enhancement with application to speech recognition,” IEEE Transaction on Signal Processing , vol. 39, no. 4, pp. 795–805, Apr. 1991
1991
Earlier work this paper cites.
J.-L. Gauvain and C.-H. Lee, “Maximum a posteriori estimation for multivariate gaussian mixture observations of markov chains,” IEEE transactions on speech and audio processing , vol. 2, no. 2, pp. 291–298, Apr. 1994
1994
Earlier work this paper cites.
Y. Gong, “Speech recognition in noisy environments: A survey,” Speech communication , vol. 16, no. 3, pp. 261–291, Apr. 1995
1995
Earlier work this paper cites.
C. J. Leggetter and P. C. Woodland, “Maximum likelihood linear regression for speaker adaptation of continuous density hidden markov models,” Computer Speech & Language , vol. 9, no. 2, pp. 171–185, Apr. 1995
1995
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural computation , vol. 9, no. 8, pp. 1735–1780, Nov. 1997
1997
Earlier work this paper cites.
J. H. Hansen and B. L. Pellom, “An effective quality evaluation protocol for speech enhancement algorithms,” in Proc. International Conference on Spoken Language Processing (ICSLP) , Sydney, Australia, 1998, pp. 2819–2822
1998
Earlier work this paper cites.
C. Marro, Y. Mahieux, and K. U. Simmer, “Analysis of noise reduction and dereverberation techniques based on microphone arrays with postfiltering,” IEEE Transactions on Speech and Audio Processing , vol. 6, no. 3, pp. 240–259, May 1998
1998
Earlier work this paper cites.
D. D. Lee and H. S. Seung, “Learning the parts of objects by non-negative matrix factorization,” Nature , vol. 401, no. 6755, pp. 788–791, Oct. 1999
1999
Earlier work this paper cites.
A. Moreno, B. Lindberg, C. Draxler, G. Richard, K. Choukri, S. Euler, and J. Allen, “SPEECHDAT-CAR. A large speech database for automotive environments,” in Proc. the Second International Conference on Language Resources and Evaluation LREC , Athens, Greece, 2000, 6 pages
2000
Earlier work this paper cites.
D. Pearce and H. Hirsch, “The aurora experimental framework for the performance evaluation of speech recognition systems under noisy conditions,” in Proc. INTERSPEECH , Beijing, China, 2000, pp. 29–32
2000
Earlier work this paper cites.
S. Sharma, D. Ellis, S. S. Kajarekar, P. Jain, and H. Hermansky, “Feature extraction using non-linear transformation for robust speech recognition on the Aurora database,” in Proc. IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP) , Istanbul, Turkey, 2000, pp. 1117–1120
2000
Earlier work this paper cites.
I.-T. R. P.862, “Perceptual evaluation of speech quality (PESQ): An objective method for end-to-end speech quality assessment of narrow-band telephone networks and speech codecs,” 2001
2001
Earlier work this paper cites.
D. Pearce and J. Picone, “Aurora working group: DSR front end LVCSR evaluation AU/384/02,” Inst. for Signal & Inform. Process., Mississippi State Univ., Tech. Rep , 2002
2002
Earlier work this paper cites.
I. McCowan and H. Bourlard, “Microphone array post-filter based on noise field coherence,” IEEE Transactions on Speech and Audio Processing , vol. 11, no. 6, pp. 709–716, Nov. 2003
2003
Earlier work this paper cites.
X. Mestre and M. A. Lagunas, “On diagonal loading for minimum variance beamformers,” in Proc. 3rd IEEE International Symposium on Signal Processing and Information Technology , Darmstadt, Germany, 2003, pp. 459–462
2003
Earlier work this paper cites.
J. Carletta, S. Ashby, S. Bourban, M. Flynn, M. Guillemot, T. Hain, J. Kadlec, V. Karaiskos, W. Kraaij, M. Kronenthal et al. , “The ami meeting corpus: A pre-announcement,” in International Workshop on Machine Learning for Multimodal Interaction , 2005, pp. 28–39
2005
Earlier work this paper cites.
D. Wang, On Ideal Binary Mask As the Computational Goal of Auditory Scene Analysis . Boston, MA: Springer US, 2005, pp. 181–197
2005
Earlier work this paper cites.
G. E. Hinton and R. R. Salakhutdinov, “Reducing the dimensionality of data with neural networks,” Science , vol. 313, no. 5786, pp. 504–507, July 2006
2006
Earlier work this paper cites.
E. Vincent, R. Gribonval, and C. Févotte, “Performance measurement in blind audio source separation,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 14, no. 4, pp. 1462–1469, July 2006
2006
Earlier work this paper cites.
S. Srinivasan, N. Roman, and D. Wang, “Binary and ratio time-frequency masks for robust speech recognition,” Speech Communication , vol. 48, no. 11, pp. 1486–1501, Nov. 2006
2006
Earlier work this paper cites.
Y. Avargel and I. Cohen, “System identification in the short-time fourier transform domain with crossband filtering,” IEEE Transactions Audio, Speech and Language Processing , vol. 15, no. 4, pp. 1305–1319, Mar. 2007
2007
Earlier work this paper cites.
H.-G. Hirsch, “Aurora-5 experimental framework for the performance evaluation of speech recognition in case of a hands-free speech input in noisy environments,” Niederrhein Univ. of Applied Sciences , 2007
2007
Earlier work this paper cites.
E. Warsitz and R. Haeb-Umbach, “Blind acoustic beamforming based on generalized eigenvalue decomposition,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 15, no. 5, pp. 1529–1539, July 2007
2007
Earlier work this paper cites.
Y. Hu and P. C. Loizou, “Evaluation of objective quality measures for speech enhancement,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 16, no. 1, pp. 229–238, Jan. 2008
2008
Earlier work this paper cites.
M. Wöllmer, F. Eyben, B. W. Schuller, Y. Sun, T. Moosmayr, and N. Nguyen-Thien, “Robust in-car spelling recognition – a tandem BLSTM-HMM approach,” in Proc. INTERSPEECH , Brighton, United Kingdom, 2009, pp. 2507–2510
2009
Earlier work this paper cites.
M. Wöllmer, F. Eyben, A. Graves, B. Schuller, and G. Rigoll, “Improving keyword spotting with a tandem BLSTM-DBN architecture,” in Proc. Advances in Non-Linear Speech Processing: International Conference on Nonlinear Speech Processing (NOLISP) , Vic, Spain, 2010, pp. 68–75
2010
Earlier work this paper cites.
M. Wöllmer, B. Schuller, F. Eyben, and G. Rigoll, “Combining Long Short-Term Memory and Dynamic Bayesian Networks for Incremental Emotion-Sensitive Artificial Listening,” IEEE Journal of Selected Topics in Signal Processing, Special Issue on Speech Processing for Natural Interaction with Intelligent Environments , vol. 4, no. 5, pp. 867–881, Oct. 2010
2010
Earlier work this paper cites.
B. Schuller, F. Weninger, M. Wöllmer, Y. Sun, and G. Rigoll, “Non-negative matrix factorization as noise-robust feature extractor for speech recognition,” in Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , Dallas, TX, 2010, pp. 4562–4565
2010
Earlier work this paper cites.
J. Schalkwyk, D. Beeferman, F. Beaufays, B. Byrne, C. Chelba, M. Cohen, M. Kamvar, and B. Strope, ““your word is my command”: Google search by voice: a case study,” in Advances in Speech Recognition . Springer, 2010, pp. 61–90
2010
Earlier work this paper cites.
M. Wöllmer, B. Schuller, F. Eyben, and G. Rigoll, “Combining long short-term memory and dynamic bayesian networks for incremental emotion-sensitive artificial listening,” IEEE Journal of Selected Topics in Signal Processing , vol. 4, no. 5, pp. 867–881, Oct. 2010
2010
Earlier work this paper cites.
L. Deng, “Front-end, back-end, and hybrid techniques for noise-robust speech recognition,” in Robust Speech Recognition of Uncertain or Missing Data . Berlin/Heidelburg, Germany: Springer, 2011, pp. 67–99
2011
Earlier work this paper cites.
K. Paliwal, K. Wójcicki, and B. Shannon, “The importance of phase in speech enhancement,” speech communication , vol. 53, no. 4, pp. 465–494, Apr. 2011
2011
Earlier work this paper cites.
N. Dehak, P. Kenny, R. Dehak, P. Dumouchel, and P. Ouellet, “Front-end factor analysis for speaker verification,” IEEE Transactions Audio, Speech and Language Processing , vol. 19, no. 4, pp. 788–798, May 2011
2011
Earlier work this paper cites.
T. Yoshioka, A. Sehr, M. Delcroix, K. Kinoshita, R. Maas, T. Nakatani, and W. Kellermann, “Making machines understand us in reverberant rooms: robustness against reverberation for automatic speech recognition,” IEEE Signal Processing Magazine , vol. 29, no. 6, pp. 114–126, Nov. 2012
2012
Earlier work this paper cites.
Y.-H. Yang and H. H. Chen, “Machine recognition of music emotion: A review,” ACM Transactions on Intelligent Systems and Technology , vol. 3, no. 3, pp. 40:1–40:30, May 2012
2012
Earlier work this paper cites.
A. L. Maas, Q. V. Le, T. M. OŃeil, O. Vinyals, P. Nguyen, and A. Y. Ng, “Recurrent neural networks for noise reduction in robust ASR,” in Proc. INTERSPEECH , Portland, OR, 2012, pp. 22–25
2012
Earlier work this paper cites.
G. E. Dahl, D. Yu, L. Deng, and A. Acero, “Context-dependent pre-trained deep neural networks for large-vocabulary speech recognition,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 20, no. 1, pp. 30–42, Jan. 2012
2012
Earlier work this paper cites.
G. Hinton, L. Deng, D. Yu, G. E. Dahl, A. rahman Mohamed, N. Jaitly, A. Senior, V. Vanhoucke, P. Nguyen, T. N. Sainath, and B. Kingsbury, “Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups,” IEEE Signal Processing Magazine , vol. 29, no. 6, pp. 82–97, Nov. 2012
2012
Earlier work this paper cites.
A. Acero, Acoustical and environmental robustness in automatic speech recognition . Berlin, Germany: Springer Science & Business Media, 2012, vol. 201
2012
Earlier work this paper cites.
T. Virtanen, R. Singh, and B. Raj, Techniques for noise robustness in automatic speech recognition . Hoboken, NJ: John Wiley & Sons, 2012
2012
Earlier work this paper cites.
F. Weninger, J. Feliu, and B. Schuller, “Supervised and semi-supervised suppression of background music in monaural speech recordings,” in Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , Kyoto, Japan, 2012, pp. 61–64
2012
Cited alongside, same era.
D. Michelsanti and Z.-H. Tan, “Conditional generative adversarial networks for speech enhancement and noise-robust speaker verification,” in Proc. INTERSPEECH , Stockholm, Sweden, 2017, pp. 2008–2012
2012
Cited alongside, same era.
A. Khabbazibasmenj, S. A. Vorobyov, and A. Hassanien, “Robust adaptive beamforming based on steering vector estimation with as little as possible prior information,” IEEE Transactions on Signal Processing , vol. 60, no. 6, pp. 2974–2987, June 2012
2012
Cited alongside, same era.
P. C. Loizou, Speech enhancement: theory and practice . Abingdon, UK: Taylor Francis, 2013
2013
Cited alongside, same era.
Y. Wang and D. Wang, “A deep neural network for time-domain signal reconstruction,” in Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , Brisbane, Australia, 2015, pp. 4390–4394
2015
Later among the works it cites.
S. Mirsamadi and J. H. Hansen, “A study on deep neural network acoustic model adaptation for robust far-field speech recognition,” in Proc. INTERSPEECH , Dresden, Germany, 2015, pp. 2430–2434
2015
Later among the works it cites.
C. Yu, A. Ogawa, M. Delcroix, T. Yoshioka, T. Nakatani, and J. H. Hansen, “Robust i-vector extraction for neural network adaptation in noisy environment,” in Proc. INTERSPEECH , Dresden, Germany, 2015, pp. 2854–2857
2015
Later among the works it cites.
R. Giri, M. L. Seltzer, J. Droppo, and D. Yu, “Improving speech recognition in reverberation using a room-aware deep neural network and multi-task learning,” in Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , Brisbane, Australia, 2015, pp. 5014–5018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Barker, E. Vincent, N. Ma, H. Christensen, and P. Green, “The PASCAL CHiME speech separation and recognition challenge,” Computer Speech & Language , vol. 27, no. 3, pp. 621–633, May 2013
2013
Cited alongside, same era.
E. Vincent, J. Barker, S. Watanabe, J. Le Roux, F. Nesta, and M. Matassoni, “The second ‘CHiME’ speech separation and recognition challenge: Datasets, tasks and baselines,” in Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , Vancouver, Canada, 2013, pp. 126–130
2013
Cited alongside, same era.
X. Lu, Y. Tsao, S. Matsuda, and C. Hori, “Speech enhancement based on deep denoising autoencoder,” in Proc. INTERSPEECH , Lyon, France, 2013, pp. 436–440
2013
Cited alongside, same era.
B. Xia and C. Bao, “Speech enhancement with weighted denoising auto-encoder,” in Proc. INTERSPEECH , Lyon, France, 2013, pp. 3444–3448
2013
Cited alongside, same era.
T. Ishii, H. Komiyama, T. Shinozaki, Y. Horiuchi, and S. Kuroiwa, “Reverberant speech recognition based on denoising autoencoder,” in Proc. INTERSPEECH , Lyon, France, 2013, pp. 3512–3516
2013
Cited alongside, same era.
2013
Cited alongside, same era.
M. Wöllmer, Z. Zhang, F. Weninger, B. Schuller, and G. Rigoll, “Feature enhancement by bidirectional LSTM networks for conversational speech recognition in highly non-stationary noise,” in Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , Vancouver, Canada, 2013, pp. 6822–6826
2013
Cited alongside, same era.
F. Weninger, J. Geiger, M. Wöllmer, B. Schuller, and G. Rigoll, “The munich feature enhancement approach to the 2nd CHiME challenge using BLSTM recurrent neural networks,” in Proc. 2nd CHiME workshop on machine listening in multisource environments , Vancouver, Canada, 2013, pp. 86–90
2013
Cited alongside, same era.
2015
Later among the works it cites.
Z. Chen, S. Watanabe, H. Erdoğan, and J. R. Hershey, “Speech enhancement and recognition using multi-task learning of long short-term memory recurrent neural networks,” in Proc. INTERSPEECH , Dresden, Germany, 2015, pp. 1–5
2015
Later among the works it cites.
T. Gao, J. Du, L. Dai, and C. Lee, “Joint training of front-end and back-end deep neural networks for robust speech recognition,” in Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , South Brisbane, Australia, 2015, pp. 4375–4379
2015
Later among the works it cites.
J. Heymann, L. Drude, A. Chinaev, and R. Haeb-Umbach, “BLSTM supported GEV beamformer front-end for the 3rd CHiME challenge,” in Proc. IEEE Workshop on Automatic Speech Recognition and Understanding (ASRU) , Scottsdale, AZ, 2015, pp. 444–451
2015
Later among the works it cites.
Y. Hoshen, R. J. Weiss, and K. W. Wilson, “Speech acoustic modeling from raw multichannel waveforms,” in Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , Brisbane, Australia, 2015, pp. 4624–4628
2015
Later among the works it cites.
A. Narayanan and D. Wang, “Improving robustness of deep neural network acoustic models via speech separation and joint adaptive training,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 23, no. 1, pp. 92–101, Jan. 2015
2015
Later among the works it cites.
G. Saon, T. Sercu, S. Rennie, and H.-K. J. Kuo, “The IBM 2016 english conversational telephone speech recognition system,” in Proc. INTERSPEECH , San Francisco, CA, 2016, pp. 7–11
2016
Later among the works it cites.
D. Amodei, R. Anubhai, E. Battenberg, C. Case, J. Casper, B. Catanzaro, J. Chen, M. Chrzanowski, A. Coates, G. Diamos et al. , “Deep speech 2: End-to-end speech recognition in english and mandarin,” in Proc. International Conference on Machine Learning (ICML) , New York City, NY, 2016, 173-182
2016
Later among the works it cites.
W. Xiong, J. Droppo, X. Huang, F. Seide, M. Seltzer, A. Stolcke, D. Yu, and G. Zweig, “Achieving human parity in conversational speech recognition,” Microsoft Research, Tech. Rep. MSR-TR-2016-71, Oct. 2016
2016
Later among the works it cites.
K. Kinoshita, M. Delcroix, S. Gannot, E. A. Habets, R. Haeb-Umbach, W. Kellermann, V. Leutnant, R. Maas, T. Nakatani, B. Raj et al. , “A summary of the REVERB challenge: state-of-the-art and remaining challenges in reverberant speech processing research,” EURASIP Journal on Advances in Signal Processing , vol. 2016, no. 1, pp. 1–19, Dec. 2016
2016
Later among the works it cites.
E. Vincent, S. Watanabe, A. A. Nugraha, J. Barker, and R. Marxer, “An analysis of environment, microphone and data simulation mismatches in robust speech recognition,” Computer Speech & Language , Dec. 2016, in press
2016
Later among the works it cites.
Y. Qian, M. Bi, T. Tan, and K. Yu, “Very deep convolutional neural networks for noise robust speech recognition,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 24, no. 12, pp. 2263–2276, Dec. 2016
2016
Later among the works it cites.
2016
Later among the works it cites.
M. Schedl, Y.-H. Yang, and P. Herrera-Boyer, “Introduction to intelligent music systems and applications,” ACM Transactions on Intelligent Systems and Technology , vol. 8, no. 2, pp. 17:1–17:8, Oct. 2016
2016
Later among the works it cites.
I. Goodfellow, Y. Bengio, and A. Courville, Deep Learning . Cambridge, MA: MIT Press, 2016
2016
Later among the works it cites.
G. Trigeorgis, F. Ringeval, R. Bruckner, E. Marchi, M. Nicolaou, B. Schuller, and S. Zafeiriou, “Adieu features? End-to-end speech emotion recognition using a Deep Convolutional Recurrent Network,” in Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , Shanghai, China, 2016, pp. 5200–5204
2016
Later among the works it cites.
Z. Zhang, F. Ringeval, J. Han, J. Deng, E. Marchi, and B. Schuller, “Facing realism in spontaneous emotion recognition from speech: Feature enhancement by autoencoder with LSTM neural networks,” in Proc. INTERSPEECH , San Francisco, CA, 2016, pp. 3593–3597
2016
Later among the works it cites.
X. Mao, C. Shen, and Y.-B. Yang, “Image restoration using very deep convolutional encoder-decoder networks with symmetric skip connections,” in Proc. Advances In Neural Information Processing Systems (NIPS) , Barcelona Spain, 2016, pp. 2802–2810
2016
Later among the works it cites.
2016
Later among the works it cites.
2016
Later among the works it cites.
E. Grais, G. Roma, A. J. Simpson, and M. D. Plumbley, “Combining mask estimates for single channel audio source separation using deep neural networks,” in Proc. INTERSPEECH , San Francisco, CA, 2016, pp. 3339–3343
2016
Later among the works it cites.
K. H. Lee, S. J. Kang, W. H. Kang, and N. S. Kim, “Two-stage noise aware training using asymmetric deep denoising autoencoder,” in Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , Shanghai, China, 2016, pp. 5765–5769
2016
Later among the works it cites.
Z. Wang and D. Wang, “A joint training framework for robust automatic speech recognition,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 24, no. 4, pp. 796–806, Apr. 2016
2016
Later among the works it cites.
M. Mimura, S. Sakai, and T. Kawahara, “Joint optimization of denoising autoencoder and DNN acoustic model based on multi-target learning for noisy speech recognition,” in Proc. INTERSPEECH , San Francisco, CA 2016, pp. 3803–3807
2016
Later among the works it cites.
J. Heymann, L. Drude, and R. Haeb-Umbach, “Neural network based spectral mask estimation for acoustic beamforming,” in Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , Shanghai, China, 2016, pp. 196–200
2016
Later among the works it cites.
T. Menne, J. Heymann, A. Alexandridis, K. Irie, A. Zeyer, M. Kitza, P. Golik, K. Ilia, L. Durde, R. Schlater, H. Ney, R. Haeb-Umbach, and A. Mouchtaris, “The RWTH /UPB/FORTH system combination for the 4th CHiME challenge evaluation,” in Proc. 4th International Workshop on Speech Processing in Everyday Environments (CHiME) , San Francisco, CA, USA, 2016, pp. 49–51
2016
Later among the works it cites.
X. Xiao, S. Watanabe, H. Erdogan, L. Lu, J. Hershey, M. L. Seltzer, G. Chen, Y. Zhang, M. Mandel, and D. Yu, “Deep beamforming networks for multi-channel speech recognition,” in Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , Shanghai, China, 2016, pp. 5745–5749
2016
Later among the works it cites.
J. Heymann, L. Drude, and R. Haeb-Umbach, “Wide residual BLSTM network with discriminative speaker adaptation for robust speech recognition,” in Proc. 4th International Workshop on Speech Processing in Everyday Environments (CHiME) , San Francisco, CA, 2016, pp. 12–17
2016
Later among the works it cites.
S. Zagoruyko and N. Komodakis, “Wide residual networks,” arXiv preprint arXiv:1605.07146 , May 2016
2016
Later among the works it cites.
S. Kundu, G. Mantena, Y. Qian, T. Tan, M. Delcroix, and K. C. Sim, “Joint acoustic factor learning for robust deep neural network based automatic speech recognition,” in Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , Shanghai, China, 2016, pp. 5025–5029
2016
Later among the works it cites.
X. Xiao, C. Xu, Z. Zhang, S. Zhao, S. Sun, and S. Watanabe, “A study of learning based beamforming methods for speech recognition,” in Proc. CHiME Workshop , San Francisco, CA, 2016, pp. 26–31
2016
Later among the works it cites.
H. Erdogan, T. Hayashi, J. R. Hershey, T. Hori, C. Hori, W.-N. Hsu, S. Kim, J. Le Roux, Z. Meng, and S. Watanabe, “Multi-channel speech recognition: LSTMs all the way through,” in Proc. CHiME-4 Workshop , San Francisco, CA, 2016
2016
Later among the works it cites.
Y. Qian and T. Tan, “The SJTU CHiME-4 system: Acoustic noise robustness for real single or multiple microphone scenarios,” in Proc. CHiME-4 Workshop , San Francisco, CA, 2016
2016
Later among the works it cites.
Z. Zhang, N. Cummins, and B. Schuller, “Advanced data exploitation for speech analysis – an overview,” IEEE Signal Processing Magazine , vol. 34, July 2017, 24 pages
2017
Closest in time.
Y. Liu, Y. Liu, S. Zhong, and S. Wu, “Implicit visual learning: Image recognition via dissipative learning model,” ACM Transactions on Intelligent Systems and Technology , vol. 8, no. 2, pp. 31:1–31:24, Jan. 2017
2017
Closest in time.
K. Qian, Y. Zhang, S. Chang, X. Yang, D. Florêncio, and M. Hasegawa-Johnson, “Speech enhancement using Bayesian WaveNet,” Proc. Interspeech , pp. 2013–2017, 2017
2017
Closest in time.
2017
Closest in time.
2017
Closest in time.
A. Creswell, T. White, V. Dumoulin, K. Arulkumaran, B. Sengupta, and A. A. Bharath, “Generative adversarial networks: An overview,” IEEE Signal Processing Magzine , 2017, 14 pages, to appear
2017
Closest in time.
D. S. Williamson and D. Wang, “Time-frequency masking in the complex domain for speech dereverberation and denoising,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 25, no. 7, pp. 1492–1501, July 2017
2017
Closest in time.
D. S. Williamson and D. Wang, “Speech dereverberation and denoising using complex ratio masks,” in Proc. IEEE International Conference on Audio, Speech, and Signal Processing (ICASSP) , New Orleans, LA, 2017, pp. 5590–5594
2017
Closest in time.
K. H. Lee, W. H. Kang, T. G. Kang, and N. S. Kim, “Integrated DNN-based model adaptation technique for noise-robust speech recognition,” in Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , New Orleans, LA, 2017, pp. 5245–5249
2017
Closest in time.
M. Ravanelli, P. Brakel, M. Omologo, and Y. Bengio, “A network of deep neural networks for distant speech recognition,” in Proc. IEEE International Conference on Audio, Speech, and Signal Processing (ICASSP) , New Orleans, LA, 2017, pp. 4880–4884
2017
Closest in time.
Z. Meng, S. Watanabe, J. R. Hershey, and H. Erdogan, “Deep long short-term memory adaptive beamforming networks for multichannel robust speech recognition,” in Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , New Orleans, LA, 2017, pp. 271–275
2017
Closest in time.
T. N. Sainath, R. J. Weiss, K. W. Wilson, B. Li, A. Narayanan, E. Variani, M. Bacchiani, I. Shafran, A. W. Senior, K. K. Chin, A. Misra, and C. Kim, “Multichannel signal processing with deep neural networks for automatic speech recognition,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 25, no. 5, pp. 965–979, May 2017
2017
Closest in time.
T. Ochiai, S. Watanabe, T. Hori, and J. R. Hershey, “Multichannel end-to-end speech recognition,” in Proc. the 34th International Conference on Machine Learning (ICML) , Sydney, Australia, 2017, pp. 2632–2641
2017
Closest in time.