Fetching the paper…
Reading the bibliography…
Although great progresses have been made in automatic speech recognition (ASR), significant performance degradation is still observed when recognizing multi-talker mixed speech.
Z. Ghahramani and M. I. Jordan, “Factorial hidden Markov models,” Machine learning (MLJ) , vol. 29, no. 2-3, pp. 245–273, 1997
1997
Earlier work this paper cites.
J. J. Godfrey and E. Holliman, “Switchboard-1 release 2,” Linguistic Data Consortium, Philadelphia , 1997
1997
Earlier work this paper cites.
D. Yu, L. Deng, and G. E. Dahl, “Roles of pre-training and fine-tuning in context-dependent DBN-HMMs for real-world speech recognition,” in NIPS Workshop on Deep Learning and Unsupervised Feature Learning , 2010
2010
Earlier work this paper cites.
M. Cooke, J. R. Hershey, and S. J. Rennie, “Monaural speech separation and recognition challenge,” Computer Speech and Language (CSL) , vol. 24, pp. 1–15, 2010
2010
Earlier work this paper cites.
J. R. Hershey, S. J. Rennie, P. A. Olsen, and T. T. Kristjansson, “Super-human multi-talker speech recognition: A graphical modeling approach,” Computer Speech and Language (CSL) , vol. 24, pp. 45 – 66, 2010
2010
Earlier work this paper cites.
S. J. Rennie, J. R. Hershey, and P. A. Olsen, “Single-channel multitalker speech recognition,” IEEE Signal Processing Magazine (SPM) , vol. 27, pp. 66–80, 2010
2010
Earlier work this paper cites.
F. Seide, G. Li, and D. Yu, “Conversational speech transcription using context-dependent deep neural networks.” in Annual Conference of International Speech Communication Association (INTERSPEECH) , 2011, pp. 437–440
2011
Earlier work this paper cites.
2011
Earlier work this paper cites.
G. E. Dahl, D. Yu, L. Deng, and A. Acero, “Context-dependent pre-trained deep neural networks for large-vocabulary speech recognition,” IEEE/ACM Transactions on Audio, Speech, and Language Processing (TASLP) , vol. 20, pp. 30–42, 2012
2012
Earlier work this paper cites.
G. Hinton, L. Deng, D. Yu, G. E. Dahl, A.-r. Mohamed, N. Jaitly, A. Senior, V. Vanhoucke, P. Nguyen, T. N. Sainath et al. , “Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups,” IEEE Signal Processing Magazine (SPM) , vol. 29, pp. 82–97, 2012
2012
Earlier work this paper cites.
O. Abdel-Hamid, A.-r. Mohamed, H. Jiang, and G. Penn, “Applying convolutional neural networks concepts to hybrid NN-HMM model for speech recognition,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2012, pp. 4277–4280
2012
Earlier work this paper cites.
T. Hain, L. Burget, J. Dines, P. N. Garner, F. Grézl, A. E. Hannani, M. Huijbregts, M. Karafiat, M. Lincoln, and V. Wan, “Transcribing meetings with the AMIDA systems,” IEEE/ACM Transactions on Audio, Speech, and Language Processing (TASLP) , vol. 20, no. 2, pp. 486–498, 2012
2012
Earlier work this paper cites.
P. Swietojanski, A. Ghoshal, and S. Renals, “Hybrid acoustic models for distant and multichannel large vocabulary speech recognition,” in IEEE Workshop on Automatic Speech Recognition and Understanding (ASRU) , 2013, pp. 285–290
2013
Earlier work this paper cites.
D. Yu and L. Deng, Automatic Speech Recognition: A Deep Learning Approach , ser. Signals and Communication Technology. Springer London, 2014. [Online]. Available: https://books.google.com/books?id=rUBTBQAAQBAJ
2014
Earlier work this paper cites.
O. Abdel-Hamid, A.-r. Mohamed, H. Jiang, L. Deng, G. Penn, and D. Yu, “Convolutional neural networks for speech recognition,” IEEE/ACM Transactions on Audio, Speech, and Language Processing (TASLP) , vol. 22, pp. 1533–1545, 2014
2014
Earlier work this paper cites.
P.-S. Huang, M. Kim, M. Hasegawa-Johnson, and P. Smaragdis, “Deep learning for monaural speech separation,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2014, pp. 1562–1566
2014
Cited alongside, same era.
Y. Wang, A. Narayanan, and D. Wang, “On training targets for supervised speech separation,” IEEE/ACM Transactions on Audio, Speech and Language Processing (TASLP) , vol. 22, pp. 1849–1858, 2014
2014
Cited alongside, same era.
Y. Xu, J. Du, L.-R. Dai, and C.-H. Lee, “An experimental study on speech enhancement based on deep neural networks,” IEEE Signal Processing Letters (SPL) , vol. 21, pp. 65–68, 2014
2014
Cited alongside, same era.
D. Yu, A. Eversole, M. Seltzer, K. Yao, Z. Huang, B. Guenter, O. Kuchaiev, Y. Zhang, F. Seide, H. Wang et al. , “An introduction to computational networks and the computational network toolkit,” Microsoft Technical Report MSR-TR-2014–112 , 2014
2014
Cited alongside, same era.
D. Amodei, R. Anubhai, E. Battenberg, C. Case, J. Casper, B. Catanzaro, J. Chen, M. Chrzanowski, A. Coates, G. Diamos et al. , “Deep speech 2: End-to-end speech recognition in English and Mandarin,” in International Conference on Machine Learning (ICML) , 2016
2016
Later among the works it cites.
S. Zhang, H. Jiang, S. Xiong, S. Wei, and L. Dai, “Compact feedforward sequential memory networks for large vocabulary continuous speech recognition,” in Annual Conference of International Speech Communication Association (INTERSPEECH) , 2016, pp. 3389–3393
2016
Later among the works it cites.
D. Yu, W. Xiong, J. Droppo, A. Stolcke, G. Ye, J. Li, and G. Zweig, “Deep convolutional neural networks with layer-wise context expansion and attention.” in Annual Conference of International Speech Communication Association (INTERSPEECH) , 2016, pp. 17–21
2016
Later among the works it cites.
J. R. Hershey, Z. Chen, J. L. Roux, and S. Watanabe, “Deep clustering: Discriminative embeddings for segmentation and separation,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2016, pp. 31–35
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
T. N. Sainath, O. Vinyals, A. Senior, and H. Sak, “Convolutional, long short-term memory, fully connected deep neural networks,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2015, pp. 4580–4584
2015
Cited alongside, same era.
M. Bi, Y. Qian, and K. Yu, “Very deep convolutional neural networks for LVCSR,” in Annual Conference of International Speech Communication Association (INTERSPEECH) , 2015, pp. 3259–3263
2015
Cited alongside, same era.
V. Mitra and H. Franco, “Time-frequency convolutional networks for robust speech recognition,” in IEEE Workshop on Automatic Speech Recognition and Understanding (ASRU) , 2015, pp. 317–323
2015
Cited alongside, same era.
V. Peddinti, D. Povey, and S. Khudanpur, “A time delay neural network architecture for efficient modeling of long temporal contexts,” in Annual Conference of International Speech Communication Association (INTERSPEECH) , 2015, pp. 3214–3218
2015
Cited alongside, same era.
C. Weng, D. Yu, M. L. Seltzer, and J. Droppo, “Deep neural networks for single-channel multi-talker speech recognition,” IEEE/ACM Transactions on Audio, Speech, and Language Processing (TASLP) , vol. 23, no. 10, pp. 1670–1679, 2015
2015
Cited alongside, same era.
F. Weninger, H. Erdogan, S. Watanabe, E. Vincent, J. Roux, J. R. Hershey, and B. Schuller, “Speech enhancement with LSTM recurrent neural networks and its application to noise-robust ASR,” in International Conference on Latent Variable Analysis and Signal Separation (LVA/ICA) . Springer-Verlag New York, Inc., 2015, pp. 91–99
2015
Cited alongside, same era.
P. S. Huang, M. Kim, M. Hasegawa-Johnson, and P. Smaragdis, “Joint optimization of masks and deep recurrent neural networks for monaural source separation,” IEEE/ACM Transactions on Audio, Speech, and Language Processing (TASLP) , vol. 23, pp. 2136–2147, Dec 2015
2015
Cited alongside, same era.
Y. Qian, M. Bi, T. Tan, and K. Yu, “Very deep convolutional neural networks for noise robust speech recognition,” IEEE/ACM Transactions on Audio, Speech, and Language Processing (TASLP) , vol. 24, no. 12, pp. 2263–2276, 2016
2016
Cited alongside, same era.
2016
Later among the works it cites.
Y. Isik, J. L. Roux, Z. Chen, S. Watanabe, and J. R. Hershey, “Single-channel multi-speaker separation using deep clustering,” in Annual Conference of International Speech Communication Association (INTERSPEECH) , 2016, pp. 545–549
2016
Later among the works it cites.
J. Du, Y. Tu, L. R. Dai, and C. H. Lee, “A regression approach to single-channel speech separation via high-resolution deep neural networks,” IEEE/ACM Transactions on Audio, Speech, and Language Processing (TASLP) , vol. 24, pp. 1424–1437, Aug 2016
2016
Later among the works it cites.
D. Povey, V. Peddinti, D. Galvez, P. Ghahremani, V. Manohar, X. Na, Y. Wang, and S. Khudanpur, “Purely sequence-trained neural networks for ASR based on lattice-free MMI,” in Annual Conference of International Speech Communication Association (INTERSPEECH) , 2016, pp. 2751–2755
2016
Later among the works it cites.
Y. Zhang, G. Chen, D. Yu, K. Yao, S. Khudanpur, and J. Glass, “Highway long short-term memory RNNs for distant speech recognition,” IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pp. 5755–5759, 2016
2016
Later among the works it cites.
2016
Later among the works it cites.
G. Saon, T. Sercu, S. Rennie, and H.-K. J. Kuo, “The IBM 2016 english conversational telephone speech recognition system,” in Annual Conference of International Speech Communication Association (INTERSPEECH) , 2016, pp. 7–11
2016
Later among the works it cites.
W. Xiong, J. Droppo, X. Huang, F. Seide, M. Seltzer, A. Stolcke, D. Yu, and G. Zweig, “The Microsoft 2016 conversational speech recognition system,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2017, pp. 5255–5259
2017
Closest in time.
D. Yu, M. Kolbæk, Z.-H. Tan, and J. Jensen, “Permutation invariant training of deep models for speaker-independent multi-talker speech separation,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2017, pp. 241–245
2017
Closest in time.
M. Kolbæk, D. Yu, Z.-H. Tan, and J. Jensen, “Multi-talker speech separation with utterance-level permutation invariant training of deep recurrent neural networks,” IEEE/ACM Transactions on Audio, Speech, and Language Processing (TASLP) , accepted, 2017
2017
Closest in time.
Z. Chen, Y. Luo, and N. Mesgarani, “Deep attractor network for single-microphone speaker separation,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2017, pp. 246–250
2017
Closest in time.