Fetching the paper…
Reading the bibliography…
Unsupervised single-channel overlapped speech recognition is one of the hardest problems in automatic speech recognition (ASR).
M. Wertheimer, “Laws of organization in perceptual forms.” 1938
1938
Earlier work this paper cites.
E. C. Cherry, “Some experiments on the recognition of speech, with one and with two ears,” The Journal of the acoustical society of America , vol. 25, no. 5, pp. 975–979, 1953
1953
Earlier work this paper cites.
A. Waibel, T. Hanazawa, G. Hinton, K. Shikano, and K. J. Lang, “Phoneme recognition using time-delay neural networks,” IEEE transactions on acoustics, speech, and signal processing , vol. 37, no. 3, pp. 328–339, 1989
1989
Earlier work this paper cites.
J. J. Godfrey, E. C. Holliman, and J. McDaniel, “Switchboard: Telephone speech corpus for research and development,” in Acoustics, Speech, and Signal Processing, 1992. ICASSP-92., 1992 IEEE International Conference on , vol. 1. IEEE, 1992, pp. 517–520
1992
Earlier work this paper cites.
A. S. Bregman, Auditory scene analysis: The perceptual organization of sound . MIT press, 1994
1994
Earlier work this paper cites.
P. C. Woodland, J. J. Odell, V. Valtchev, and S. J. Young, “Large vocabulary continuous speech recognition using htk,” in Acoustics, Speech, and Signal Processing, 1994. ICASSP-94., 1994 IEEE International Conference on , vol. 2. Ieee, 1994, pp. II–125
1994
Earlier work this paper cites.
L. Deng, A. Acero, M. Plumpe, and X. Huang, “Large-vocabulary speech recognition under adverse acoustic environments.” in INTERSPEECH , 2000, pp. 806–809
2000
Earlier work this paper cites.
D. Povey, “Discriminative training for large vocabulary speech recognition,” Ph.D. dissertation, University of Cambridge, 2005
2005
Earlier work this paper cites.
D. Wang and G. J. Brown, Computational auditory scene analysis: Principles, algorithms, and applications . Wiley-IEEE press, 2006
2006
Earlier work this paper cites.
M. N. Schmidt and R. K. Olsson, “Single-channel speech separation using sparse non-negative matrix factorization,” in Interspeech 2006 , 2006
2006
Earlier work this paper cites.
S. F. Chen, B. Kingsbury, L. Mangu, D. Povey, G. Saon, H. Soltau, and G. Zweig, “Advances in speech transcription at ibm under the darpa ears program,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 14, no. 5, pp. 1596–1608, 2006
2006
Earlier work this paper cites.
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, “Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,” in Proceedings of the 23rd international conference on Machine learning . ACM, 2006, pp. 369–376
2006
Earlier work this paper cites.
D. Povey, D. Kanevsky, B. Kingsbury, B. Ramabhadran, G. Saon, and K. Visweswariah, “Boosted mmi for model and feature-space discriminative training,” in Acoustics, Speech and Signal Processing, 2008. ICASSP 2008. IEEE International Conference on . IEEE, 2008, pp. 4057–4060
2008
Earlier work this paper cites.
Y. Bengio, J. Louradour, R. Collobert, and J. Weston, “Curriculum learning,” in Proceedings of the 26th annual international conference on machine learning . ACM, 2009, pp. 41–48
2009
Earlier work this paper cites.
R. Hadsell, P. Sermanet, J. Ben, A. Erkan, M. Scoffier, K. Kavukcuoglu, U. Muller, and Y. LeCun, “Learning long-range vision for autonomous off-road driving,” Journal of Field Robotics , vol. 26, no. 2, pp. 120–144, 2009
2009
Earlier work this paper cites.
M. Cooke, J. R. Hershey, and S. J. Rennie, “Monaural speech separation and recognition challenge,” Computer Speech & Language , vol. 24, no. 1, pp. 1–15, 2010
2010
Earlier work this paper cites.
S. J. Pan and Q. Yang, “A survey on transfer learning,” IEEE Transactions on knowledge and data engineering , vol. 22, no. 10, pp. 1345–1359, 2010
2010
Cited alongside, same era.
T. Hain, L. Burget, J. Dines, P. N. Garner, F. Grézl, A. El Hannani, M. Huijbregts, M. Karafiat, M. Lincoln, and V. Wan, “Transcribing meetings with the amida systems,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 20, no. 2, pp. 486–498, 2012
2012
Cited alongside, same era.
A. Graves, N. Jaitly, and A.-r. Mohamed, “Hybrid speech recognition with deep bidirectional lstm,” in Automatic Speech Recognition and Understanding (ASRU), 2013 IEEE Workshop on . IEEE, 2013, pp. 273–278
2013
Cited alongside, same era.
D. Yu, K. Yao, H. Su, G. Li, and F. Seide, “Kl-divergence regularized deep neural network adaptation for improved large vocabulary speech recognition,” in Acoustics, Speech and Signal Processing (ICASSP), 2013 IEEE International Conference on . IEEE, 2013, pp. 7893–7897
2013
D. Povey, V. Peddinti, D. Galvez, P. Ghahremani, V. Manohar, X. Na, Y. Wang, and S. Khudanpur, “Purely sequence-trained neural networks for asr based on lattice-free mmi.” in INTERSPEECH , 2016, pp. 2751–2755
2016
Later among the works it cites.
2016
Later among the works it cites.
F. Seide and A. Agarwal, “Cntk: Microsoft’s open-source deep-learning toolkit,” in Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining . ACM, 2016, pp. 2135–2135
2016
Later among the works it cites.
D. Yu, W. Xiong, J. Droppo, A. Stolcke, G. Ye, J. Li, and G. Zweig, “Deep convolutional neural networks with layer-wise context expansion and attention.” in INTERSPEECH , 2016, pp. 17–21
2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
K. Veselỳ, A. Ghoshal, L. Burget, and D. Povey, “Sequence-discriminative training of deep neural networks.” in Interspeech , 2013, pp. 2345–2349
2013
Cited alongside, same era.
M. L. Seltzer, D. Yu, and Y. Wang, “An investigation of deep neural networks for noise robust speech recognition,” in Acoustics, Speech and Signal Processing (ICASSP), 2013 IEEE International Conference on . IEEE, 2013, pp. 7398–7402
2013
Cited alongside, same era.
J. Du, Y. Tu, Y. Xu, L. Dai, and C.-H. Lee, “Speech separation of a target speaker based on deep neural networks,” in Signal Processing (ICSP), 2014 12th International Conference on . IEEE, 2014, pp. 473–477
2014
Cited alongside, same era.
J. Ba and R. Caruana, “Do deep nets really need to be deep?” in Advances in neural information processing systems , 2014, pp. 2654–2662
2014
Cited alongside, same era.
J. Li, R. Zhao, J.-T. Huang, and Y. Gong, “Learning small-size dnn with output-distribution-based criteria,” in Fifteenth Annual Conference of the International Speech Communication Association , 2014
2014
Cited alongside, same era.
2014
Cited alongside, same era.
C. Weng, D. Yu, M. L. Seltzer, and J. Droppo, “Deep neural networks for single-channel multi-talker speech recognition,” IEEE/ACM Transactions on Audio, Speech and Language Processing (TASLP) , vol. 23, no. 10, pp. 1670–1679, 2015
2015
Cited alongside, same era.
T. N. Sainath, O. Vinyals, A. Senior, and H. Sak, “Convolutional, long short-term memory, fully connected deep neural networks,” in Acoustics, Speech and Signal Processing (ICASSP), 2015 IEEE International Conference on . IEEE, 2015, pp. 4580–4584
2015
Cited alongside, same era.
Z. Chen, Y. Luo, and N. Mesgarani, “Deep attractor network for single-microphone speaker separation,” in Acoustics, Speech and Signal Processing (ICASSP), 2017 IEEE International Conference on . IEEE, 2017, pp. 246–250
2017
Closest in time.
Y. Wang, J. Du, L.-R. Dai, and C.-H. Lee, “A gender mixture detection approach to unsupervised single-channel speech separation based on deep neural networks,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 25, no. 7, pp. 1535–1546, 2017
2017
Closest in time.
D. Yu, M. Kolbæk, Z.-H. Tan, and J. Jensen, “Permutation invariant training of deep models for speaker-independent multi-talker speech separation,” in Acoustics, Speech and Signal Processing (ICASSP), 2017 IEEE International Conference on . IEEE, 2017, pp. 241–245
2017
Closest in time.
2017
Closest in time.
J. Li, M. L. Seltzer, X. Wang, R. Zhao, and Y. Gong, “Large-scale domain adaptation via teacher-student learning,” in Interspeech 2017 , 2017
2017
Closest in time.
S. Watanabe, T. Hori, J. Le Roux, and J. R. Hershey, “Student-teacher network learning with enhanced features,” in Acoustics, Speech and Signal Processing (ICASSP), 2017 IEEE International Conference on . IEEE, 2017, pp. 5275–5279
2017
Closest in time.
S. Wang, K. Li, Z. Huang, S. M. Siniscalchi, and C.-H. Lee, “A transfer learning and progressive stacking approach to reducing deep model sizes with an application to speech enhancement,” in Acoustics, Speech and Signal Processing (ICASSP), 2017 IEEE International Conference on . IEEE, 2017, pp. 5575–5579
2017
Closest in time.
2017
Closest in time.
L. Lu and S. Renals, “Small-footprint highway deep neural networks for speech recognition,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 25, no. 7, pp. 1502–1511, July 2017
2017
Closest in time.
X. Zhang and D. Wang, “Binaural reverberant speech separation based on deep neural networks,” Proc. Interspeech 2017 , pp. 2018–2022, 2017
2017
Closest in time.
L. Drude and R. Haeb-Umbach, “Tight integration of spatial and spectral features for bss with deep clustering embeddings,” Proc. Interspeech 2017 , pp. 2650–2654, 2017
2017
Closest in time.