Fetching the paper…
Reading the bibliography…
Recently, there has been growing interest in multi-speaker speech recognition, where the utterances of multiple speakers are recognized from their mixture.
Multitalker speech separation with utterance-level permutation invariant training of deep recurrent neural networks
Morten Kolbæk, Dong Yu, Zheng-Hua Tan, and Jesper Jensen. 2017 · 1913
Earlier work this paper cites.
CSR-II (wsj1) complete
Linguistic Data Consortium. 1994 · 1994
Earlier work this paper cites.
Corpus of Spontaneous Japanese: Its design and evaluation
Kikuo Maekawa. 2003 · 2003
Earlier work this paper cites.
Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber. 2006 · 2006
Earlier work this paper cites.
CSR-I (wsj0) complete
John Garofalo, David Graff, Doug Paul, and David Pallett. 2007 · 2007
Earlier work this paper cites.
Monaural speech separation and recognition challenge
Martin Cooke, John R Hershey, and Steven J Rennie. 2009 · 2009
Earlier work this paper cites.
The kaldi speech recognition toolkit
Daniel Povey, Arnab Ghoshal, Gilles Boulianne, Lukas Burget, Ondrej Glembek, Nagendra Goel, Mirko Hannemann, Petr Motlicek, Yanmin Qian, Petr Schwarz, Jan Silovsky, Georg Stemmer, and Karel Vesely. 2011 · 2011
Earlier work this paper cites.
ADADELTA: an adaptive learning rate method
Matthew D Zeiler. 2012 · 2012
Earlier work this paper cites.
On the difficulty of training recurrent neural networks
Razvan Pascanu, Tomas Mikolov, and Yoshua Bengio. 2013 · 2013
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman. 2014 · 2014
Cited alongside, same era.
Attention-based models for speech recognition
Jan K Chorowski, Dzmitry Bahdanau, Dmitriy Serdyuk, Kyunghyun Cho, and Yoshua Bengio. 2015 · 2015
Cited alongside, same era.
Kaldi recipe for Japanese spontaneous speech recognition and its evaluation
Takafumi Moriya, Takahiro Shinozaki, and Shinji Watanabe. 2015 · 2015
Cited alongside, same era.
Chainer: a next-generation open source framework for deep learning
Seiya Tokui, Kenta Oono, Shohei Hido, and Justin Clayton. 2015 · 2015
Cited alongside, same era.
End-to-end attention-based large vocabulary speech recognition
ChainerMN: Scalable Distributed Deep Learning Framework
Takuya Akiba, Keisuke Fukuda, and Shuji Suzuki. 2017 · 2017
Later among the works it cites.
Deep attractor network for single-microphone speaker separation
Zhuo Chen, Yi Luo, and Nima Mesgarani. 2017 · 2017
Later among the works it cites.
Joint CTC-attention based end-to-end speech recognition using multi-task learning
Suyoun Kim, Takaaki Hori, and Shinji Watanabe. 2017 · 2017
Later among the works it cites.
Single-channel multi-talker speech recognition with permutation invariant training
Yanmin Qian, Xuankai Chang, and Dong Yu. 2017 · 2017
Later among the works it cites.
Permutation invariant training of deep models for speaker-independent multi-talker speech separation
Dong Yu, Morten Kolbæk, Zheng-Hua Tan, and Jesper Jensen. 2017 · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dzmitry Bahdanau, Jan Chorowski, Dmitriy Serdyuk, Philemon Brakel, and Yoshua Bengio. 2016 · 2016
Cited alongside, same era.
Deep clustering: Discriminative embeddings for segmentation and separation
John R Hershey, Zhuo Chen, Jonathan Le Roux, and Shinji Watanabe. 2016 · 2016
Cited alongside, same era.
Single-channel multi-speaker separation using deep clustering
Yusuf Isik, Jonathan Le Roux, Zhuo Chen, Shinji Watanabe, and John R. Hershey. 2016 · 2016
Cited alongside, same era.
Joint CTC/attention decoding for end-to-end speech recognition
Takaaki Hori, Shinji Watanabe, and John R Hershey. 2017a
Cited in the paper.
Multi-level language modeling and decoding for open vocabulary end-to-end speech recognition
Takaaki Hori, Shinji Watanabe, and John R Hershey. 2017b
Cited in the paper.
Advances in joint CTC-Attention based end-to-end speech recognition with a deep CNN encoder and RNN-LM
Takaaki Hori, Shinji Watanabe, Yu Zhang, and Chan William. 2017c
Cited in the paper.
Progressive joint modeling in unsupervised single-channel overlapped speech recognition
Zhehuai Chen, Jasha Droppo, Jinyu Li, and Wayne Xiong. 2018 · 2018
Closest in time.
End-to-end multi-speaker speech recognition
Shane Settle, Jonathan Le Roux, Takaaki Hori, Shinji Watanabe, and John R. Hershey. 2018 · 2018
Closest in time.