Fetching the paper…
Reading the bibliography…
This paper proposes serialized output training (SOT), a novel framework for multi-speaker overlapped speech recognition based on an attention-based encoder-decoder approach.
H. W. Kuhn, “The Hungarian method for the assignment problem,”
1955
Earlier work this paper cites.
F. Seide, G. Li, and D. Yu, “Conversational speech transcription using context-dependent deep neural networks,” in
2011
Earlier work this paper cites.
D. Povey, A. Ghoshal, G. Boulianne, L. Burget, O. Glembek, N. Goel, M. Hannemann, P. Motlicek, Y. Qian, P. Schwarz
2011
Earlier work this paper cites.
G. Hinton, L. Deng, D. Yu, G. E. Dahl, A.-r. Mohamed, N. Jaitly, A. Senior, V. Vanhoucke, P. Nguyen, T. N. Sainath
2012
Earlier work this paper cites.
2014
Earlier work this paper cites.
J. Chorowski, D. Bahdanau, K. Cho, and Y. Bengio, “End-to-end continuous speech recognition using attention-based recurrent NN: First results,” in
2014
Earlier work this paper cites.
C. Weng, D. Yu, M. L. Seltzer, and J. Droppo, “Deep neural networks for single-channel multi-talker speech recognition,”
2015
Earlier work this paper cites.
J. K. Chorowski, D. Bahdanau, D. Serdyuk, K. Cho, and Y. Bengio, “Attention-based models for speech recognition,” in
2015
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: an ASR corpus based on public domain audio books,” in
2015
Earlier work this paper cites.
T. Ko, V. Peddinti, D. Povey, and S. Khudanpur, “Audio augmentation for speech recognition,” in
2015
Earlier work this paper cites.
2016
Earlier work this paper cites.
D. Amodei, S. Ananthanarayanan, R. Anubhai, J. Bai, E. Battenberg, C. Case, J. Casper, B. Catanzaro, Q. Cheng, G. Chen
2016
Earlier work this paper cites.
D. Povey, V. Peddinti, D. Galvez, P. Ghahremani, V. Manohar, X. Na, Y. Wang, and S. Khudanpur, “Purely sequence-trained neural networks for ASR based on lattice-free MMI,” in
2016
Cited alongside, same era.
W. Chan, N. Jaitly, Q. Le, and O. Vinyals, “Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,” in
2016
Cited alongside, same era.
J. L. Ba, J. R. Kiros, and G. E. Hinton, “Layer normalization,”
2016
Cited alongside, same era.
G. Saon, G. Kurata, T. Sercu, K. Audhkhasi, S. Thomas, D. Dimitriadis, X. Cui, B. Ramabhadran, M. Picheny, L.-L. Lim
2017
Cited alongside, same era.
D. Yu, M. Kolbæk, Z.-H. Tan, and J. Jensen, “Permutation invariant training of deep models for speaker-independent multi-talker speech separation,” in
2017
C. Lüscher, E. Beck, K. Irie, M. Kitza, W. Michel, A. Zeyer, R. Schlüter, and H. Ney, “RWTH ASR systems for LibriSpeech: Hybrid vs attention,” in
2019
Later among the works it cites.
D. S. Park, W. Chan, Y. Zhang, C.-C. Chiu, B. Zoph, E. D. Cubuk, and Q. V. Le, “Specaugment: A simple data augmentation method for automatic speech recognition,” in
2019
Later among the works it cites.
2019
Later among the works it cites.
N. Kanda, S. Horiguchi, R. Takashima, Y. Fujita, K. Nagamatsu, and S. Watanabe, “Auxiliary interference speaker loss for target-speaker speech recognition,” in
2019
Later among the works it cites.
X. Chang, Y. Qian, K. Yu, and S. Watanabe, “End-to-end monaural multi-speaker ASR system without pretraining,” in
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
D. Yu, X. Chang, and Y. Qian, “Recognizing multi-talker speech with permutation invariant training,”
2017
Cited alongside, same era.
S. Watanabe, T. Hori, and J. R. Hershey, “Language independent end-to-end architecture for joint language identification and speech recognition,” in
2017
Cited alongside, same era.
S. Kim, T. Hori, and S. Watanabe, “Joint CTC-attention based end-to-end speech recognition using multi-task learning,” in
2017
Cited alongside, same era.
N. Kanda, Y. Fujita, and K. Nagamatsu, “Lattice-free state-level minimum Bayes risk training of acoustic models.” in
2018
Cited alongside, same era.
H. Seki, T. Hori, S. Watanabe, J. Le Roux, and J. R. Hershey, “A purely end-to-end system for multi-speaker speech recognition,” in
2018
Cited alongside, same era.
H. Seki, S. Watanabe, T. Hori, J. Le Roux, and J. R. Hershey, “An end-to-end language-tracking speech recognizer for mixed-language speech,” in
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2019
Later among the works it cites.
X. Chang, W. Zhang, Y. Qian, J. L. Roux, and S. Watanabe, “MIMO-SPEECH: End-to-end multi-channel multi-speaker speech recognition,” in
2019
Later among the works it cites.
L. El Shafey, H. Soltau, and I. Shafran, “Joint speech recognition and speaker diarization via sequence transduction,” in
2019
Later among the works it cites.
K. Irie, A. Zeyer, R. Schlüter, and H. Ney, “Language modeling with deep transformers,” in
2019
Later among the works it cites.
A. Zeyer, P. Bahar, K. Irie, R. Schlüter, and H. Ney, “A comparison of transformer and LSTM encoder decoder models for ASR,” in
2019
Later among the works it cites.
A. Rosenberg, B. Ramabhadran, P. Moreno, Y. Jia, Y. Wu, Y. Zhang, and Z. Wu, “Speech recognition with augmented synthesized speech,” 2019, pp. 996–1002
2019
Later among the works it cites.
A. Tripathi, H. Lu, and H. Sak, “End-to-end multi-talker overlapping speech recognition,” in
2020
Closest in time.