Fetching the paper…
Reading the bibliography…
Automatic speech recognition (ASR) of single channel far-field recordings with an unknown number of speakers is traditionally tackled by cascaded modules.
L. Rabiner and B.-H. Juang,
1993
Earlier work this paper cites.
X. Huang, A. Acero, H.-W. Hon, and R. Reddy,
2001
Earlier work this paper cites.
T. Hain, L. Burget, J. Dines, G. Garau, V. Wan, M. Karafiat, J. Vepa, and M. Lincoln, “The AMI system for the transcription of speech in meetings,” in
2007
Earlier work this paper cites.
A. Graves, “Sequence transduction with recurrent neural networks,”
2012
Earlier work this paper cites.
C. Fox, Y. Liu, E. Zwyssig, and T. Hain, “The Sheffield wargames corpus,” in
2013
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: An ASR corpus based on public domain audio books,” in
2015
Earlier work this paper cites.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in
2015
Earlier work this paper cites.
Y. Liu, C. Fox, M. Hasan, and T. Hain, “The sheffield wargame corpus - day two and day three,”
2016
Earlier work this paper cites.
Y. Isik, J. L. Roux, Z. Chen, S. Watanabe, and J. R. Hershey, “Single-channel multi-speaker separation using deep clustering,”
2016
Earlier work this paper cites.
J. L. Ba, J. R. Kiros, and G. E. Hinton, “Layer normalization,” in
2016
Earlier work this paper cites.
D. Yu, M. Kolbaek, Z.-H. Tan, and J. Jensen, “Permutation invariant training of deep models for speaker-independent multi-talker speech separation,”
2017
Earlier work this paper cites.
D. Yu, X. Chang, and Y. Qian, “Recognizing multi-talker speech with permutation invariant training,”
2017
Earlier work this paper cites.
J. Barker, S. Watanabe, E. Vincent, and J. Trmal, “The fifth ‘CHiME’ speech separation and recognition challenge: Dataset, task and baselines,”
2018
Cited alongside, same era.
S. Settle, J. L. Roux, T. Hori, S. Watanabe, and J. R. Hershey, “End-to-end multi-speaker speech recognition,” in
2018
Cited alongside, same era.
Y. Qian, X. Chang, and D. Yu, “Single-channel multi-talker speech recognition with permutation invariant training,”
2018
Cited alongside, same era.
H. Seki, T. Hori, S. Watanabe, J. Le Roux, and J. R. Hershey, “A purely end-to-end system for multi-speaker speech recognition,”
2018
Cited alongside, same era.
T. Kudo and J. Richardson, “SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing,” in
2018
Cited alongside, same era.
N. Kanda, Y. Gaur, X. Wang, Z. Meng, Z. Chen, T. Zhou, and T. Yoshioka, “Joint speaker counting, speech recognition, and speaker identification for overlapped speech of any number of speakers,”
2020
Later among the works it cites.
D. S. Park, Y. Zhang, C.-C. Chiu, Y. Chen, B. Li, W. Chan, Q. V. Le, and Y. Wu, “SpecAugment on large scale datasets,”
2020
Later among the works it cites.
Z. Chen, T. Yoshioka, L. Lu, T. Zhou, Z. Meng, Y. Luo, J. Wu, X. Xiao, and J. Li, “Continuous speech separation: Dataset and analysis,” in
2020
Later among the works it cites.
N. Kanda, G. Ye, Y. Wu, Y. Gaur, X. Wang, Z. Meng, Z. Chen, and T. Yoshioka, “Large-scale pre-training of end-to-end multi-talker ASR for meeting transcription with single distant microphone,” in
2021
Closest in time.
N. Kanda, X. Chang, Y. Gaur, X. Wang, Z. Meng, Z. Chen, and T. Yoshioka, “Investigation of end-to-end speaker-attributed ASR for continuous multi-talker recordings,” in
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
T. Menne, I. Sklyar, R. Schlüter, and H. Ney, “Analysis of deep clustering as preprocessing for automatic speech recognition of sparsely overlapping speech,”
2019
Cited alongside, same era.
X. Chang, Y. Qian, K. Yu, and S. Watanabe, “End-to-end monaural multi-speaker ASR system without pretraining,”
2019
Cited alongside, same era.
D. S. Park, W. Chan, Y. Zhang, C.-C. Chiu, B. Zoph, E. D. Cubuk, and Q. V. Le, “SpecAugment: A simple data augmentation method for automatic speech recognition,”
2019
Cited alongside, same era.
T. von Neumann, K. Kinoshita, L. Drude, C. Boeddeker, M. Delcroix, T. Nakatani, and R. Haeb-Umbach, “End-to-end training of time domain audio separation and recognition,”
2020
Cited alongside, same era.
T. v. Neumann, C. Boeddeker, L. Drude, K. Kinoshita, M. Delcroix, T. Nakatani, and R. Haeb-Umbach, “Multi-talker ASR for an unknown number of sources: Joint training of source counting, separation and ASR,”
2020
Cited alongside, same era.
A. Tripathi, H. Lu, and H. Sak, “End-to-end multi-talker overlapping speech recognition,” in
2020
Cited alongside, same era.
N. Kanda, Y. Gaur, X. Wang, Z. Meng, and T. Yoshioka, “Serialized output training for end-to-end overlapped speech recognition,”
2020
Cited alongside, same era.
2021
Closest in time.
N. Kanda, G. Ye, Y. Gaur, X. Wang, Z. Meng, Z. Chen, and T. Yoshioka, “End-to-end speaker-attributed ASR with transformer,”
2021
Closest in time.
N. Kanda, X. Xiao, J. Wu, T. Zhou, Y. Gaur, X. Wang, Z. Meng, Z. Chen, and T. Yoshioka, “A comparative study of modular and joint approaches for speaker-attributed asr on monaural long-form audio,”
2021
Closest in time.
I. Sklyar, A. Piunova, and Y. Liu, “Streaming multi-speaker ASR with RNN-T,” in
2021
Closest in time.
L. Lu, N. Kanda, J. Li, and Y. Gong, “Streaming end-to-end multi-talker speech recognition,”
2021
Closest in time.
L. Lu, N. Kanda, J. Li, and Y.Gong, “Streaming multi-talker speech recognition with joint speaker identification,” in
2021
Closest in time.
T. von Neumann, K. Kinoshita, C. Boeddeker, M. Delcroix, and R. Haeb-Umbach, “Graph-PIT: Generalized permutation invariant training for continuous separation of arbitrary numbers of speakers,” in
2021
Closest in time.