Fetching the paper…
Reading the bibliography…
This paper proposes a token-level serialized output training (t-SOT), a novel framework for streaming multi-talker automatic speech recognition (ASR).
Ö. Çetin and E. Shriberg, “Analysis of overlaps in meetings by dialog factors, hot spots, speakers, and collection site: Insights for automatic speech recognition,” in
2006
Earlier work this paper cites.
J. Carletta
2006
Earlier work this paper cites.
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, “Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,” in
2006
Earlier work this paper cites.
J. G. Fiscus, J. Ajot, N. Radde, and C. Laprun, “Multiple dimension Levenshtein edit distance calculations for evaluating automatic speech recognition systems during simultaneous speech,” in
2006
Earlier work this paper cites.
A. Graves, “Sequence transduction with recurrent neural networks,”
2012
Earlier work this paper cites.
J. Chorowski, D. Bahdanau, K. Cho, and Y. Bengio, “End-to-end continuous speech recognition using attention-based recurrent NN: First results,” in
2014
Earlier work this paper cites.
J. K. Chorowski
2015
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: an ASR corpus based on public domain audio books,” in
2015
Earlier work this paper cites.
J. R. Hershey
2016
Earlier work this paper cites.
Y. Isik, J. L. Roux, Z. Chen, S. Watanabe, and J. R. Hershey, “Single-channel multi-speaker separation using deep clustering,” in
2016
Earlier work this paper cites.
D. Yu, X. Chang, and Y. Qian, “Recognizing multi-talker speech with permutation invariant training,”
2017
Earlier work this paper cites.
D. Yu, M. Kolbæk, Z.-H. Tan, and J. Jensen, “Permutation invariant training of deep models for speaker-independent multi-talker speech separation,” in
2017
Earlier work this paper cites.
S. Watanabe, T. Hori, and J. R. Hershey, “Language independent end-to-end architecture for joint language identification and speech recognition,” in
2017
Earlier work this paper cites.
M. McAuliffe, M. Socolof, S. Mihuc, M. Wagner, and M. Sonderegger, “Montreal forced aligner: Trainable text-speech alignment using Kaldi,” in
2017
Earlier work this paper cites.
T. Yoshioka
2018
Cited alongside, same era.
J. Barker, S. Watanabe, E. Vincent, and J. Trmal, “The fifth ’CHiME’ speech separation and recognition challenge: Dataset, task and baselines,”
2018
Cited alongside, same era.
H. Seki, T. Hori, S. Watanabe, J. Le Roux, and J. R. Hershey, “A purely end-to-end system for multi-speaker speech recognition,” in
2018
Cited alongside, same era.
H. Inaguma
2018
Cited alongside, same era.
N. Kanda
2019
Cited alongside, same era.
T. Yoshioka
2019
Cited alongside, same era.
X. Chang, Y. Qian, K. Yu, and S. Watanabe, “End-to-end monaural multi-speaker ASR system without pretraining,” in
N. Kanda
2020
Later among the works it cites.
D. S. Park
2020
Later among the works it cites.
S. Chen, Y. Wu, Z. Chen, J. Wu, J. Li, T. Yoshioka, C. Wang, S. Liu, and M. Zhou, “Continuous speech separation with Conformer,” in
2021
Later among the works it cites.
N. Kanda
2021
Later among the works it cites.
N. Kanda
2021
Later among the works it cites.
N. Kanda
2021
Later among the works it cites.
L. Lu, N. Kanda, J. Li, and Y. Gong, “Streaming end-to-end multi-talker speech recognition,”
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2019
Cited alongside, same era.
X. Chang
2019
Cited alongside, same era.
L. El Shafey, H. Soltau, and I. Shafran, “Joint speech recognition and speaker diarization via sequence transduction,” in
2019
Cited alongside, same era.
S. Watanabe
2020
Cited alongside, same era.
Z. Chen, T. Yoshioka, L. Lu, T. Zhou, Z. Meng, Y. Luo, J. Wu, and J. Li, “Continuous speech separation: dataset and analysis,” in
2020
Cited alongside, same era.
A. Tripathi, H. Lu, and H. Sak, “End-to-end multi-talker overlapping speech recognition,” in
2020
Cited alongside, same era.
I. Sklyar, A. Piunova, and Y. Liu, “Streaming multi-speaker ASR with RNN-T,” in
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
X. Chen, Y. Wu, Z. Wang, S. Liu, and J. Li, “Developing real-time streaming transformer transducer for speech recognition on large-scale dataset,” in
2021
Later among the works it cites.
2021
Later among the works it cites.
2022
Closest in time.