Fetching the paper…
Reading the bibliography…
Streaming recognition of multi-talker conversations has so far been evaluated only for 2-speaker single-turn sessions.
“The hungarian method for the assignment problem,”
H. Kuhn, · 1955
Earlier work this paper cites.
“Observations on overlap: findings and implications for automatic processing of multi-party conversation,”
E. Shriberg, A. Stolcke, and D. Baron, · 2001
Earlier work this paper cites.
“The AMI meeting corpus: A pre-announcement,”
J. Carletta et al., · 2005
Earlier work this paper cites.
“Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,”
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, · 2006
Earlier work this paper cites.
“Sequence transduction with recurrent neural networks,”
A. Graves, · 2012
Earlier work this paper cites.
“The REVERB challenge: A common evaluation framework for dereverberation and recognition of reverberant speech,”
K. Kinoshita, M. Delcroix, T. Yoshioka, T. Nakatani, A. Sehr, W. Kellermann, and R. Maas, · 2013
Earlier work this paper cites.
“The third ‘CHiME’ speech separation and recognition challenge: Dataset, task and baselines,”
J. Barker, R. Marxer, E. Vincent, and S. Watanabe, · 2015
Earlier work this paper cites.
“Attention-based models for speech recognition,”
J. Chorowski, D. Bahdanau, D. Serdyuk, K. Cho, and Y. Bengio, · 2015
Earlier work this paper cites.
“Deep speech 2 : End-to-end speech recognition in english and mandarin,”
D. Amodei et al., · 2016
Earlier work this paper cites.
“On training the recurrent neural network encoder-decoder for large vocabulary end-to-end speech recognition,”
L. Lu, X. Zhang, and S. Renals, · 2016
Earlier work this paper cites.
“Toward human parity in conversational speech recognition,”
W. Xiong, J. Droppo, X. Huang, F. Seide, M. Seltzer, A. Stolcke, D. Yu, and G. Zweig, · 2017
Earlier work this paper cites.
“State-of-the-art speech recognition with sequence-to-sequence models,”
C. Chiu, T. Sainath, Y. Wu, R. Prabhavalkar, P. Nguyen, Z. Chen, A. Kannan, R. J. Weiss, K. Rao, K. Gonina, N. Jaitly, B. Li, J. Chorowski, and M. Bacchiani, · 2018
Cited alongside, same era.
“End-to-end multi-speaker speech recognition,”
S. Settle, J. Le Roux, T. Hori, S. Watanabe, and J. Hershey, · 2018
Cited alongside, same era.
“Meeting transcription using asynchronous distant microphones,”
T. Yoshioka, D. Dimitriadis, A. Stolcke, W. Hinthorn, Z. Chen, M. Zeng, and X. Huang, · 2019
Cited alongside, same era.
“Streaming end-to-end speech recognition for mobile devices,”
Y. He, T. N. Sainath, R. Prabhavalkar, I. McGraw, R. Alvarez, D. Zhao, D. Rybach, A. Kannan, Y. Wu, R. Pang, et al., · 2019
Cited alongside, same era.
“Improving RNN transducer modeling for end-to-end speech recognition,”
J. Li, R. Zhao, H. Hu, and Y. Gong, · 2019
Cited alongside, same era.
“End-to-end monaural multi-speaker ASR system without pretraining,”
“Dual-path RNN: Efficient long sequence modeling for time-domain single-channel speech separation,”
Y. Luo, Z. Chen, and T. Yoshioka, · 2020
Later among the works it cites.
“Dual-path transformer network: Direct context-aware modeling for end-to-end monaural speech separation,”
J-J. Chen, Q. Mao, and D. Liu, · 2020
Later among the works it cites.
“Continuous speech separation: Dataset and analysis,”
Z. Chen, T. Yoshioka, L. Lu, T. Zhou, Z. Meng, Y. Luo, J. Wu, and J. Li, · 2020
Later among the works it cites.
“Blockwise self-attention for long document understanding,”
J. Qiu, H. Ma, O. Levy, W-T. Yih, S. Wang, and J. Tang, · 2020
Later among the works it cites.
“O(n) connections are expressive enough: Universal approximability of sparse transformers,”
C. Yun, Y-W. Chang, S. Bhojanapalli, A. S. Rawat, S. Reddi, and S. Kumar, · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
X. Chang, Y. Qian, K. Yu, and S. Watanabe, · 2019
Cited alongside, same era.
“Generating long sequences with sparse transformers,”
R. Child, S. Gray, A. Radford, and I. Sutskever, · 2019
Cited alongside, same era.
“Axial attention in multidimensional transformers,”
J. Ho, N. Kalchbrenner, D. Weissenborn, and T. Salimans, · 2019
Cited alongside, same era.
“CHiME-6 challenge: Tackling multispeaker speech recognition for unsegmented recordings,”
S. Watanabe, M. Mandel, J. Barker, and E. Vincent, · 2020
Cited alongside, same era.
“Serialized output training for end-to-end overlapped speech recognition,”
N. Kanda, Y. Gaur, X. Wang, Z. Meng, and T. Yoshioka, · 2020
Cited alongside, same era.
“End-to-end multi-talker overlapping speech recognition,”
A. Tripathi, H. Lu, and H. Sak, · 2020
Cited alongside, same era.
C. Wang, Y. Wu, Y. Du, J. Li, S. Liu, L. Lu, S. Ren, G. Ye, S. Zhao, and M. Zhou, · 2020
Later among the works it cites.
“Recent advances in end-to-end automatic speech recognition,”
J. Li, · 2021
Closest in time.
“Streaming multi-speaker ASR with RNN-T,”
I. Sklyar, A. Piunova, and Y. Liu, · 2021
Closest in time.
“Streaming end-to-end multi-talker speech recognition,”
L. Lu, N. Kanda, J. Li, and Y. Gong, · 2021
Closest in time.
“End-to-end speaker-attributed ASR with transformer,”
N. Kanda, G. Ye, Y. Gaur, X. Wang, Z. Meng, Z. Chen, and T. Yoshioka, · 2021
Closest in time.
“Continuous speech separation with conformer,”
S. Chen, Y. Wu, Z. Chen, J. Li, C. Wang, S. Liu, and M. Zhou, · 2021
Closest in time.