Fetching the paper…
Reading the bibliography…
Continuous speech separation plays a vital role in complicated speech related tasks such as conversation transcription.
“Image method for efficiently simulating small-room acoustics,”
J. Allen and D. Berkley, · 1979
Earlier work this paper cites.
“CSR-II (WSJ1) Complete,” 1994,
Linguistic Data Consortium Philadelphia, · 1994
Earlier work this paper cites.
“Generating sensor signals in isotropic noise fields,”
Emanuël AP Habets and Sharon Gannot, · 2007
Earlier work this paper cites.
“A multichannel mmse-based framework for speech source separation and noise reduction,”
M. Souden, S. Araki, K. Kinoshita, T. Nakatani, and H. Sawada, · 2013
Earlier work this paper cites.
“On training targets for supervised speech separation,”
Yuxuan Wang, Arun Narayanan, and DeLiang Wang, · 2014
Earlier work this paper cites.
“Batch normalization: Accelerating deep network training by reducing internal covariate shift,”
Sergey Ioffe and Christian Szegedy, · 2015
Earlier work this paper cites.
“Deep clustering: Discriminative embeddings for segmentation and separation,”
John R Hershey, Zhuo Chen, Jonathan Le Roux, and Shinji Watanabe, · 2016
Earlier work this paper cites.
“Permutation invariant training of deep models for speaker-independent multi-talker speech separation,”
Dong Yu, Morten Kolbæk, Zheng-Hua Tan, and Jesper Jensen, · 2017
Earlier work this paper cites.
“Multitalker speech separation with utterance-level permutation invariant training of deep recurrent neural networks,”
Morten Kolbæk, Dong Yu, and Others, · 2017
Earlier work this paper cites.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin, · 2017
Earlier work this paper cites.
“Deep recurrent networks for separation and recognition of single-channel speech in nonstationary background audio,”
Hakan Erdogan, John R Hershey, Shinji Watanabe, and Jonathan Le Roux, · 2017
Earlier work this paper cites.
“Developing far-field speaker system via teacher-student learning,”
Jinyu Li, Rui Zhao, Zhuo Chen, Changliang Liu, Xiong Xiao, Guoli Ye, and Yifan Gong, · 2018
Cited alongside, same era.
“Speech-transformer: a no-recurrence sequence-to-sequence model for speech recognition,”
Linhao Dong, Shuang Xu, and Bo Xu, · 2018
Cited alongside, same era.
“Multi-microphone neural speech separation for far-field multi-talker speech recognition,”
Takuya Yoshioka, Hakan Erdogan, Zhuo Chen, and Fil Alleva, · 2018
Cited alongside, same era.
“Self-attention with relative position representations,”
Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani, · 2018
Cited alongside, same era.
“Recognizing overlapped speech in meetings: A multichannel separation approach using neural networks,”
Takuya Yoshioka, Hakan Erdogan, Zhuo Chen, Xiong Xiao, and Fil Alleva, · 2018
Cited alongside, same era.
“Decoupled weight decay regularization,”
“Transformer-xl: Attentive language models beyond a fixed-length context,”
Zihang Dai, Zhilin Yang, and Others, · 2019
Later among the works it cites.
“On the comparison of popular end-to-end models for large scale speech recognition,”
Jinyu Li, Yu Wu, Yashesh Gaur, Chengyi Wang, Rui Zhao, and Shujie Liu, · 2020
Closest in time.
“Chime-6 challenge: Tackling multispeaker speech recognition for unsegmented recordings,”
Shinji Watanabe, Michael Mandel, Jon Barker, and Emmanuel Vincent, · 2020
Closest in time.
“Dual-path rnn: efficient long sequence modeling for time-domain single-channel speech separation,”
Yi Luo, Zhuo Chen, and Takuya Yoshioka, · 2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ilya Loshchilov and Frank Hutter, · 2018
Cited alongside, same era.
“Improving noise robustness of automatic speech recognition via parallel data and teacher-student learning,”
Ladislav Mošner, Minhua Wu, et al., · 2019
Cited alongside, same era.
“A speaker-dependent approach to separation of far-field multi-talker microphone array speech for front-end processing in the chime-5 challenge,”
Lei Sun, Jun Du, et al., · 2019
Cited alongside, same era.
“Advances in online audio-visual meeting transcription,”
Takuya Yoshioka, Igor Abramovski, et al., · 2019
Cited alongside, same era.
“Conv-tasnet: Surpassing ideal time–frequency magnitude masking for speech separation,”
Yi Luo and Nima Mesgarani, · 2019
Cited alongside, same era.
“Streaming end-to-end speech recognition for mobile devices,”
Yanzhang He, Tara N Sainath, et al., · 2019
Cited alongside, same era.
Jingjing Chen, Qirong Mao, and Dong Liu, · 2020
Closest in time.
“End-to-end multi-speaker speech recognition with transformer,”
Xuankai Chang, Wangyou Zhang, Yanmin Qian, Jonathan Le Roux, and Shinji Watanabe, · 2020
Closest in time.
“Continuous speech separation: Dataset and analysis,”
Zhuo Chen, Takuya Yoshioka, et al., · 2020
Closest in time.
“Transformer transducer: A streamable speech recognition model with transformer encoders and RNN-T loss,”
Qian Zhang, Han Lu, Hasim Sak, et al., · 2020
Closest in time.
“Conformer: Convolution-augmented transformer for speech recognition,”
Anmol Gulati, James Qin, et al., · 2020
Closest in time.
“Semantic mask for transformer based end-to-end speech recognition,”
Chengyi Wang, Yu Wu, Yujiao Du, Jinyu Li, Shujie Liu, Liang Lu, Shuo Ren, Guoli Ye, Sheng Zhao, and Ming Zhou, · 2020
Closest in time.