Fetching the paper…
Reading the bibliography…
With its strong modeling capacity that comes from a multi-head and multi-layer structure, Transformer is a very powerful model for learning a sequential representation and has been successfully applied to speech separation recently.
“CSR-II (WSJ1) Complete,” 1994,
Linguistic Data Consortium Philadelphia, · 1994
Earlier work this paper cites.
“Room impulse response generator,”
Emanuel AP Habets, · 2006
Earlier work this paper cites.
“Generating sensor signals in isotropic noise fields,”
Emanuël AP Habets and Sharon Gannot, · 2007
Earlier work this paper cites.
“On training targets for supervised speech separation,”
Yuxuan Wang, Arun Narayanan, and DeLiang Wang, · 2014
Earlier work this paper cites.
“Adam: A method for stochastic optimization,”
Diederik P Kingma and Jimmy Ba, · 2014
Earlier work this paper cites.
“Deep clustering: Discriminative embeddings for segmentation and separation,”
John R Hershey, Zhuo Chen, Jonathan Le Roux, and Shinji Watanabe, · 2016
Earlier work this paper cites.
“Single-channel multi-speaker separation using deep clustering,”
Yusuf Isik, Jonathan Le Roux, Zhuo Chen, Shinji Watanabe, and John R Hershey, · 2016
Earlier work this paper cites.
“Permutation invariant training of deep models for speaker-independent multi-talker speech separation,”
Dong Yu, Morten Kolbæk, Zheng-Hua Tan, and Jesper Jensen, · 2017
Earlier work this paper cites.
“Multitalker speech separation with utterance-level permutation invariant training of deep recurrent neural networks,”
Morten Kolbæk, Dong Yu, Zheng-Hua Tan, and Jesper Jensen, · 2017
Earlier work this paper cites.
“Deep recurrent networks for separation and recognition of single-channel speech in nonstationary background audio,”
Hakan Erdogan, John R Hershey, Shinji Watanabe, and Jonathan Le Roux, · 2017
Earlier work this paper cites.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin, · 2017
Cited alongside, same era.
“Speech-transformer: a no-recurrence sequence-to-sequence model for speech recognition,”
Linhao Dong, Shuang Xu, and Bo Xu, · 2018
Cited alongside, same era.
“Self-attention with relative position representations,”
Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani, · 2018
Cited alongside, same era.
“Multi-microphone neural speech separation for far-field multi-talker speech recognition,”
Takuya Yoshioka, Hakan Erdogan, Zhuo Chen, and Fil Alleva, · 2018
Cited alongside, same era.
“Decoupled weight decay regularization,”
Ilya Loshchilov and Frank Hutter, · 2018
Cited alongside, same era.
“Advances in online audio-visual meeting transcription,”
“A comparative study on transformer vs rnn in speech applications,”
Shigeki Karita, Nanxin Chen, Tomoki Hayashi, Takaaki Hori, Hirofumi Inaguma, Ziyan Jiang, Masao Someki, Nelson Enrique Yalta Soplin, Ryuichi Yamamoto, Xiaofei Wang, et al., · 2019
Later among the works it cites.
“Semantic mask for transformer based end-to-end speech recognition,”
Chengyi Wang, Yu Wu, Yujiao Du, Jinyu Li, Shujie Liu, Liang Lu, Shuo Ren, Guoli Ye, Sheng Zhao, and Ming Zhou, · 2019
Later among the works it cites.
“Shallow-deep networks: Understanding and mitigating network overthinking,”
Yigitcan Kaya, Sanghyun Hong, and Tudor Dumitras, · 2019
Later among the works it cites.
“Dual-path rnn: efficient long sequence modeling for time-domain single-channel speech separation,”
Yi Luo, Zhuo Chen, and Takuya Yoshioka, · 2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Takuya Yoshioka, Igor Abramovski, Cem Aksoylar, Zhuo Chen, Moshe David, Dimitrios Dimitriadis, Yifan Gong, Ilya Gurvich, Xuedong Huang, Yan Huang, et al., · 2019
Cited alongside, same era.
“Low-latency speaker-independent continuous speech separation,”
Takuya Yoshioka, Zhuo Chen, Changliang Liu, Xiong Xiao, Hakan Erdogan, and Dimitrios Dimitriadis, · 2019
Cited alongside, same era.
“Conv-tasnet: Surpassing ideal time–frequency magnitude masking for speech separation,”
Yi Luo and Nima Mesgarani, · 2019
Cited alongside, same era.
“Very deep self-attention networks for end-to-end speech recognition,”
Ngoc-Quan Pham, Thai-Son Nguyen, Jan Niehues, Markus Müller, Sebastian Stüker, and Alexander Waibel, · 2019
Cited alongside, same era.
Ziqiang Shi, Rujie Liu, and Jiqing Han, · 2020
Closest in time.
“On the comparison of popular end-to-end models for large scale speech recognition,”
Jinyu Li, Yu Wu, Yashesh Gaur, Chengyi Wang, Rui Zhao, and Shujie Liu, · 2020
Closest in time.
“End-to-end multi-speaker speech recognition with transformer,”
Xuankai Chang, Wangyou Zhang, Yanmin Qian, Jonathan Le Roux, and Shinji Watanabe, · 2020
Closest in time.
“Continuous speech separation with conformer,”
Sanyuan Chen, Yu Wu, Zhuo Chen, Jinyu Li, Chengyi Wang, Shujie Liu, and Ming Zhou, · 2020
Closest in time.
“Continuous speech separation: Dataset and analysis,”
Zhuo Chen, Takuya Yoshioka, Liang Lu, Tianyan Zhou, Zhong Meng, Yi Luo, Jian Wu, Xiong Xiao, and Jinyu Li, · 2020
Closest in time.