Fetching the paper…
Reading the bibliography…
Transformers have been the most successful architecture for various speech modeling tasks, including speech separation.
“A new approach to linear filtering and prediction problems,”
Rudolph Emil Kalman, · 1960
Earlier work this paper cites.
“Backpropagation applied to handwritten zip code recognition,”
Y. LeCun, B. Boser, J. S. Denker, D. Henderson, R. E. Howard, W. Hubbard, and L. D. Jackel, · 1989
Earlier work this paper cites.
“Long short-term memory,”
Long Short-Term Memory, · 2010
Earlier work this paper cites.
“Adam: A method for stochastic optimization,”
Diederik P. Kingma and Jimmy Ba, · 2015
Earlier work this paper cites.
“Quantifying the vanishing gradient and long distance dependency problem in recursive neural networks and recursive LSTMs,”
Phong Le and Willem Zuidema, · 2016
Earlier work this paper cites.
“Deep clustering: Discriminative embeddings for segmentation and separation,”
John R. Hershey, Zhuo Chen, Jonathan Le Roux, and Shinji Watanabe, · 2016
Earlier work this paper cites.
“Gaussian error linear units (gelus),”
Dan Hendrycks and Kevin Gimpel, · 2016
Earlier work this paper cites.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin, · 2017
Earlier work this paper cites.
“Searching for activation functions,”
Prajit Ramachandran, Barret Zoph, and Quoc V Le, · 2017
Earlier work this paper cites.
“End-to-end source separation with adaptive front-ends,”
Shrikant Venkataramani, Jonah Casebeer, and Paris Smaragdis, · 2018
Earlier work this paper cites.
“Tasnet: time-domain audio separation network for real-time, single-channel speech separation,”
Yi Luo and Nima Mesgarani, · 2018
Earlier work this paper cites.
“Conv-tasnet: Surpassing ideal time–frequency magnitude masking for speech separation,”
Yi Luo and Nima Mesgarani, · 2019
Earlier work this paper cites.
“Root mean square layer normalization,”
Biao Zhang and Rico Sennrich, · 2019
Earlier work this paper cites.
“Sdr–half-baked or well done?,”
Jonathan Le Roux, Scott Wisdom, Hakan Erdogan, and John R Hershey, · 2019
Cited alongside, same era.
“Dual-path rnn: efficient long sequence modeling for time-domain single-channel speech separation,”
Yi Luo, Zhuo Chen, and Takuya Yoshioka, · 2020
Cited alongside, same era.
“Voice separation with an unknown number of multiple speakers,”
Eliya Nachmani, Yossi Adi, and Lior Wolf, · 2020
Cited alongside, same era.
“Dual-Path Transformer Network: Direct Context-Aware Modeling for End-to-End Monaural Speech Separation,”
Jingjing Chen, Qirong Mao, and Dong Liu, · 2020
Cited alongside, same era.
“Sudo rm-rf: Efficient networks for universal audio source separation,”
Efthymios Tzinis, Zhepei Wang, and Paris Smaragdis, · 2020
Cited alongside, same era.
“Attention is all you need in speech separation,”
Cem Subakan, Mirco Ravanelli, Samuele Cornell, Mirko Bronzi, and Jianyuan Zhong, · 2021
“Mamba: Linear-time sequence modeling with selective state spaces,”
Albert Gu and Tri Dao, · 2023
Later among the works it cites.
“On time domain conformer models for monaural speech separation in noisy reverberant acoustic environments,”
William Ravenscroft, Stefan Goetze, and Thomas Hain, · 2023
Later among the works it cites.
“Mossformer: Pushing the performance limit of monaural speech separation using gated single-head transformer with convolution-augmented joint self-attentions,”
Shengkui Zhao and Bin Ma, · 2023
Later among the works it cites.
“Mossformer2: Combining transformer and rnn-free recurrent network for enhanced time-domain monaural speech separation,” 2023
Shengkui Zhao, Yukun Ma, Chongjia Ni, Chong Zhang, Hao Wang, Trung Hieu Nguyen, Kun Zhou, Jiaqi Yip, Dianwen Ng, and Bin Ma, · 2023
Later among the works it cites.
“A neural state-space model approach to efficient speech separation,” 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Combining recurrent, convolutional, and continuous-time models with linear state space layers,”
Albert Gu, Isys Johnson, Karan Goel, Khaled Saab, Tri Dao, Atri Rudra, and Christopher Ré, · 2021
Cited alongside, same era.
“Wavesplit: End-to-end speech separation by speaker clustering,”
Neil Zeghidour and David Grangier, · 2021
Cited alongside, same era.
“SpeechBrain: A general-purpose speech toolkit,” 2021,
Mirco Ravanelli, Titouan Parcollet, Peter Plantinga, Aku Rouhe, Samuele Cornell, Loren Lugosch, Cem Subakan, Nauman Dawalatabad, Abdelwahab Heba, Jianyuan Zhong, Ju-Chieh Chou, Sung-Lin Yeh, Szu-Wei Fu, Chien-Feng Liao, Elena Rastorgueva, François Grondin, William Aris, Hwidong Na, Yan Gao, Renato De Mori, and Yoshua Bengio, · 2021
Cited alongside, same era.
“Exploring self-attention mechanisms for speech separation,”
Cem Subakan, Mirco Ravanelli, Samuele Cornell, François Grondin, and Mirko Bronzi, · 2022
Cited alongside, same era.
“QDPN - Quasi-dual-path Network for single-channel Speech Separation,”
Joel Rixen and Matthias Renz, · 2022
Cited alongside, same era.
“Efficiently modeling long sequences with structured state spaces,”
Albert Gu, Karan Goel, and Christopher Ré, · 2022
Cited alongside, same era.
Chen Chen, Chao-Han Huck Yang, Kai Li, Yuchen Hu, Pin-Jui Ku, and Eng Siong Chng, · 2023
Later among the works it cites.
“Vision mamba: Efficient visual representation learning with bidirectional state space model,”
Lianghui Zhu, Bencheng Liao, Qian Zhang, Xinlong Wang, Wenyu Liu, and Xinggang Wang, · 2024
Closest in time.
“Separate and diffuse: Using a pretrained diffusion model for better source separation,”
Shahar Lutati, Eliya Nachmani, and Lior Wolf, · 2024
Closest in time.
“U-mamba: Enhancing long-range dependency for biomedical image segmentation,”
Jun Ma, Feifei Li, and Bo Wang, · 2024
Closest in time.
“Graph-mamba: Towards long-range graph sequence modeling with selective state spaces,”
Chloe Wang, Oleksii Tsepa, Jun Ma, and Bo Wang, · 2024
Closest in time.
Zeyu Zhang, Akide Liu, Ian Reid, Richard Hartley, Bohan Zhuang, and Hao Tang, · 2024
Closest in time.
“Vivim: a video vision mamba for medical video object segmentation,”
Yijun Yang, Zhaohu Xing, and Lei Zhu, · 2024
Closest in time.
“Pointmamba: A simple state space model for point cloud analysis,”
Dingkang Liang, Xin Zhou, Xinyu Wang, Xingkui Zhu, Wei Xu, Zhikang Zou, Xiaoqing Ye, and Xiang Bai, · 2024
Closest in time.
“Multichannel long-term streaming neural speech enhancement for static and moving speakers,”
Changsheng Quan and Xiaofei Li, · 2024
Closest in time.