Fetching the paper…
Reading the bibliography…
It is too early to conclude that Mamba is a better alternative to transformers for speech before comparing Mamba with transformers in terms of both performance and efficiency in multiple speech-related tasks.
“Long short-term memory,”
Sepp Hochreiter and Jürgen Schmidhuber, · 1997
Earlier work this paper cites.
“Librispeech: An asr corpus based on public domain audio books,”
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“Deep clustering: Discriminative embeddings for segmentation and separation,”
John R Hershey, Zhuo Chen, Jonathan Le Roux, and Shinji Watanabe, · 2016
Earlier work this paper cites.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin, · 2017
Earlier work this paper cites.
“Conv-tasnet: Surpassing ideal time–frequency magnitude masking for speech separation,”
Yi Luo and Nima Mesgarani, · 2019
Earlier work this paper cites.
“Libritts: A corpus derived from librispeech for text-to-speech,”
Heiga Zen, Viet Dang, Robert A. J. Clark, Yu Zhang, Ron J. Weiss, Ye Jia, Z. Chen, and Yonghui Wu, · 2019
Earlier work this paper cites.
“Sdr–half-baked or well done?,”
Jonathan Le Roux, Scott Wisdom, Hakan Erdogan, and John R Hershey, · 2019
Earlier work this paper cites.
“Dual-path rnn: efficient long sequence modeling for time-domain single-channel speech separation,”
Yi Luo, Zhuo Chen, and Takuya Yoshioka, · 2020
Earlier work this paper cites.
“Dual-Path Transformer Network: Direct Context-Aware Modeling for End-to-End Monaural Speech Separation,”
Jingjing Chen, Qirong Mao, and Dong Liu, · 2020
Earlier work this paper cites.
“Conformer: Convolution-augmented Transformer for Speech Recognition,”
Anmol Gulati, James Qin, Chung-Cheng Chiu, Niki Parmar, Yu Zhang, Jiahui Yu, Wei Han, Shibo Wang, Zhengdong Zhang, Yonghui Wu, and Ruoming Pang, · 2020
Earlier work this paper cites.
“Attention is all you need in speech separation,”
Cem Subakan, Mirco Ravanelli, Samuele Cornell, Mirko Bronzi, and Jianyuan Zhong, · 2021
Earlier work this paper cites.
“High fidelity neural audio compression,”
Alexandre Défossez, Jade Copet, Gabriel Synnaeve, and Yossi Adi, · 2022
Earlier work this paper cites.
“QDPN - Quasi-dual-path Network for single-channel Speech Separation,”
Joel Rixen and Matthias Renz, · 2022
Cited alongside, same era.
“Mamba: Linear-time sequence modeling with selective state spaces,”
Albert Gu and Tri Dao, · 2023
Cited alongside, same era.
“Mossformer: Pushing the performance limit of monaural speech separation using gated single-head transformer with convolution-augmented joint self-attentions,”
Shengkui Zhao and Bin Ma, · 2023
Cited alongside, same era.
“Neural codec language models are zero-shot text to speech synthesizers,”
Chengyi Wang, Sanyuan Chen, Yu Wu, Ziqiang Zhang, Long Zhou, Shujie Liu, Zhuo Chen, Yanqing Liu, Huaming Wang, Jinyu Li, et al., · 2023
Cited alongside, same era.
“Vision mamba: Efficient visual representation learning with bidirectional state space model,”
“U-mamba: Enhancing long-range dependency for biomedical image segmentation,”
Jun Ma, Feifei Li, and Bo Wang, · 2024
Closest in time.
“An investigation of incorporating mamba for speech enhancement,”
Rong Chao, Wen-Huang Cheng, Moreno La Quatra, Sabato Marco Siniscalchi, Chao-Han Huck Yang, Szu-Wei Fu, and Yu Tsao, · 2024
Closest in time.
Yueyuan Sui, Minghui Zhao, Junxi Xia, Xiaofan Jiang, and Stephen Xia, · 2024
Closest in time.
“Spmamba: State-space model is all you need in speech separation,”
Kai Li and Guo Chen, · 2024
Closest in time.
“Rawbmamba: End-to-end bidirectional state space model for audio deepfake detection,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lianghui Zhu, Bencheng Liao, Qian Zhang, Xinlong Wang, Wenyu Liu, and Xinggang Wang, · 2024
Cited alongside, same era.
“Multichannel long-term streaming neural speech enhancement for static and moving speakers,”
Changsheng Quan and Xiaofei Li, · 2024
Cited alongside, same era.
Xilin Jiang, Cong Han, and Nima Mesgarani, · 2024
Cited alongside, same era.
“Audio mamba: Bidirectional state space model for audio representation learning,”
Mehmet Hamza Erol, Arda Senocak, Jiu Feng, and Joon Son Chung, · 2024
Cited alongside, same era.
“An empirical study of mamba-based language models,” 2024
Roger Waleffe, Wonmin Byeon, Duncan Riach, Brandon Norick, Vijay Korthikanti, Tri Dao, Albert Gu, Ali Hatamizadeh, Sudhakar Singh, Deepak Narayanan, Garvit Kulshreshtha, Vartika Singh, Jared Casper, Jan Kautz, Mohammad Shoeybi, and Bryan Catanzaro, · 2024
Cited alongside, same era.
“A survey on vision mamba: Models, applications and challenges,”
Rui Xu, Shu Yang, Yihui Wang, Bo Du, and Hao Chen, · 2024
Cited alongside, same era.
“Vivim: a video vision mamba for medical video object segmentation,”
Yijun Yang, Zhaohu Xing, and Lei Zhu, · 2024
Cited alongside, same era.
“Video mamba suite: State space model as a versatile alternative for video understanding,”
Guo Chen, Yifei Huang, Jilan Xu, Baoqi Pei, Zhe Chen, Zhiqi Li, Jiahao Wang, Kunchang Li, Tong Lu, and Limin Wang, · 2024
Cited alongside, same era.
Yujie Chen, Jiangyan Yi, Jun Xue, Chenglong Wang, Xiaohui Zhang, Shunbo Dong, Siding Zeng, Jianhua Tao, Lv Zhao, and Cunhang Fan, · 2024
Closest in time.
“Audio mamba: Pretrained audio state space model for audio tagging,”
Jiaju Lin and Haoxuan Hu, · 2024
Closest in time.
“Audio mamba: Selective state spaces for self-supervised audio representations,”
Sarthak Yadav and Zheng-Hua Tan, · 2024
Closest in time.
“Ssamba: Self-supervised audio representation learning with mamba state space model,”
Siavash Shams, Sukru Samet Dindar, Xilin Jiang, and Nima Mesgarani, · 2024
Closest in time.
“Mamba in speech: Towards an alternative to self-attention,”
Xiangyu Zhang, Qiquan Zhang, Hexin Liu, Tianyi Xiao, Xinyuan Qian, Beena Ahmed, Eliathamby Ambikairajah, Haizhou Li, and Julien Epps, · 2024
Closest in time.
“Exploring the capability of mamba in speech applications,”
Koichi Miyazaki, Yoshiki Masuyama, and Masato Murata, · 2024
Closest in time.
“Jamba: A hybrid transformer-mamba language model,”
Opher Lieber, Barak Lenz, Hofit Bata, Gal Cohen, Jhonathan Osin, Itay Dalmedigos, Erez Safahi, Shaked Meirom, Yonatan Belinkov, Shai Shalev-Shwartz, et al., · 2024
Closest in time.