Fetching the paper…
Reading the bibliography…
We propose Samba ASR,the first state of the art Automatic Speech Recognition(ASR)model leveraging the novel Mamba architecture as both encoder and decoder,built on the foundation of state space models(SSMs).Unlike transformerbased ASR models,which rely on self-attention mechanisms to capture dependencies,Samba ASR effectively models both local and global temporal dependencies using efficient statespace dynamics,achieving remarkable performance gains.By addressing the limitations of transformers,such as quadratic scaling with input length and difficulty in handling longrange dependencies,Samba ASR achieves superior accuracy and efficiency.Experimental results demonstrate that Samba ASR surpasses existing opensource transformerbased ASR models across various standard benchmarks,establishing it as the new state of theart in ASR.Extensive evaluations on the benchmark dataset show significant improvements in Word Error Rate(WER),with competitive performance even in lowresource scenarios.Furthermore,the inherent computational efficiency and parameter optimization of the Mamba architecture make Samba ASR a scalable and robust solution for diverse ASR tasks.Our contributions include the development of a new Samba ASR architecture for automatic speech recognition(ASR),demonstrating the superiority of structured statespace models(SSMs)over transformer based models for speech sequence processing.We provide a comprehensive evaluation on public benchmarks,showcasing stateoftheart(SOTA)performance,and present an indepth analysis of computational efficiency,robustness to noise,and sequence generalization.This work highlights the viability of Mamba SSMs as a transformerfree alternative for efficient and accurate ASR.By leveraging the advancements of statespace modeling,Samba ASR redefines ASR performance standards and sets a new benchmark for future research in this field.
Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber · 2006
Earlier work this paper cites.
Speech recognition with deep recurrent neural networks, 2013
Alex Graves, Abdel rahman Mohamed, and Geoffrey Hinton · 2013
Earlier work this paper cites.
Librispeech: An asr corpus based on public domain audio books
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur · 2015
Earlier work this paper cites.
Neural machine translation of rare words with subword units, 2016
Rico Sennrich, Barry Haddow, and Alexandra Birch · 2016
Earlier work this paper cites.
Dnn-based source enhancement to increase objective sound quality assessment score
Yuma Koizumi, Kenta Niwa, Yusuke Hioka, Kazunori Kobayashi, and Yoichi Haneda · 2018
Earlier work this paper cites.
Decoupled weight decay regularization, 2019
Ilya Loshchilov and Frank Hutter · 2019
Earlier work this paper cites.
wav2vec 2.0: A framework for self-supervised learning of speech representations, 2020
Alexei Baevski, Henry Zhou, Abdelrahman Mohamed, and Michael Auli · 2020
Earlier work this paper cites.
Conformer: Convolution-augmented transformer for speech recognition, 2020
Anmol Gulati, James Qin, Chung-Cheng Chiu, Niki Parmar, Yu Zhang, Jiahui Yu, Wei Han, Shibo Wang, Zhengdong Zhang, Yonghui Wu, and Ruoming Pang · 2020
Earlier work this paper cites.
Massively multilingual asr: 50 languages, 1 model, 1 billion parameters, 2020
Vineel Pratap, Anuroop Sriram, Paden Tomasello, Awni Hannun, Vitaliy Liptchinsky, Gabriel Synnaeve, and Ronan Collobert · 2020
Earlier work this paper cites.
Gigaspeech: An evolving, multi-domain asr corpus with 10,000 hours of transcribed audio, 2021
Guoguo Chen, Shuzhou Chai, Guanbo Wang, Jiayu Du, Wei-Qiang Zhang, Chao Weng, Dan Su, Daniel Povey, Jan Trmal, Junbo Zhang, Mingjie Jin, Sanjeev Khudanpur, Shinji Watanabe, Shuaijiang Zhao, Wei Zou, Xiangang Li, Xuchen Yao, Yongqing Wang, Yujun Wang, Zhao You, and Zhiyong Yan · 2021
Earlier work this paper cites.
Spgispeech: 5,000 hours of transcribed financial audio for fully formatted end-to-end speech recognition, 2021
Patrick K. O’Neill, Vitaly Lavrukhin, Somshubra Majumdar, Vahid Noroozi, Yuekai Zhang, Oleksii Kuchaiev, Jagadeesh Balam, Yuliya Dovzhenko, Keenan Freyberg, Michael D. Shulman, Boris Ginsburg, Shinji Watanabe, and Georg Kucsko · 2021
Cited alongside, same era.
Training data-efficient image transformers & distillation through attention, 2021
Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Hervé Jégou · 2021
Cited alongside, same era.
Robust speech recognition via large-scale weak supervision, 2022
Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever · 2022
Cited alongside, same era.
Efficiently modeling long sequences with structured state spaces, 2022
Albert Gu, Karan Goel, and Christopher Ré · 2022
Cited alongside, same era.
istftnet: Fast and lightweight mel-spectrogram vocoder incorporating inverse short-time fourier transform, 2022
Takuhiro Kaneko, Kou Tanaka, Hirokazu Kameoka, and Shogo Seki · 2022
Zamba: A compact 7b ssm hybrid model, 2024
Paolo Glorioso, Quentin Anthony, Yury Tokpanov, James Whittington, Jonathan Pilault, Adam Ibrahim, and Beren Millidge · 2024
Later among the works it cites.
Mamba: Linear-time sequence modeling with selective state spaces, 2024
Albert Gu and Tri Dao · 2024
Later among the works it cites.
Jamba: A hybrid transformer-mamba language model, 2024
Opher Lieber, Barak Lenz, Hofit Bata, Gal Cohen, Jhonathan Osin, Itay Dalmedigos, Erez Safahi, Shaked Meirom, Yonatan Belinkov, Shai Shalev-Shwartz, Omri Abend, Raz Alon, Tomer Asida, Amir Bergman, Roman Glozman, Michael Gokhman, Avashalom Manevich, Nir Ratner, Noam Rozen, Erez Shwartz, Mor Zusman, and Yoav Shoham · 2024
Later among the works it cites.
Vision mamba: Efficient visual representation learning with bidirectional state space model, 2024
Lianghui Zhu, Bencheng Liao, Qian Zhang, Xinlong Wang, Wenyu Liu, and Xinggang Wang · 2024
Later among the works it cites.
Less is more: Accurate speech recognition & translation without web-scale data, 2024
Krishna C. Puvvada, Piotr Żelasko, He Huang, Oleksii Hrinchuk, Nithin Rao Koluguri, Kunal Dhawan, Somshubra Majumdar, Elena Rastorgueva, Zhehuai Chen, Vitaly Lavrukhin, Jagadeesh Balam, and Boris Ginsburg · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Fastspeech 2: Fast and high-quality end-to-end text to speech, 2022
Yi Ren, Chenxu Hu, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu · 2022
Cited alongside, same era.
Attention is all you need, 2023
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2023
Cited alongside, same era.
Mac: A unified framework boosting low resource automatic speech recognition, 2023
Zeping Min, Qian Ge, Zhong Li, and Weinan E · 2023
Cited alongside, same era.
Falcon mamba: The first competitive attention-free 7b language model, 2024
Jingwei Zuo, Maksim Velikanov, Dhia Eddine Rhaiem, Ilyas Chahed, Younes Belkada, Guillaume Kunsch, and Hakim Hacid · 2024
Cited alongside, same era.
Later among the works it cites.
Mamba in speech: Towards an alternative to self-attention, 2024
Xiangyu Zhang, Qiquan Zhang, Hexin Liu, Tianyi Xiao, Xinyuan Qian, Beena Ahmed, Eliathamby Ambikairajah, Haizhou Li, and Julien Epps · 2024
Later among the works it cites.
Exploring the capability of mamba in speech applications, 2024
Koichi Miyazaki, Yoshiki Masuyama, and Masato Murata · 2024
Later among the works it cites.
Speech slytherin: Examining the performance and efficiency of mamba for speech separation, recognition, and synthesis, 2024
Xilin Jiang, Yinghao Aaron Li, Adrian Nicolas Florea, Cong Han, and Nima Mesgarani · 2024
Later among the works it cites.
Efficient whisper on streaming speech, 2024
Rongxiang Wang, Zhiming Xu, and Felix Xiaozhu Lin · 2024
Later among the works it cites.