Fetching the paper…
Reading the bibliography…
This paper explores the capability of Mamba, a recently proposed architecture based on state space models (SSMs), as a competitive alternative to Transformer-based models.
G. E. Blelloch, “Prefix sums and their applications,” 1990
1990
Earlier work this paper cites.
M. Schuster and K. K. Paliwal, “Bidirectional recurrent neural networks,” IEEE Trans. Signal Process. , vol. 45, no. 11, pp. 2673–2681, 1997
1997
Earlier work this paper cites.
K. Maekawa, H. Koiso et al. , “Spontaneous speech corpus of Japanese,” in Proc. LREC , 2000
2000
Earlier work this paper cites.
C.-Y. Lin, “ROUGE: A package for automatic evaluation of summaries,” in Proc. Summarization Branches Out , 2004, pp. 74–81
2004
Earlier work this paper cites.
S. Banerjee and A. Lavie, “METEOR: An automatic metric for MT evaluation with improved correlation with human judgments,” in Proc. ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization , 2005, pp. 65–72
2005
Earlier work this paper cites.
A. Rousseau, P. Deléglise et al. , “Enhancing the TED-LIUM corpus with selected data for language modeling and more TED talks.” in Proc. LREC , 2014, pp. 3935–3939
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
V. Panayotov, G. Chen et al. , “LibriSpeech: An ASR corpus based on public domain audio books,” in Proc. ICASSP , 2015
2015
Earlier work this paper cites.
J. L. Ba, J. R. Kiros et al. , “Layer normalization,” arXiv preprint arXiv:1607.06450 , 2016
2016
Earlier work this paper cites.
R. J. Weiss, J. Chorowski et al. , “Sequence-to-sequence models can directly translate foreign speech,” in Proc. Interspeech , 2017, pp. 2625–2629
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer et al. , “Attention is all you need,” in Proc. NeurIPS , 2017
2017
Earlier work this paper cites.
H. Bu, J. Du et al. , “AISHELL-1: An open-source Mandarin speech corpus and a speech recognition baseline,” in Proc. O-COCOSDA , 2017, pp. 1–5
2017
Earlier work this paper cites.
K. Ito and L. Johnson, “The LJ speech dataset,” https://keithito.com/LJ-Speech-Dataset/ , 2017
2017
Earlier work this paper cites.
M. McAuliffe, M. Socolof et al. , “Montreal forced aligner: Trainable text-speech alignment using Kaldi,” in Proc. Interspeech , 2017, pp. 498–502
2017
Earlier work this paper cites.
R. Sanabria, O. Caglayan et al. , “How2: a large-scale dataset for multimodal language understanding,” in Proc. ViGIL , 2018
2018
Earlier work this paper cites.
S. Karita, N. Chen et al. , “A comparative study on Transformer vs RNN in speech applications,” in Proc. ASRU , 2019, pp. 449–456
2019
Earlier work this paper cites.
I. Loshchilov and F. Hutter, “Decoupled weight decay regularization,” in Proc. ICLR , 2019
2019
Earlier work this paper cites.
T. Zhang, V. Kishore et al. , “BERTScore: Evaluating text generation with BERT,” in Proc. ICLR , 2019
2019
Cited alongside, same era.
A. Gulati, J. Qin et al. , “Conformer: Convolution-augmented transformer for speech recognition,” in Proc. Interspeech , 2020, pp. 5036–5040
2020
Cited alongside, same era.
Y. Ren, C. Hu et al. , “FastSpeech 2: Fast and high-quality end-to-end text to speech,” in Proc. ICLR , 2020
2020
Cited alongside, same era.
J. Kong, J. Kim et al. , “HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis,” Proc. NeurIPS , vol. 33, pp. 17 022–17 033, 2020
2020
Cited alongside, same era.
E. Bastianelli, A. Vanzo et al. , “SLURP: A spoken language understanding resource package,” in Proc. EMNLP , 2020
2020
Cited alongside, same era.
2023
Later among the works it cites.
Y. Peng, K. Kim et al. , “A Comparative Study on E-Branchformer vs Conformer in Speech Recognition, Translation, and Understanding Tasks,” in Proc. Interspeech , 2023, pp. 2208–2212
2023
Later among the works it cites.
2023
Later among the works it cites.
G. Saon, A. Gupta et al. , “Diagonal state space augmented transformers for speech recognition,” in Proc. ICASSP , 2023, pp. 1–5
2023
Later among the works it cites.
K. Miyazaki, M. Murata et al. , “Structured state space decoder for speech recognition and synthesis,” in Proc. ICASSP , 2023, pp. 1–5
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
G. Chen, S. Chai et al. , “GigaSpeech: An evolving, multi-domain ASR corpus with 10,000 hours of transcribed audio,” in Proc. Interspeech , 2021, pp. 4376–4380
2021
Cited alongside, same era.
2021
Cited alongside, same era.
K. Kim, F. Wu et al. , “E-Branchformer: Branchformer with enhanced merging for speech recognition,” in Proc. SLT , 2022, pp. 84–91
2022
Cited alongside, same era.
Y. Peng et al. , “Branchformer: Parallel MLP-attention architectures to capture local and global context for speech recognition and understanding,” in Proc. ICML , 2022
2022
Cited alongside, same era.
K. Goel, A. Gu et al. , “It’s raw! audio generation with state-space models,” in Proc. ICML , 2022, pp. 7616–7633
2022
Cited alongside, same era.
A. Gu, K. Goel et al. , “Efficiently modeling long sequences with structured state spaces,” in Proc. ICLR , 2022
2022
Cited alongside, same era.
A. Gu, K. Goel et al. , “On the parameterization and initialization of diagonal state space models,” Proc. NeurIPS , vol. 35, pp. 35 971–35 983, 2022
2022
Cited alongside, same era.
2023
Later among the works it cites.
Y. Fathullah, C. Wu et al. , “Multi-head state space model for speech recognition,” in Proc. Interspeech , 2023, pp. 241–245
2023
Later among the works it cites.
C. Chen, C.-H. H. Yang et al. , “A neural state-space modeling approach to efficient speech separation,” in Proc. Interspeech , 2023, pp. 3784–3788
2023
Later among the works it cites.
P.-J. Ku, C.-H. H. Yang et al. , “A multi-dimensional deep structured state space approach to speech enhancement using small-footprint models,” in Proc. Interspeech , 2023, pp. 2453–2457
2023
Later among the works it cites.
2023
Later among the works it cites.
J. T. Smith, A. Warrington et al. , “Simplified state space layers for sequence modeling,” in Proc. ICLR , 2023
2023
Later among the works it cites.
T. Kano, A. Ogawa et al. , “Speech summarization of long spoken document: Improving memory efficiency of speech/text encoders,” in Proc. ICASSP , 2023, pp. 1–5
2023
Later among the works it cites.
R. Sharma, W. Chen et al. , “Espnet-Summ: Introducing a novel large dataset, toolkit, and a cross-corpora evaluation of speech summarization systems,” in Proc. ASRU , 2023, pp. 1–8
2023
Later among the works it cites.
R. Prabhavalkar, T. Hori et al. , “End-to-end speech recognition: A survey,” IEEE/ACM Trans. Audio, Speech, Lang. Process. , vol. 32, pp. 325–351, 2024
2024
Closest in time.
Z. Yao, L. Guo et al. , “Zipformer: A faster and better encoder for automatic speech recognition,” in Proc. ICLR , 2024
2024
Closest in time.
H. Shan, A. Gu et al. , “Augmenting conformers with structured state-space sequence models for online speech recognition,” in Proc. ICASSP , 2024, pp. 12 221–12 225
2024
Closest in time.
2024
Closest in time.