Fetching the paper…
Reading the bibliography…
Deep State Space Models (SSMs), such as Mamba (Gu & Dao, 2024), have become powerful tools for language modeling, offering high performance and linear scalability with sequence length.
Learning multiple layers of features from tiny images
Krizhevsky, A., Hinton, G., et al · 2009
Earlier work this paper cites.
Deep learning face attributes in the wild
Liu, Z., Luo, P., Wang, X., and Tang, X · 2015
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Spider: A large-scale human-labeled dataset for complex and cross-domain semantic parsing and text-to-sql task
Yu, T., Zhang, R., Yang, K., Yasunaga, M., Wang, D., Li, Z., Ma, J., Li, I., Yao, Q., Roman, S., et al · 2018
Earlier work this paper cites.
SAMSum corpus: A human-annotated dialogue dataset for abstractive summarization
Gliwa, B., Mochol, I., Biesek, M., and Wawer, A · 2019
Earlier work this paper cites.
Parameter-efficient transfer learning for NLP
Houlsby, N., Giurgiu, A., Jastrzebski, S., Morrone, B., De Laroussilhe, Q., Gesmundo, A., Attariyan, M., and Gelly, S · 2019
Earlier work this paper cites.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Hippo: Recurrent memory with optimal polynomial projections
Gu, A., Dao, T., Ermon, S., Rudra, A., and Ré, C · 2020
Earlier work this paper cites.
Transformers are RNNs: Fast autoregressive transformers with linear attention
Katharopoulos, A., Vyas, A., Pappas, N., and Fleuret, F · 2020
Earlier work this paper cites.
Combining recurrent, convolutional, and continuous-time models with linear state space layers
Gu, A., Johnson, I., Goel, K., Saab, K., Dao, T., Rudra, A., and Ré, C · 2021
Earlier work this paper cites.
Towards a unified view of parameter-efficient transfer learning
He, J., Zhou, C., Ma, X., Berg-Kirkpatrick, T., and Neubig, G · 2021
Earlier work this paper cites.
LoRA: Low-rank adaptation of large language models
Hu, E. J., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., Chen, W., et al · 2021
Earlier work this paper cites.
The power of scale for parameter-efficient prompt tuning
Lester, B., Al-Rfou, R., and Constant, N · 2021
Earlier work this paper cites.
Prefix-Tuning: Optimizing Continuous Prompts for Generation
Li, X. L. and Liang, P · 2021
Earlier work this paper cites.
Liu, X., Zheng, Y., Du, Z., Ding, M., Qian, Y., Yang, Z., and Tang, J · 2021
Cited alongside, same era.
DART: Open-Domain Structured Data Record to Text Generation
Nan, L., Radev, D., Zhang, R., Rau, A., Sivaprasad, A., Hsieh, C., Tang, X., Vyas, A., Verma, N., Krishna, P., et al · 2021
Cited alongside, same era.
PICARD: Parsing incrementally for constrained auto-regressive decoding from language models
Scholak, T., Schucher, N., and Bahdanau, D · 2021
Cited alongside, same era.
Pufferfish: Communication-efficient models at no extra cost
Wang, H., Agarwal, S., and Papailiopoulos, D · 2021
Cited alongside, same era.
LIFT: Language-interfaced fine-tuning for non-language machine learning tasks
Dinh, T., Zeng, Y., Zhang, R., Lin, Z., Gira, M., Rajput, S., yong Sohn, J., Papailiopoulos, D., and Lee, K · 2022
Cited alongside, same era.
Retentive network: A successor to transformer for large language models
Sun, Y., Dong, L., Huang, S., Ma, S., Xia, Y., Xue, J., Wang, J., and Wei, F · 2023
Later among the works it cites.
Transformers are SSMs: Generalized models and efficient algorithms through structured state space duality
Dao, T. and Gu, A · 2024
Closest in time.
Mamba: Linear-time sequence modeling with selective state spaces
Gu, A. and Dao, T · 2024
Closest in time.
Mamba state-space models can be strong downstream learners
Halloran, J. T., Gulati, M., and Roysdon, P. F · 2024
Closest in time.
Lora+ efficient low rank adaptation of large models
Hayou, S., Ghosh, N., and Yu, B · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hungry hungry hippos: Towards language modeling with state space models
Fu, D. Y., Dao, T., Saab, K. K., Thomas, A. W., Rudra, A., and Re, C · 2022
Cited alongside, same era.
Diagonal state spaces are as effective as structured state spaces
Gupta, A., Gu, A., and Berant, J · 2022
Cited alongside, same era.
P-Tuning: Prompt Tuning Can Be Comparable to Fine-tuning Across Scales and Tasks
Liu, X., Ji, K., Fu, Y., Tam, W., Du, Z., Yang, Z., and Tang, J · 2022
Cited alongside, same era.
BitFit: Simple parameter-efficient fine-tuning for transformer-based masked language-models
Zaken, E. B., Goldberg, Y., and Ravfogel, S · 2022
Cited alongside, same era.
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al · 2023
Cited alongside, same era.
The expressive power of tuning only the normalization layers
Giannou, A., Rajput, S., and Papailiopoulos, D · 2023
Cited alongside, same era.
LLM-adapters: An adapter family for parameter-efficient fine-tuning of large language models
Hu, Z., Wang, L., Lan, Y., Xu, W., Lim, E.-P., Bing, L., Xu, X., Poria, S., and Lee, R · 2023
Cited alongside, same era.
Jang, U., Lee, J. D., and Ryu, E. K · 2024
Closest in time.
Dora: Weight-decomposed low-rank adaptation
Liu, S.-Y., Wang, C.-Y., Yin, H., Molchanov, P., Wang, Y.-C. F., Cheng, K.-T., and Chen, M.-H · 2024
Closest in time.
Can Mamba learn how to learn? a comparative study on in-context learning tasks
Park, J., Park, J., Xiong, Z., Lee, N., Cho, J., Oymak, S., Lee, K., and Papailiopoulos, D · 2024
Closest in time.
When do prompting and prefix-tuning work? a theory of capabilities and limitations
Petrov, A., Torr, P. H., and Bibi, A · 2024
Closest in time.
Sparse is enough in fine-tuning pre-trained large language models
Song, W., Li, Z., Zhang, L., Zhao, H., and Du, B · 2024
Closest in time.
The expressive power of low-rank adaptation
Zeng, Y. and Lee, K · 2024
Closest in time.
State-offset tuning: State-based parameter-efficient fine-tuning for state space models
Kang, W., Galim, K., Zeng, Y., Lee, M., Koo, H. I., and Cho, N. I · 2025
Closest in time.
Jamba: Hybrid transformer-mamba language models
Lieber, O., Lenz, B., Bata, H., Cohen, G., Osin, J., Dalmedigos, I., Safahi, E., Meirom, S., Belinkov, Y., Shalev-Shwartz, S., et al · 2025
Closest in time.
MambaPEFT: Exploring parameter-efficient fine-tuning for mamba
Yoshimura, M., Hayashi, T., and Maeda, Y · 2025
Closest in time.