Fetching the paper…
Reading the bibliography…
Transformers with linear recurrent modeling offer linear-time training and constant-memory inference.
Intrinsic dimensionality explains the effectiveness of language model fine-tuning, 2020
Aghajanyan, A., Zettlemoyer, L., and Gupta, S · 2012
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge
Clark, P., Cowhey, I., Etzioni, O., Khot, T., Sabharwal, A., Schoenick, C., and Tafjord, O · 2018
Earlier work this paper cites.
Winogrande: An adversarial winograd schema challenge at scale
Sakaguchi, K., Bras, R. L., Bhagavatula, C., and Choi, Y · 2019
Earlier work this paper cites.
Hellaswag: Can a machine really finish your sentence?
Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y · 2019
Earlier work this paper cites.
Longformer: The long-document transformer, 2020
Beltagy, I., Peters, M. E., and Cohan, A · 2020
Earlier work this paper cites.
Piqa: Reasoning about physical commonsense in natural language
Bisk, Y., Zellers, R., Bras, R. L., Gao, J., and Choi, Y · 2020
Earlier work this paper cites.
Transformers are rnns: Fast autoregressive transformers with linear attention
Katharopoulos, A., Vyas, A., Pappas, N., and Fleuret, F · 2020
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models, 2021
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2021
Earlier work this paper cites.
Finetuning pretrained transformers into rnns, 2021
Kasai, J., Peng, H., Zhang, Y., Yogatama, D., Ilharco, G., Pappas, N., Mao, Y., Chen, W., and Smith, N. A · 2021
Earlier work this paper cites.
Gpt-4 moe, June 2023
Chintala, S · 2023
Earlier work this paper cites.
Mamba: Linear-time sequence modeling with selective state spaces
Gu, A. and Dao, T · 2023
Earlier work this paper cites.
Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., de las Casas, D., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., Lavaud, L. R., Lachaux, M.-A., Stock, P., Scao, T. L., Lavril, T., Wang, T., Lacroix, T., and Sayed, W. E · 2023
Earlier work this paper cites.
Efficient memory management for large language model serving with pagedattention
Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., Gonzalez, J., Zhang, H., and Stoica, I · 2023
Earlier work this paper cites.
Cmmlu: Measuring massive multitask language understanding in chinese, 2023
Li, H., Zhang, Y., Koto, F., Yang, Y., Zhao, H., Gong, Y., Duan, N., and Baldwin, T · 2023
Earlier work this paper cites.
RWKV: Reinventing RNNs for the transformer era
Peng, B., Alcaide, E., Anthony, Q., Albalak, A., Arcadinho, S., Biderman, S., Cao, H., Cheng, X., Chung, M., Derczynski, L., Du, X., Grella, M., Gv, K., He, X., Hou, H., Kazienko, P., Kocon, J., Kong, J., Koptyra, B., Lau, H., Lin, J., Mantri, K. S. I., Mom, F., Saito, A., Song, G., Tang, X., Wind, J., Woźniak, S., Zhang, Z., Zhou, Q., Zhu, J., and Zhu, R.-J · 2023
Cited alongside, same era.
Stripedhyena: Moving beyond transformers with hybrid signal processing models
Poli, M., Wang, J., Massaroli, S., Quesnelle, J., Carlow, R., Nguyen, E., and Thomas, A · 2023
Cited alongside, same era.
Transnormerllm: A faster and better large language model with improved transnormer
Qin, Z., Li, D., Sun, W., Sun, W., Shen, X., Han, X., Wei, Y., Lv, B., Luo, X., Qiao, Y., et al · 2023
Cited alongside, same era.
Retentive network: A successor to transformer for large language models
Sun, Y., Dong, L., Huang, S., Ma, S., Xia, Y., Xue, J., Wang, J., and Wei, F · 2023
Cited alongside, same era.
Stanford alpaca: An instruction-following llama model
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al · 2024
Later among the works it cites.
A framework for few-shot language model evaluation, 07 2024
Gao, L., Tow, J., Abbasi, B., Biderman, S., Black, S., DiPofi, A., Foster, C., Golding, L., Hsu, J., Le Noac’h, A., Li, H., McDonell, K., Muennighoff, N., Ociepa, C., Phang, J., Reynolds, L., Schoelkopf, H., Skowron, A., Sutawika, L., Tang, E., Thite, A., Wang, B., Wang, K., and Zou, A · 2024
Later among the works it cites.
Zamba: A compact 7b ssm hybrid model, 2024
Glorioso, P., Anthony, Q., Tokpanov, Y., Whittington, J., Pilault, J., Ibrahim, A., and Millidge, B · 2024
Later among the works it cites.
Agent attention: On the integration of softmax and linear attention, 2024
Han, D., Ye, T., Han, Y., Xia, Z., Pan, S., Wan, P., Song, S., and Huang, G · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Taori, R., Gulrajani, I., Zhang, T., Dubois, Y., Li, X., Guestrin, C., Liang, P., and Hashimoto, T. B · 2023
Cited alongside, same era.
Internlm: A multilingual language model with progressively enhanced capabilities, 2023
Team, I · 2023
Cited alongside, same era.
Llama: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., Rodriguez, A., Joulin, A., Grave, E., and Lample, G · 2023
Cited alongside, same era.
Gated linear attention transformers with hardware-efficient training
Yang, S., Wang, B., Shen, Y., Panda, R., and Kim, Y · 2023
Cited alongside, same era.
Simple linear attention language models balance the recall-throughput tradeoff
Arora, S., Eyuboglu, S., Zhang, M., Timalsina, A., Alberti, S., Dylan Zinsley, J. Z., Rudra, A., and Ré, C · 2024
Cited alongside, same era.
xlstm: Extended long short-term memory, 2024
Beck, M., Pöppel, K., Spanring, M., Auer, A., Prudnikova, O., Kopp, M., Klambauer, G., Brandstetter, J., and Hochreiter, S · 2024
Cited alongside, same era.
Titans: Learning to memorize at test time
Behrouz, A., Zhong, P., and Mirrokni, V · 2024
Cited alongside, same era.
Transformers to ssms: Distilling quadratic knowledge to subquadratic models, 2024
Bick, A., Li, K. Y., Xing, E. P., Kolter, J. Z., and Gu, A · 2024
Cited alongside, same era.
Lieber, O., Lenz, B., Bata, H., Cohen, G., Osin, J., Dalmedigos, I., Safahi, E., Meirom, S., Belinkov, Y., Shalev-Shwartz, S., et al · 2024
Later among the works it cites.
Linearizing large language models, 2024
Mercat, J., Vasiljevic, I., Keh, S., Arora, K., Dave, A., Gaidon, A., and Kollar, T · 2024
Later among the works it cites.
Eagle and finch: Rwkv with matrix-valued states and dynamic recurrence
Peng, B., Goldstein, D., Anthony, Q., Albalak, A., Alcaide, E., Biderman, S., Cheah, E., Ferdinan, T., Hou, H., Kazienko, P., et al · 2024
Later among the works it cites.
Llama-moe v2: Exploring sparsity of llama from perspective of mixture-of-experts with post-training
Qu, X., Dong, D., Hu, X., Zhu, T., Sun, W., and Cheng, Y · 2024
Later among the works it cites.
The mamba in the llama: Distilling and accelerating hybrid models, 2024
Wang, J., Paliotta, D., May, A., Rush, A. M., and Dao, T · 2024
Later among the works it cites.
Language models are super mario: Absorbing abilities from homologous models as a free lunch
Yu, L., Yu, B., Yu, H., Huang, F., and Li, Y · 2024
Later among the works it cites.
Llama-moe: Building mixture-of-experts from llama with continual pre-training
Zhu, T., Qu, X., Dong, D., Ruan, J., Tong, J., He, C., and Cheng, Y · 2024
Later among the works it cites.
Mom: Linear sequence modeling with mixture-of-memories
Du, J., Sun, W., Lan, D., Hu, J., and Cheng, Y · 2025
Closest in time.
Gershman, S. J., Fiete, I., and Irie, K · 2025
Closest in time.
Minimax-01: Scaling foundation models with lightning attention, 2025
MiniMax, Li, A., Gong, B., Yang, B., Shan, B., Liu, C., Zhu, C., Zhang, C., Guo, C., Chen, D., Li, D., Jiao, E., Li, G., Zhang, G., Sun, H., Dong, H., Zhu, J., Zhuang, J., Song, J., Zhu, J., Han, J., Li, J., Xie, J., Xu, J., Yan, J., Zhang, K., Xiao, K., Kang, K., Han, L., Wang, L., Yu, L., Feng, L., Zheng, L., Chai, L., Xing, L., Ju, M., Chi, M., Zhang, M., Huang, P., Niu, P., Li, P., Zhao, P., Yang, Q., Xu, Q., Wang, Q., Wang, Q., Li, Q., Leng, R., Shi, S., Yu, S., Li, S., Zhu, S., Huang, T., Liang, T., Sun, W., Sun, W., Cheng, W., Li, W., Song, X., Su, X., Han, X., Zhang, X., Hou, X., Min, X., Zou, X., Shen, X., Gong, Y., Zhu, Y., Zhou, Y., Zhong, Y., Hu, Y., Fan, Y., Yu, Y., Yang, Y., Li, Y., Huang, Y., Li, Y., Huang, Y., Xu, Y., Mao, Y., Li, Z., Li, Z., Tao, Z., Ying, Z., Cong, Z., Qin, Z., Fan, Z., Yu, Z., Jiang, Z., and Wu, Z · 2025
Closest in time.
Lasp-2: Rethinking sequence parallelism for linear attention and its hybrid
Sun, W., Lan, D., Zhong, Y., Qu, X., and Cheng, Y · 2025
Closest in time.