Fetching the paper…
Reading the bibliography…
Hybrid models that combine the language modeling capabilities of Attention layers with the efficiency of Recurrent layers (e.g., State Space Models) have gained traction in practically supporting long contexts in Large Language Model serving.
Improving WWW proxies performance with greedy-dual-size-frequency caching policy
Cherkasova, L · 1998
Earlier work this paper cites.
The design and operation of cloudlab
Duplyakin, D., Ricci, R., Maricq, A., Wong, G., Duerig, J., Eide, E., Stoller, L., Hibler, M., Johnson, D., Webb, K., et al · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T. B · 2020
Earlier work this paper cites.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2020
Earlier work this paper cites.
https://openai.com/index/chatgpt/ , 2022
Introducing ChatGPT · 2022
Earlier work this paper cites.
https://github.com/features/copilot , 2022
Github Copilot · 2022
Earlier work this paper cites.
Flashattention: Fast and memory-efficient exact attention with io-awareness
Dao, T., Fu, D., Ermon, S., Rudra, A., and Ré, C · 2022
Earlier work this paper cites.
Self-consistency improves chain of thought reasoning in language models
Wang, X., Wei, J., Schuurmans, D., Le, Q., Chi, E., Narang, S., Chowdhery, A., and Zhou, D · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al · 2022
Earlier work this paper cites.
React: Synergizing reasoning and acting in language models
Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K., and Cao, Y · 2022
Earlier work this paper cites.
Orca: A distributed serving system for transformer-based generative models
Yu, G.-I., Jeong, J. S., Kim, G.-W., Kim, S., and Chun, B.-G · 2022
Earlier work this paper cites.
https://www.perplexity.ai/ , 2023
Perplexity · 2023
Earlier work this paper cites.
Flashattention-2: Faster attention with better parallelism and work partitioning
Dao, T · 2023
Earlier work this paper cites.
Mamba: Linear-time sequence modeling with selective state spaces (2023)
Gu, A. and Dao, T · 2023
Earlier work this paper cites.
Swe-bench: Can language models resolve real-world github issues?
Jimenez, C. E., Yang, J., Wettig, A., Yao, S., Pei, K., Press, O., and Narasimhan, K · 2023
Earlier work this paper cites.
Dspy: Compiling declarative language model calls into self-improving pipelines
Khattab, O., Singhvi, A., Maheshwari, P., Zhang, Z., Santhanam, K., Vardhamanan, S., Haq, S., Sharma, A., Joshi, T. T., Moazam, H., et al · 2023
Earlier work this paper cites.
Efficient memory management for large language model serving with pagedattention
Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., Gonzalez, J., Zhang, H., and Stoica, I · 2023
Earlier work this paper cites.
Loogle: Can long-context language models understand long contexts?
Li, J., Wang, M., Zheng, Z., and Zhang, M · 2023
Earlier work this paper cites.
Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing
Liu, P., Yuan, W., Fu, J., Jiang, Z., Hayashi, H., and Neubig, G · 2023
Earlier work this paper cites.
Rwkv: Reinventing rnns for the transformer era
Peng, B., Alcaide, E., Anthony, Q., Albalak, A., Arcadinho, S., Biderman, S., Cao, H., Cheng, X., Chung, M., Grella, M., et al · 2023
Cited alongside, same era.
Retentive network: A successor to transformer for large language models
Sun, Y., Dong, L., Huang, S., Ma, S., Xia, Y., Xue, J., Wang, J., and Wei, F · 2023
Cited alongside, same era.
Gemini: a family of highly capable multimodal models
Team, G., Anil, R., Borgeaud, S., Wu, Y., Alayrac, J.-B., Yu, J., Soricut, R., Schalkwyk, J., Dai, A. M., Hauth, A., et al · 2023
Cited alongside, same era.
Gated linear attention transformers with hardware-efficient training
Yang, S., Wang, B., Shen, Y., Panda, R., and Kim, Y · 2023
Cited alongside, same era.
Zamba: A compact 7b ssm hybrid model
Glorioso, P., Anthony, Q., Tokpanov, Y., Whittington, J., Pilault, J., Ibrahim, A., and Millidge, B · 2024
Closest in time.
Infinigen: Efficient generative inference of large language models with dynamic kv cache management
Lee, W., Lee, J., Seo, J., and Sim, J · 2024
Closest in time.
Jamba: A hybrid transformer-mamba language model
Lieber, O., Lenz, B., Bata, H., Cohen, G., Osin, J., Dalmedigos, I., Safahi, E., Meirom, S., Belinkov, Y., Shalev-Shwartz, S., et al · 2024
Closest in time.
Can mamba learn how to learn? a comparative study on in-context learning tasks
Park, J., Park, J., Xiong, Z., Lee, N., Cho, J., Oymak, S., Lee, K., and Papailiopoulos, D · 2024
Closest in time.
Mooncake: Kimi’s kvcache-centric architecture for llm serving
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yu, L. and Li, J · 2023
Cited alongside, same era.
https://cartesia.ai/blog/2024-08-27-on-device , 2024
Rene: An Open-Source 1.3B SSM Language Model · 2024
Cited alongside, same era.
https://research.character.ai/optimizing-inference/ , 2024
Optimizing AI Inference at Character.AI · 2024
Cited alongside, same era.
https://github.com/gkamradt/LLMTest_NeedleInAHaystack , 2024
Needle In A Haystack - Pressure Testing LLMs · 2024
Cited alongside, same era.
https://platform.openai.com/docs/guides/prompt-caching , 2024
Prompt caching - OpenAI API · 2024
Cited alongside, same era.
https://sharegpt.com/ , 2024
ShareGPT · 2024
Cited alongside, same era.
https://docs.vllm.ai/en/stable/models/engine_args.html , 2024
Engine Arguments – vLLM · 2024
Cited alongside, same era.
Taming throughput-latency tradeoff in llm inference with sarathi-serve
Agrawal, A., Kedia, N., Panwar, A., Mohan, J., Kwatra, N., Gulavani, B., Tumanov, A., and Ramjee, R · 2024
Cited alongside, same era.
Qin, R., Li, Z., He, W., Zhang, M., Wu, Y., Zheng, W., and Xu, X · 2024
Closest in time.
Samba: Simple hybrid state space models for efficient unlimited context language modeling
Ren, L., Liu, Y., Lu, Y., Shen, Y., Liang, C., and Chen, W · 2024
Closest in time.
Flashattention-3: Fast and accurate attention with asynchrony and low-precision
Shah, J., Bikshandi, G., Zhang, Y., Thakkar, V., Ramani, P., and Dao, T · 2024
Closest in time.
Preble: Efficient distributed prompt scheduling for llm serving
Srivatsa, V., He, Z., Abhyankar, R., Li, D., and Zhang, Y · 2024
Closest in time.
Learning to (learn at test time): Rnns with expressive hidden states
Sun, Y., Li, X., Dalal, K., Xu, J., Vikram, A., Zhang, G., Dubois, Y., Chen, X., Wang, X., Koyejo, S., et al · 2024
Closest in time.
Jamba-1.5: Hybrid transformer-mamba models at scale
Team, J., Lenz, B., Arazi, A., Bergman, A., Manevich, A., Peleg, B., Aviram, B., Almagor, C., Fridman, C., Padnos, D., et al · 2024
Closest in time.
An empirical study of mamba-based language models
Waleffe, R., Byeon, W., Riach, D., Norick, B., Korthikanti, V., Dao, T., Gu, A., Hatamizadeh, A., Singh, S., Narayanan, D., et al · 2024
Closest in time.
Opendevin: An open platform for ai software developers as generalist agents
Wang, X., Li, B., Song, Y., Xu, F. F., Tang, X., Zhuge, M., Pan, J., Song, Y., Li, B., Singh, J., et al · 2024
Closest in time.
Loongserve: Efficiently serving long-context large language models with elastic sequence parallelism
Wu, B., Liu, S., Zhong, Y., Sun, P., Liu, X., and Jin, X · 2024
Closest in time.
Cacheblend: Fast large language model serving with cached knowledge fusion
Yao, J., Li, H., Liu, Y., Ray, S., Cheng, Y., Zhang, Q., Du, K., Lu, S., and Jiang, J · 2024
Closest in time.
B’mojo: Hybrid state space realizations of foundation models with eidetic and fading memory
Zancato, L., Seshadri, A., Dukler, Y., Golatkar, A., Shen, Y., Bowman, B., Trager, M., Achille, A., and Soatto, S · 2024
Closest in time.
https://llm.hunyuan.tencent.com/#/blog/hy-t1?lang=en , 2025
Reasoning Efficiency Redefined! Meet Tencent’s ‘Hunyuan-T1’—The First Mamba-Powered Ultra-Large Model · 2025
Closest in time.
https://research.nvidia.com/labs/adlr/nemotronh/ , 2025
Nemotron-H: A Family of Accurate, Efficient Hybrid Mamba-Transformer Models · 2025
Closest in time.
Minimax-01: Scaling foundation models with lightning attention
Li, A., Gong, B., Yang, B., Shan, B., Liu, C., Zhu, C., Zhang, C., Guo, C., Chen, D., Li, D., et al · 2025
Closest in time.