Fetching the paper…
Reading the bibliography…
Scaling the effective context length is essential for advancing large language models (LLMs) toward artificial general intelligence (AGI).
“Outrageously large neural networks: The sparsely-gated mixture-of-experts layer”
Noam Shazeer et al · 2017
Earlier work this paper cites.
“Attention is all you need”
A Waswani et al · 2017
Earlier work this paper cites.
“Online normalizer calculation for softmax”
Maxim Milakov and Natalia Gimelshein · 2018
Earlier work this paper cites.
“Generating long sequences with sparse transformers”
Rewon Child et al · 2019
Earlier work this paper cites.
Qipeng Guo et al · 2019
Earlier work this paper cites.
“Axial attention in multidimensional transformers”
Jonathan Ho et al · 2019
Earlier work this paper cites.
“Blockwise self-attention for long document understanding”
Jiezhong Qiu et al · 2019
Earlier work this paper cites.
“Megatron-lm: Training multi-billion parameter language models using model parallelism”
Mohammad Shoeybi et al · 2019
Earlier work this paper cites.
“ETC: Encoding long and structured inputs in transformers”
Joshua Ainslie et al · 2020
Earlier work this paper cites.
“Longformer: The long-document transformer”
Iz Beltagy, Matthew Peters and Arman Cohan · 2020
Earlier work this paper cites.
“Rethinking attention with performers”
Krzysztof Choromanski et al · 2020
Earlier work this paper cites.
“Gmat: Global memory augmentation for transformers”
Ankit Gupta and Jonathan Berant · 2020
Earlier work this paper cites.
“Transformers are rnns: Fast autoregressive transformers with linear attention”
Angelos Katharopoulos et al · 2020
Earlier work this paper cites.
“Reformer: The efficient transformer”
Nikita Kitaev, Łukasz Kaiser and Anselm Levskaya · 2020
Earlier work this paper cites.
“Gshard: Scaling giant models with conditional computation and automatic sharding”
Dmitry Lepikhin et al · 2020
Earlier work this paper cites.
“Sparse sinkhorn attention”
Yi Tay et al · 2020
Earlier work this paper cites.
“Efficient transformers: A survey. arXiv”
Yi Tay et al · 2020
Earlier work this paper cites.
“Linformer: Self-attention with linear complexity”
Sinong Wang et al · 2020
Earlier work this paper cites.
“Big bird: Transformers for longer sequences”
Manzil Zaheer et al · 2020
Earlier work this paper cites.
“LongT5: Efficient text-to-text transformer for long sequences”
Mandy Guo et al · 2021
Earlier work this paper cites.
“Efficient content-based sparse attention with routing transformers”
Aurko Roy et al · 2021
Earlier work this paper cites.
“Flashattention: Fast and memory-efficient exact attention with io-awareness”
Tri Dao et al · 2022
Earlier work this paper cites.
“Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity”
William Fedus, Barret Zoph and Noam Shazeer · 2022
Cited alongside, same era.
“Training compute-optimal large language models”
Jordan Hoffmann et al · 2022
Cited alongside, same era.
“Deepspeed-moe: Advancing mixture-of-experts inference and training to power next-generation ai scale”
Samyam Rajbhandari et al · 2022
Cited alongside, same era.
Yuhuai Wu et al · 2022
Cited alongside, same era.
“St-moe: Designing stable and transferable sparse expert models”
Barret Zoph et al · 2022
Cited alongside, same era.
“Moa: Mixture of sparse attention for automatic large language model compression”
Tianyu Fu et al · 2024
Later among the works it cites.
“SeerAttention: Learning Intrinsic Sparse Attention in Your LLMs”
Yizhao Gao et al · 2024
Later among the works it cites.
“Deliberative alignment: Reasoning enables safer language models”
Melody Guan et al · 2024
Later among the works it cites.
“Minference 1.0: Accelerating pre-filling for long-context llms via dynamic sparse attention”
Huiqiang Jiang et al · 2024
Later among the works it cites.
“Retrievalattention: Accelerating long-context llm inference via vector retrieval”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Colt5: Faster long-range transformers with conditional computation”
Joshua Ainslie et al · 2023
Cited alongside, same era.
“Introducing 100K Context Windows”, https://www.anthropic.com/news/100k-context-windows , 2023
Anthropic · 2023
Cited alongside, same era.
“Extending Context Window of Large Language Models via Positional Interpolation”
Shouyuan Chen et al · 2023
Cited alongside, same era.
“Longnet: Scaling transformers to 1,000,000,000 tokens”
Jiayu Ding et al · 2023
Cited alongside, same era.
“Model tells you what to discard: Adaptive kv cache compression for llms”
Suyu Ge et al · 2023
Cited alongside, same era.
“Mamba: Linear-time sequence modeling with selective state spaces”
Albert Gu and Tri Dao · 2023
Cited alongside, same era.
“Blockwise Parallel Transformer for Large Context Models”
Hao Liu and Pieter Abbeel · 2023
Cited alongside, same era.
Di Liu et al · 2024
Later among the works it cites.
“LongHeads: Multi-Head Attention is Secretly a Long Context Processor”
Yi Lu et al · 2024
Later among the works it cites.
“Linearizing Large Language Models”
Jean Mercat et al · 2024
Later among the works it cites.
“Transformers are multi-state rnns”
Matanel Oren et al · 2024
Later among the works it cites.
“Eagle and finch: Rwkv with matrix-valued states and dynamic recurrence”
Bo Peng et al · 2024
Later among the works it cites.
“Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context”
Machel Reid et al · 2024
Later among the works it cites.
“Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference”
Jiaming Tang et al · 2024
Later among the works it cites.
“The mamba in the llama: Distilling and accelerating hybrid models”
Junxiong Wang et al · 2024
Later among the works it cites.
An Yang et al · 2024
Later among the works it cites.
“LoLCATs: On Low-Rank Linearizing of Large Language Models”
Michael Zhang et al · 2024
Later among the works it cites.
“Simlayerkv: A simple framework for layer-level KV cache reduction”
Xuan Zhang et al · 2024
Later among the works it cites.
“H2o: Heavy-hitter oracle for efficient generative inference of large language models”
Zhenyu Zhang et al · 2024
Later among the works it cites.
“Open-Sora: Democratizing Efficient Video Production for All”, 2024
Zangwei Zheng et al · 2024
Later among the works it cites.
“Transformers to ssms: Distilling quadratic knowledge to subquadratic models”
Aviv Bick et al · 2025
Closest in time.
“Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning”
Daya Guo et al · 2025
Closest in time.
“Minimax-01: Scaling foundation models with lightning attention”
Aonian Li et al · 2025
Closest in time.
“Kimi k1. 5: Scaling Reinforcement Learning with LLMs”
Kimi Team et al · 2025
Closest in time.
“Human hippocampal CA3 uses specific functional connectivity rules for efficient associative memory”
Jake Watson et al · 2025
Closest in time.