Fetching the paper…
Reading the bibliography…
Long-context modeling is crucial for next-generation language models, yet the high computational cost of standard attention mechanisms poses significant computational challenges.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. u. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Drop: A reading comprehension benchmark requiring discrete reasoning over paragraphs
D. Dua, Y. Wang, P. Dasigi, G. Stanovsky, S. Singh, and M. Gardner · 2019
Earlier work this paper cites.
Fast transformer decoding: One write-head is all you need
N. Shazeer · 2019
Earlier work this paper cites.
Triton: an intermediate language and compiler for tiled neural network computations
P. Tillet, H.-T. Kung, and D. Cox · 2019
Earlier work this paper cites.
Longformer: The long-document transformer
I. Beltagy, M. E. Peters, and A. Cohan · 2020
Earlier work this paper cites.
Measuring massive multitask language understanding
D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt · 2020
Earlier work this paper cites.
Big bird: Transformers for longer sequences
M. Zaheer, G. Guruganesh, K. A. Dubey, J. Ainslie, C. Alberti, S. Ontanon, P. Pham, A. Ravula, Q. Wang, L. Yang, et al · 2020
Earlier work this paper cites.
Program synthesis with large language models
J. Austin, A. Odena, M. Nye, M. Bosma, H. Michalewski, D. Dohan, E. Jiang, C. Cai, M. Terry, Q. Le, et al · 2021
Earlier work this paper cites.
Evaluating large language models trained on code
M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. D. O. Pinto, J. Kaplan, H. Edwards, Y. Burda, N. Joseph, G. Brockman, et al · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems, 2021
K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, et al · 2021
Earlier work this paper cites.
Challenging big-bench tasks and whether chain-of-thought can solve them
M. Suzgun, N. Scales, N. Schärli, S. Gehrmann, Y. Tay, H. W. Chung, A. Chowdhery, Q. V. Le, E. H. Chi, D. Zhou, et al · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou, et al · 2022
Earlier work this paper cites.
Star: Bootstrapping reasoning with reasoning
E. Zelikman, Y. Wu, J. Mu, and N. D. Goodman · 2022
Earlier work this paper cites.
Gqa: Training generalized multi-query transformer models from multi-head checkpoints
J. Ainslie, J. Lee-Thorp, M. de Jong, Y. Zemlyanskiy, F. Lebrón, and S. Sanghai · 2023
Cited alongside, same era.
Longbench: A bilingual, multitask benchmark for long context understanding
Y. Bai, X. Lv, J. Zhang, H. Lyu, J. Tang, Z. Huang, Z. Du, X. Liu, A. Zeng, L. Hou, et al · 2023
Cited alongside, same era.
Model tells you what to discard: Adaptive kv cache compression for llms
S. Ge, Y. Zhang, L. Liu, M. Zhang, J. Han, and J. Gao · 2023
Cited alongside, same era.
Llmlingua: Compressing prompts for accelerated inference of large language models
H. Jiang, Q. Wu, C.-Y. Lin, Y. Yang, and L. Qiu · 2023
Cited alongside, same era.
LLMTest NeedleInAHaystack
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
G. T. Google, P. Georgiev, V. I. Lei, R. Burnell, L. Bai, A. Gulati, G. Tanzer, D. Vincent, Z. Pan, S. Wang, et al · 2024
Later among the works it cites.
Minference 1.0: Accelerating pre-filling for long-context llms via dynamic sparse attention
H. Jiang, Y. Li, C. Zhang, Q. Wu, X. Luo, S. Ahn, Z. Han, A. H. Abdi, D. Li, C.-Y. Lin, et al · 2024
Later among the works it cites.
Snapkv: Llm knows what you are looking for before generation
Y. Li, Y. Huang, B. Yang, B. Venkitesh, A. Locatelli, H. Ye, T. Cai, P. Lewis, and D. Chen · 2024
Later among the works it cites.
Clusterkv: Manipulating llm kv cache in semantic space for recallable compression
G. Liu, C. Li, J. Zhao, C. Zhang, and M. Guo · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
G. Kamradt · 2023
Cited alongside, same era.
Cmmlu: Measuring massive multitask language understanding in chinese
H. Li, Y. Zhang, F. Koto, Y. Yang, H. Zhao, Y. Gong, N. Duan, and T. Baldwin · 2023
Cited alongside, same era.
Generative agents: Interactive simulacra of human behavior
J. S. Park, J. C. O’Brien, C. J. Cai, M. R. Morris, P. Liang, and M. S. Bernstein · 2023
Cited alongside, same era.
Efficient streaming language models with attention sinks
G. Xiao, Y. Tian, B. Chen, S. Han, and M. Lewis · 2023
Cited alongside, same era.
Repocoder: Repository-level code completion through iterative retrieval and generation
F. Zhang, B. Chen, Y. Zhang, J. Keung, J. Liu, D. Zan, Y. Mao, J. Lou, and W. Chen · 2023
Cited alongside, same era.
Deepseekmoe: Towards ultimate expert specialization in mixture-of-experts language models
D. Dai, C. Deng, C. Zhao, R. Xu, H. Gao, D. Chen, J. Li, W. Zeng, X. Yu, Y. Wu, et al · 2024
Cited alongside, same era.
Deepseek-v2: A strong, economical, and efficient mixture-of-experts language model
DeepSeek-AI · 2024
Cited alongside, same era.
Hashattention: Semantic sparsity for faster inference
A. Desai, S. Yang, A. Cuadron, A. Klimovic, M. Zaharia, J. E. Gonzalez, and I. Stoica · 2024
Cited alongside, same era.
Yarn: Efficient context window extension of large language models
B. Peng, J. Quesnelle, H. Fan, and E. Shippole · 2024
Later among the works it cites.
Quest: Query-aware sparsity for efficient long-context llm inference
J. Tang, Y. Zhao, K. Zhu, G. Xiao, B. Kasikci, and S. Han · 2024
Later among the works it cites.
Mmlu-pro: A more robust and challenging multi-task language understanding benchmark
Y. Wang, X. Ma, G. Zhang, Y. Ni, A. Chandra, S. Guo, W. Ren, A. Arulraj, X. He, Z. Jiang, et al · 2024
Later among the works it cites.
W. Wu, Z. Pan, C. Wang, L. Chen, Y. Bai, K. Fu, Z. Wang, and H. Xiong · 2024
Later among the works it cites.
Codeagent: Enhancing code generation with tool-integrated agent systems for real-world repo-level coding challenges
K. Zhang, J. Li, G. Li, X. Shi, and Z. Jin · 2024
Later among the works it cites.
Buzz: Beehive-structured sparse kv cache with segmented heavy hitters for efficient llm inference
J. Zhao, Z. Fang, S. Li, S. Yang, and S. He · 2024
Later among the works it cites.
Llm × \times mapreduce: Simplified long-sequence processing using large language models
Z. Zhou, C. Li, X. Chen, S. Wang, Y. Chao, Z. Li, H. Wang, R. An, Q. Shi, Z. Tan, et al · 2024
Later among the works it cites.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025
DeepSeek-AI · 2025
Closest in time.