Fetching the paper…
Reading the bibliography…
Although large language models (LLMs) have achieved significant success in natural language processing, they still struggle with long-context comprehension.
Compressive transformers for long-range sequence modelling
Rae, J. W., Potapenko, A., Jayakumar, S. M., Hillier, C., and Lillicrap, T. P · 1911
Earlier work this paper cites.
Some studies in machine learning using the game of checkers
Samuel, A. L · 1959
Earlier work this paper cites.
The need for biases in learning generalizations
Mitchell, T. M · 1980
Earlier work this paper cites.
Mechanisms of skill acquisition and the law of practice
Newell, A. and Rosenbloom, P. S · 1981
Earlier work this paper cites.
Machine Learning: An Artificial Intelligence Approach, Vol. I
Michalski, R. S., Carbonell, J. G., and Mitchell, T. M. (eds.) · 1983
Earlier work this paper cites.
Computational Complexity of Machine Learning
Kearns, M. J · 1989
Earlier work this paper cites.
Pattern Classification
Duda, R. O., Hart, P. E., and Stork, D. G · 2000
Earlier work this paper cites.
Crafting papers on machine learning
Langley, P · 2000
Earlier work this paper cites.
Learning question classifiers
Li, X. and Roth, D · 2002
Earlier work this paper cites.
Task Definition for Large Scale Text Categorization at NLPCC 2014, 2014
NLPCC · 2014
Earlier work this paper cites.
TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension
Joshi, M., Choi, E., Weld, D., and Zettlemoyer, L · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Earlier work this paper cites.
The NarrativeQA reading comprehension challenge
Kočiský, T., Schwarz, J., Blunsom, P., Dyer, C., Hermann, K. M., Melis, G., and Grefenstette, E · 2018
Earlier work this paper cites.
Don‘t give me the details, just the summary! Topic-aware convolutional neural networks for extreme summarization
Narayan, S., Cohen, S. B., and Lapata, M · 2018
Earlier work this paper cites.
Hotpotqa: A dataset for diverse, explainable multi-hop question answering, 2018
Yang, Z., Qi, P., Zhang, S., Bengio, Y., Cohen, W. W., Salakhutdinov, R., and Manning, C. D · 2018
Earlier work this paper cites.
SAMSum corpus: A human-annotated dialogue dataset for abstractive summarization
Gliwa, B., Mochol, I., Biesek, M., and Wawer, A · 2019
Earlier work this paper cites.
Dense passage retrieval for open-domain question answering
Karpukhin, V., Oguz, B., Min, S., Lewis, P., Wu, L., Edunov, S., Chen, D., and Yih, W.-t · 2020
Earlier work this paper cites.
Linformer: Self-attention with linear complexity
Wang, S., Li, B. Z., Khabsa, M., Fang, H., and Ma, H · 2020
Earlier work this paper cites.
Suppressed for anonymity, 2021
Author, N. N · 2021
Earlier work this paper cites.
A dataset of information-seeking questions and answers anchored in research papers, 2021
Dasigi, P., Lo, K., Beltagy, I., Cohan, A., Smith, N. A., and Gardner, M · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2021
Earlier work this paper cites.
Efficient attentions for long document summarization
Huang, L., Cao, S., Parulian, N., Ji, H., and Wang, L · 2021
Earlier work this paper cites.
Deep reinforcement learning of transition states
Zhang, J., Lei, Y.-K., Zhang, Z., Han, X., Li, M., Yang, L., Yang, Y. I., and Gao, Y. Q · 2021
Earlier work this paper cites.
Flashattention: Fast and memory-efficient exact attention with io-awareness
Dao, T., Fu, D., Ermon, S., Rudra, A., and Ré, C · 2022
Earlier work this paper cites.
Llm.int8(): 8-bit matrix multiplication for transformers at scale, 2022
Dettmers, T., Lewis, M., Belkada, Y., and Zettlemoyer, L · 2022
Cited alongside, same era.
Retrieval as attention: End-to-end learning of retrieval and reading within a single transformer
Jiang, Z., Gao, L., Wang, Z., Araki, J., Ding, H., Callan, J., and Neubig, G · 2022
Cited alongside, same era.
Sentence-T5: Scalable sentence encoders from pre-trained text-to-text models
Ni, J., Hernandez Abrego, G., Constant, N., Ma, J., Hall, K., Cer, D., and Yang, Y · 2022
Cited alongside, same era.
Show your work: Scratchpads for intermediate computation with language models
Nye, M., Andreassen, A. J., Gur-Ari, G., Michalewski, H., Austin, J., Bieber, D., Dohan, D., Lewkowycz, A., Bosma, M., Luan, D., Sutton, C., and Odena, A · 2022
Cited alongside, same era.
Train short, test long: Attention with linear biases enables input length extrapolation
Press, O., Smith, N., and Lewis, M · 2022
Cited alongside, same era.
Kdeformer: Accelerating transformers via kernel density estimation
Zandieh, A., Han, I., Daliri, M., and Karbasi, A · 2023
Closest in time.
Introducing the next generation of Claude
Anthropic · 2024
Closest in time.
Longbench v2: Towards deeper understanding and reasoning on realistic long-context multitasks
Bai, Y., Tu, S., Zhang, J., Peng, H., Wang, X., Lv, X., Cao, S., Xu, J., Hou, L., Dong, Y., et al · 2024
Closest in time.
Deepseek-v2: A strong, economical, and efficient mixture-of-experts language model, 2024
DeepSeek-AI, Liu, A., Feng, B., Wang, B., Wang, B., Liu, B., Zhao, C., Dengr, C., Ruan, C., Dai, D., Guo, D., Yang, D., Chen, D., Ji, D., Li, E., Lin, F., Luo, F., Hao, G., Chen, G., Li, G., Zhang, H., Xu, H., Yang, H., Zhang, H., Ding, H., Xin, H., Gao, H., Li, H., Qu, H., Cai, J. L., Liang, J., Guo, J., Ni, J., Li, J., Chen, J., Yuan, J., Qiu, J., Song, J., Dong, K., Gao, K., Guan, K., Wang, L., Zhang, L., Xu, L., Xia, L., Zhao, L., Zhang, L., Li, M., Wang, M., Zhang, M., Zhang, M., Tang, M., Li, M., Tian, N., Huang, P., Wang, P., Zhang, P., Zhu, Q., Chen, Q., Du, Q., Chen, R. J., Jin, R. L., Ge, R., Pan, R., Xu, R., Chen, R., Li, S. S., Lu, S., Zhou, S., Chen, S., Wu, S., Ye, S., Ma, S., Wang, S., Zhou, S., Yu, S., Zhou, S., Zheng, S., Wang, T., Pei, T., Yuan, T., Sun, T., Xiao, W. L., Zeng, W., An, W., Liu, W., Liang, W., Gao, W., Zhang, W., Li, X. Q., Jin, X., Wang, X., Bi, X., Liu, X., Wang, X., Shen, X., Chen, X., Chen, X., Nie, X., Sun, X., Wang, X., Liu, X., Xie, X., Yu, X., Song, X., Zhou, X., Yang, X., Lu, X., Su, X., Wu, Y., Li, Y. K., Wei, Y. X., Zhu, Y. X., Xu, Y., Huang, Y., Li, Y., Zhao, Y., Sun, Y., Li, Y., Wang, Y., Zheng, Y., Zhang, Y., Xiong, Y., Zhao, Y., He, Y., Tang, Y., Piao, Y., Dong, Y., Tan, Y., Liu, Y., Wang, Y., Guo, Y., Zhu, Y., Wang, Y., Zou, Y., Zha, Y., Ma, Y., Yan, Y., You, Y., Liu, Y., Ren, Z. Z., Ren, Z., Sha, Z., Fu, Z., Huang, Z., Zhang, Z., Xie, Z., Hao, Z., Shao, Z., Wen, Z., Xu, Z., Zhang, Z., Li, Z., Wang, Z., Gu, Z., Li, Z., and Xie, Z · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chain-of-thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al · 2022
Cited alongside, same era.
Alpa: Automating inter- and Intra-Operator parallelism for distributed deep learning
Zheng, L., Li, Z., Zhang, H., Zhuang, Y., Chen, Z., Huang, Y., Wang, Y., Xu, Y., Zhuo, D., Xing, E. P., Gonzalez, J. E., and Stoica, I · 2022
Cited alongside, same era.
GQA: Training generalized multi-query transformer models from multi-head checkpoints
Ainslie, J., Lee-Thorp, J., de Jong, M., Zemlyanskiy, Y., Lebron, F., and Sanghai, S · 2023
Cited alongside, same era.
Quantizable transformers: Removing outliers by helping attention heads do nothing
Bondarenko, Y., Nagel, M., and Blankevoort, T · 2023
Cited alongside, same era.
Extending context window of large language models via positional interpolation
Chen, S., Wong, S., Chen, L., and Tian, Y · 2023
Cited alongside, same era.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, March 2023
Chiang, W.-L., Li, Z., Lin, Z., Sheng, Y., Wu, Z., Zhang, H., Zheng, L., Zhuang, S., Zhuang, Y., Gonzalez, J. E., Stoica, I., and Xing, E. P · 2023
Cited alongside, same era.
Dynamically Scaled RoPE further increases performance of long context LLaMA with zero fine-tuning, 2023
Emozilla · 2023
Cited alongside, same era.
Closest in time.
Marlin: a fast 4-bit inference kernel for medium batchsizes
Frantar, E. and Alistarh, D · 2024
Closest in time.
Model tells you what to discard: Adaptive KV cache compression for LLMs
Ge, S., Zhang, Y., Liu, L., Zhang, M., Han, J., and Gao, J · 2024
Closest in time.
Hyperattention: Long-context attention in near-linear time
Han, I., Jayaram, R., Karbasi, A., Mirrokni, V., Woodruff, D., and Zandieh, A · 2024
Closest in time.
Training compute-optimal large language models
Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., de Las Casas, D., and Hendricks · 2024
Closest in time.
RULER: What’s the real context size of your long-context language models?
Hsieh, C.-P., Sun, S., Kriman, S., Acharya, S., Rekesh, D., Jia, F., and Ginsburg, B · 2024
Closest in time.
The impact of positional encoding on length generalization in transformers
Kazemnejad, A., Padhi, I., Natesan Ramamurthy, K., Das, P., and Reddy, S · 2024
Closest in time.
Awq: Activation-aware weight quantization for llm compression and acceleration, 2024
Lin, J., Tang, J., Tang, H., Yang, S., Chen, W.-M., Wang, W.-C., Xiao, G., Dang, X., Gan, C., and Han, S · 2024
Closest in time.
Random-access infinite context length for transformers
Mohtashami, A. and Jaggi, M · 2024
Closest in time.
Nvidia ada lovelace professional gpu architecture
NVIDIA · 2024
Closest in time.
Nvbench: Nvidia’s benchmarking tool for gpus, 2024
NVIDIA · 2024
Closest in time.
New models and developer products announced at devday
OpenAI · 2024
Closest in time.
Introducing GPT-4o: our fastest and most affordable flagship model
OpenAI · 2024
Closest in time.
Transformers are multi-state RNNs
Oren, M., Hassid, M., Yarden, N., Adi, Y., and Schwartz, R · 2024
Closest in time.
Roformer: Enhanced transformer with rotary position embedding
Su, J., Ahmed, M., Lu, Y., Pan, S., Bo, W., and Liu, Y · 2024
Closest in time.
QUEST: Query-aware sparsity for efficient long-context LLM inference
Tang, J., Zhao, Y., Zhu, K., Xiao, G., Kasikci, B., and Han, S · 2024
Closest in time.
Efficient streaming language models with attention sinks
Xiao, G., Tian, Y., Chen, B., Han, S., and Lewis, M · 2024
Closest in time.
Cascade inference: Memory bandwidth efficient shared prefix batch decoding
Ye, Z., Lai, R., Lu, R., Lin, C.-Y., Zheng, S., Chen, L., Chen, T., and Ceze, L · 2024
Closest in time.
Infinitebench: Extending long context evaluation beyond 100k tokens
Zhang, X., Chen, Y., Hu, S., Xu, Z., Chen, J., Hao, M., Han, X., Thai, Z., Wang, S., Liu, Z., et al · 2024
Closest in time.
Atom: Low-bit quantization for efficient and accurate llm serving, 2024
Zhao, Y., Lin, C.-Y., Zhu, K., Ye, Z., Chen, L., Zheng, S., Ceze, L., Krishnamurthy, A., Chen, T., and Kasikci, B · 2024
Closest in time.
Sparq attention: bandwidth-efficient llm inference
Ribar, L., Chelombiev, I., Hudlass-Galley, L., Blake, C., Luschi, C., and Orr, D · 2025
Closest in time.