Fetching the paper…
Reading the bibliography…
We extend the context length of Llama-3-8B-Instruct from 8K to 80K via QLoRA fine-tuning.
Measuring massive multitask language understanding, 2021
D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt · 2021
Earlier work this paper cites.
https://github.com/unslothai/unsloth , 2023
Unsloth.ai · 2023
Earlier work this paper cites.
Longbench: A bilingual, multitask benchmark for long context understanding, 2023
Y. Bai, X. Lv, J. Zhang, H. Lyu, J. Tang, Z. Huang, Z. Du, X. Liu, A. Zeng, L. Hou, Y. Dong, J. Tang, and J. Li · 2023
Earlier work this paper cites.
Extending context window of large language models via positional interpolation, 2023
S. Chen, S. Wong, L. Chen, and Y. Tian · 2023
Earlier work this paper cites.
Redpajama: An open source recipe to reproduce llama training dataset, 2023
T. Computer · 2023
Earlier work this paper cites.
Qlora: Efficient finetuning of quantized llms, 2023
T. Dettmers, A. Pagnoni, A. Holtzman, and L. Zettlemoyer · 2023
Earlier work this paper cites.
Mistral 7b, 2023
A. Q. Jiang, A. Sablayrolles, A. Mensch, C. Bamford, D. S. Chaplot, D. de las Casas, F. Bressand, G. Lengyel, G. Lample, L. Saulnier, L. R. Lavaud, M.-A. Lachaux, P. Stock, T. L. Scao, T. Lavril, T. Wang, T. Lacroix, and W. E. Sayed · 2023
Cited alongside, same era.
How long can open-source llms truly promise on context length?, June 2023
D. Li*, R. Shao*, A. Xie, Y. Sheng, L. Zheng, J. E. Gonzalez, I. Stoica, X. Ma, , and H. Zhang · 2023
Cited alongside, same era.
Yarn: Efficient context window extension of large language models, 2023
B. Peng, J. Quesnelle, H. Fan, and E. Shippole · 2023
Cited alongside, same era.
Llama 2: Open foundation and fine-tuned chat models, 2023
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, D. Bikel, L. Blecher, C. C. Ferrer, M. Chen, G. Cucurull, D. Esiobu, J. Fernandes, J. Fu, W. Fu, B. Fuller, C. Gao, V. Goswami, N. Goyal, A. Hartshorn, S. Hosseini, R. Hou, H. Inan, M. Kardas, V. Kerkez, M. Khabsa, I. Kloumann, A. Korenev, P. S. Koura, M.-A. Lachaux, T. Lavril, J. Lee, D. Liskovich, Y. Lu, Y. Mao, X. Martinet, T. Mihaylov, P. Mishra, I. Molybog, Y. Nie, A. Poulton, J. Reizenstein, R. Rungta, K. Saladi, A. Schelten, R. Silva, E. M. Smith, R. Subramanian, X. E. Tan, B. Tang, R. Taylor, A. Williams, J. X. Kuan, P. Xu, Z. Yan, I. Zarov, Y. Zhang, A. Fan, M. Kambadur, S. Narang, A. Rodriguez, R. Stojnic, S. Edunov, and T. Scialom · 2023
Cited alongside, same era.
Longrope: Extending llm context window beyond 2 million tokens, 2024
Y. Ding, L. L. Zhang, C. Zhang, Y. Xu, N. Shang, J. Xu, F. Yang, and M. Yang · 2024
Closest in time.
Data engineering for scaling language models to 128k context, 2024
Y. Fu, R. Panda, X. Niu, X. Yue, H. Hajishirzi, Y. Kim, and H. Peng · 2024
Closest in time.
Gpt-4 technical report, 2024
OpenAI · 2024
Closest in time.
Soaring from 4k to 400k: Extending llm’s context with activation beacon, 2024
P. Zhang, Z. Liu, S. Xiao, N. Shao, Q. Ye, and Z. Dou · 2024
Closest in time.
∞ \infty bench: Extending long context evaluation beyond 100k tokens, 2024
X. Zhang, Y. Chen, S. Hu, Z. Xu, J. Chen, M. K. Hao, X. Han, Z. L. Thai, S. Wang, Z. Liu, and M. Sun · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Make your llm fully utilize the context, 2024
S. An, Z. Ma, Z. Lin, N. Zheng, and J.-G. Lou · 2024
Cited alongside, same era.
Longlora: Efficient fine-tuning of long-context large language models, 2024
Y. Chen, S. Qian, H. Tang, X. Lai, Z. Liu, S. Han, and J. Jia · 2024
Cited alongside, same era.