Fetching the paper…
Reading the bibliography…
We study the continual pretraining recipe for scaling language models' context lengths to 128K, with a focus on data engineering.
Compressive transformers for long-range sequence modelling
Rae, J. W., Potapenko, A., Jayakumar, S. M., Hillier, C., and Lillicrap, T. P · 1911
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2020
Earlier work this paper cites.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Earlier work this paper cites.
Zero: Memory optimizations toward training trillion parameter models
Rajbhandari, S., Rasley, J., Ruwase, O., and He, Y · 2020
Earlier work this paper cites.
Evaluating large language models trained on code
Chen, M., Tworek, J., Jun, H., Yuan, Q., Pinto, H. P. d. O., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., et al · 2021
Earlier work this paper cites.
Sequence parallelism: Long sequence training from system perspective
Li, S., Xue, F., Baranwal, C., Li, Y., and You, Y · 2021
Earlier work this paper cites.
Train short, test long: Attention with linear biases enables input length extrapolation
Press, O., Smith, N. A., and Lewis, M · 2021
Earlier work this paper cites.
Training compute-optimal large language models
Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., Casas, D. d. L., Hendricks, L. A., Welbl, J., Clark, A., et al · 2022
Earlier work this paper cites.
A length-extrapolatable transformer
Sun, Y., Dong, L., Patra, B., Ma, S., Huang, S., Benhaim, A., Chaudhary, V., Song, X., and Wei, F · 2022
Earlier work this paper cites.
L-eval: Instituting standardized evaluation for long context language models
An, C., Gong, S., Zhong, M., Li, M., Zhang, J., Kong, L., and Qiu, X · 2023
Earlier work this paper cites.
Model card and evaluations for claude models, July 2023
Anthropic · 2023
Cited alongside, same era.
Llemma: An open language model for mathematics
Azerbayev, Z., Schoelkopf, H., Paster, K., Santos, M. D., McAleer, S., Jiang, A. Q., Deng, J., Biderman, S., and Welleck, S · 2023
Cited alongside, same era.
Codeplan: Repository-level coding using llms and planning
Bairi, R., Sonwane, A., Kanade, A., Iyer, A., Parthasarathy, S., Rajamani, S., Ashok, B., Shet, S., et al · 2023
Cited alongside, same era.
Peek across: Improving multi-document modeling via cross-document question-answering
Caciularu, A., Peters, M. E., Goldberger, J., Dagan, I., and Cohan, A · 2023
Cited alongside, same era.
Flashattention-2: Faster attention with better parallelism and work partitioning
Hierarchically gated recurrent neural network for sequence modeling
Qin, Z., Yang, S., and Zhong, Y · 2023
Later among the works it cites.
Zeroscrolls: A zero-shot benchmark for long text understanding, 2023
Shaham, U., Ivgi, M., Efrat, A., Berant, J., and Levy, O · 2023
Later among the works it cites.
SlimPajama: A 627B token cleaned and deduplicated version of RedPajama
Soboleva, D., Al-Khateeb, F., Myers, R., Steeves, J. R., Hestness, J., and Dey, N · 2023
Later among the works it cites.
Retentive network: A successor to transformer for large language models
Sun, Y., Dong, L., Huang, S., Ma, S., Xia, Y., Xue, J., Wang, J., and Wei, F · 2023
Later among the works it cites.
Introducing mpt-7b: A new standard for open-source, commercially usable llms, 2023
Team, M. N · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dao, T · 2023
Cited alongside, same era.
Openmoe: Open mixture-of-experts language models
Fuzhao Xue, Zian Zheng, Y. F. J. N. Z. Z. W. Z. and You, Y · 2023
Cited alongside, same era.
Mamba: Linear-time sequence modeling with selective state spaces
Gu, A. and Dao, T · 2023
Cited alongside, same era.
Jacobs, S. A., Tanaka, M., Zhang, C., Zhang, M., Song, L., Rajbhandari, S., and He, Y · 2023
Cited alongside, same era.
Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., Casas, D. d. l., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., et al · 2023
Cited alongside, same era.
Needle in a haystack - pressure testing llms
Kamradt, G · 2023
Cited alongside, same era.
How long can open-source llms truly promise on context length?, June 2023a
Li, D., Shao, R., Xie, A., Sheng, Y., Zheng, L., Gonzalez, J. E., Stoica, I., Ma, X., and Zhang, H · 2023
Cited alongside, same era.
Yarn: Efficient context window extension of large language models, 2023
Peng, B., Quesnelle, J., Fan, H., and Shippole, E · 2023
Cited alongside, same era.
Llama-2-7b-32k-instruct — and fine-tuning for llama-2 models with together api, August 2023
Together · 2023
Later among the works it cites.
Llm-powered autonomous agents
Weng, L · 2023
Later among the works it cites.
Efficient streaming language models with attention sinks
Xiao, G., Tian, Y., Chen, B., Han, S., and Lewis, M · 2023
Later among the works it cites.
Effective long-context scaling of foundation models
Xiong, W., Liu, J., Molybog, I., Zhang, H., Bhargava, P., Hou, R., Martin, L., Rungta, R., Sankararaman, K. A., Oguz, B., et al · 2023
Later among the works it cites.
Infinitebench: 128k long-context benchmark for language models, 2023
Zhang, X., Chen, Y., Hu, S., Wu, Q., Chen, J., Xu, Z., Dai, Z., Han, X., Wang, S., Liu, Z., and Sun, M · 2023
Later among the works it cites.
Lifelong and Continual Learning Dialogue Systems
Mazumder, S. and Liu, B · 2024
Closest in time.
Xverse-13b, January 2024
XVerse · 2024
Closest in time.