Fetching the paper…
Reading the bibliography…
Large language models (LLMs) with extended context windows have made significant strides yet remain a challenge due to the scarcity of long documents.
Some studies in machine learning using the game of checkers
Samuel, A. L · 1959
Earlier work this paper cites.
The need for biases in learning generalizations
Mitchell, T. M · 1980
Earlier work this paper cites.
Mechanisms of skill acquisition and the law of practice
Newell, A. and Rosenbloom, P. S · 1981
Earlier work this paper cites.
Machine Learning: An Artificial Intelligence Approach, Vol. I
Michalski, R. S., Carbonell, J. G., and Mitchell, T. M. (eds.) · 1983
Earlier work this paper cites.
Computational Complexity of Machine Learning
Kearns, M. J · 1989
Earlier work this paper cites.
Pattern Classification
Duda, R. O., Hart, P. E., and Stork, D. G · 2000
Earlier work this paper cites.
Crafting papers on machine learning
Langley, P · 2000
Earlier work this paper cites.
On the integration of structure indexes and inverted lists
Kaushik, R., Krishnamurthy, R., Naughton, J. F., and Ramakrishnan, R · 2004
Earlier work this paper cites.
A fast learning algorithm for deep belief nets
Hinton, G. E., Osindero, S., and Teh, Y. W · 2006
Earlier work this paper cites.
Scaling learning algorithms towards AI
Bengio, Y. and LeCun, Y · 2007
Earlier work this paper cites.
Visualizing data using t-sne
van der Maaten, L. and Hinton, G · 2008
Earlier work this paper cites.
The inverted multi-index
Babenko, A. and Lempitsky, V · 2014
Earlier work this paper cites.
Deep learning , volume 1
Goodfellow, I., Bengio, Y., Courville, A., and Bengio, Y · 2016
Earlier work this paper cites.
The lambada dataset: Word prediction requiring a broad discourse context
Paperno, D., Kruszewski, G., Lazaridou, A., Pham, Q. N., Bernardi, R., Pezzelle, S., Baroni, M., Boleda, G., and Fernández, R · 2016
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2017
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge
Clark, P., Cowhey, I., Etzioni, O., Khot, T., Sabharwal, A., Schoenick, C., and Tafjord, O · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Radford, A · 2018
Earlier work this paper cites.
Document expansion by query prediction
Nogueira, R., Yang, W., Lin, J., and Cho, K · 2019
Earlier work this paper cites.
Sagawa, S., Koh, P. W., Hashimoto, T. B., and Liang, P · 2019
Earlier work this paper cites.
Hellaswag: Can a machine really finish your sentence?
Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y · 2019
Earlier work this paper cites.
Piqa: Reasoning about physical commonsense in natural language
Bisk, Y., Zellers, R., Gao, J., Choi, Y., et al · 2020
Earlier work this paper cites.
The pile: An 800gb dataset of diverse text for language modeling
Gao, L., Biderman, S., Black, S., Golding, L., Hoppe, T., Foster, C., Phang, J., He, H., Thite, A., Nabeshima, N., et al · 2020
Earlier work this paper cites.
Retrieval augmented language model pre-training
Guu, K., Lee, K., Tung, Z., Pasupat, P., and Chang, M · 2020
Earlier work this paper cites.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2020
Earlier work this paper cites.
Scaling laws for autoregressive generative modeling
Henighan, T., Kaplan, J., Katz, M., Chen, M., Hesse, C., Jackson, J., Jun, H., Brown, T. B., Dhariwal, P., Gray, S., et al · 2020
Earlier work this paper cites.
Hard negative mixing for contrastive learning
Kalantidis, Y., Sariyildiz, M. B., Pion, N., Weinzaepfel, P., and Larlus, D · 2020
Earlier work this paper cites.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W.-t., Rocktäschel, T., et al · 2020
Earlier work this paper cites.
Logiqa: A challenge dataset for machine reading comprehension with logical reasoning
Liu, J., Cui, L., Liu, H., Huang, D., Wang, Y., and Zhang, Y · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Earlier work this paper cites.
Zero: Memory optimizations toward training trillion parameter models
Rajbhandari, S., Rasley, J., Ruwase, O., and He, Y · 2020
Earlier work this paper cites.
Contrastive learning with hard negative samples
Robinson, J., Chuang, C.-Y., Sra, S., and Jegelka, S · 2020
Earlier work this paper cites.
Approximate nearest neighbor negative contrastive learning for dense text retrieval
Xiong, L., Xiong, C., Li, Y., Tang, K.-F., Liu, J., Bennett, P., Ahmed, J., and Overwijk, A · 2020
Earlier work this paper cites.
Suppressed for anonymity, 2021
Author, N. N · 2021
Earlier work this paper cites.
Incremental false negative detection for contrastive learning
Chen, T.-S., Hung, W.-C., Tseng, H.-Y., Chien, S.-Y., and Yang, M.-H · 2021
Cited alongside, same era.
The inductive bias of in-context learning: Rethinking pretraining example design
Levine, Y., Wies, N., Jannai, D., Navon, D., Hoshen, Y., and Shashua, A · 2021
Cited alongside, same era.
Learning passage impacts for inverted indexes
Mallia, A., Khattab, O., Suel, T., and Tonellotto, N · 2021
Cited alongside, same era.
Winogrande: An adversarial winograd schema challenge at scale
Sakaguchi, K., Bras, R. L., Bhagavatula, C., and Choi, Y · 2021
Cited alongside, same era.
Roformer: Enhanced transformer with rotary position embedding, 2021
Su, J., Lu, Y., Pan, S., Wen, B., and Liu, Y · 2021
Cited alongside, same era.
Trams: Training-free memory selection for long-range language modeling
Yu, H., Zhang, Y., Bi, W., et al · 2023
Later among the works it cites.
Pose: Efficient context window extension of llms via positional skip-wise training
Zhu, D., Yang, N., Wang, L., Song, Y., Wu, W., Wei, F., and Li, S · 2023
Later among the works it cites.
Agarwal, R., Singh, A., Zhang, L. M., Bohnet, B., Rosias, L., Chan, S., Zhang, B., Anand, A., Abbas, Z., Nova, A., et al · 2024
Later among the works it cites.
Yi: Open foundation models by 01.ai, 2024
AI, ., :, Young, A., Chen, B., Li, C., Huang, C., Zhang, G., Zhang, G., Li, H., Zhu, J., Chen, J., Chang, J., Yu, K., Liu, P., Liu, Q., Yue, S., Yang, S., Yang, S., Yu, T., Xie, W., Huang, W., Hu, X., Ren, X., Niu, X., Nie, P., Xu, Y., Liu, Y., Wang, Y., Cai, Y., Gu, Z., Liu, Z., and Dai, Z · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Optimizing dense retrieval model training with hard negatives
Zhan, J., Mao, J., Liu, Y., Guo, J., Zhang, M., and Ma, S · 2021
Cited alongside, same era.
Revisiting neural scaling laws in language and vision
Alabdulmohsin, I. M., Neyshabur, B., and Zhai, X · 2022
Cited alongside, same era.
Training compute-optimal large language models
Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., Casas, D. d. L., Hendricks, L. A., Welbl, J., Clark, A., et al · 2022
Cited alongside, same era.
The stack: 3 tb of permissively licensed source code
Kocetkov, D., Li, R., Allal, L. B., Li, J., Mou, C., Ferrandis, C. M., Jernite, Y., Mitchell, M., Hughes, S., Wolf, T., et al · 2022
Cited alongside, same era.
Lifelong and continual learning dialogue systems
Mazumder, S. and Liu, B · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Cited alongside, same era.
Query-as-context pre-training for dense passage retrieval
Wu, X., Ma, G., Qian, W., Lin, Z., and Hu, S · 2022
Cited alongside, same era.
An, C., Huang, F., Zhang, J., Gong, S., Qiu, X., Zhou, C., and Kong, L · 2024
Later among the works it cites.
Many-shot jailbreaking
Anil, C., Durmus, E., Sharma, M., Benton, J., Kundu, S., Batson, J., Rimsky, N., Tong, M., Mu, J., Ford, D., et al · 2024
Later among the works it cites.
Longbench v2: Towards deeper understanding and reasoning on realistic long-context multitasks
Bai, Y., Tu, S., Zhang, J., Peng, H., Wang, X., Lv, X., Cao, S., Xu, J., Hou, L., Dong, Y., Tang, J., and Li, J · 2024
Later among the works it cites.
Smollm-corpus, 2024
Ben Allal, L., Lozhkov, A., Penedo, G., Wolf, T., and von Werra, L · 2024
Later among the works it cites.
In-context learning with long-context models: An in-depth exploration
Bertsch, A., Ivgi, M., Alon, U., Berant, J., Gormley, M. R., and Neubig, G · 2024
Later among the works it cites.
Deepseek llm: Scaling open-source language models with longtermism
Bi, X., Chen, D., Chen, G., Chen, S., Dai, D., Deng, C., Ding, H., Dong, K., Du, Q., Fu, Z., et al · 2024
Later among the works it cites.
Language models as science tutors
Chevalier, A., Geng, J., Wettig, A., Chen, H., Mizera, S., Annala, T., Aragon, M. J., Fanlo, A. R., Frieder, S., Machado, S., et al · 2024
Later among the works it cites.
Deepseek-v2: A strong, economical, and efficient mixture-of-experts language model, 2024
DeepSeek-AI · 2024
Later among the works it cites.
Fewer truncations improve language modeling
Ding, H., Wang, Z., Paolini, G., Kumar, V., Deoras, A., Roth, D., and Soatto, S · 2024
Later among the works it cites.
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al · 2024
Later among the works it cites.
Data engineering for scaling language models to 128k context
Fu, Y., Panda, R., Niu, X., Yue, X., Hajishirzi, H., Kim, Y., and Peng, H · 2024
Later among the works it cites.
Ruler: What’s the real context size of your long-context language models?
Hsieh, C.-P., Sun, S., Kriman, S., Acharya, S., Rekesh, D., Jia, F., and Ginsburg, B · 2024
Later among the works it cites.
Understanding the effects of language-specific class imbalance in multilingual fine-tuning
Jung, V. and van der Plas, L · 2024
Later among the works it cites.
World model on million-length video and language with ringattention
Liu, H., Yan, W., Zaharia, M., and Abbeel, P · 2024
Later among the works it cites.
Fineweb-edu, May 2024
Lozhkov, A., Ben Allal, L., von Werra, L., and Wolf, T · 2024
Later among the works it cites.
Introducing meta llama 3: The most capable openly available llm to date, 2024
Meta · 2024
Later among the works it cites.
Dolma: An open corpus of three trillion tokens for language model pretraining research
Soldaini, L., Kinney, R., Bhagia, A., Schwenk, D., Atkinson, D., Authur, R., Bogin, B., Chandu, K., Dumas, J., Elazar, Y., et al · 2024
Later among the works it cites.
Unraveling the mystery of scaling laws: Part i
Su, H., Tian, Z., Shen, X., and Cai, X · 2024
Later among the works it cites.
Qwen2.5: A party of foundation models, September 2024
Team, Q · 2024
Later among the works it cites.
Tian, J., Zheng, D., Cheng, Y., Wang, R., Zhang, C., and Zhang, D · 2024
Later among the works it cites.
Focused transformer: Contrastive training for context scaling
Tworkowski, S., Staniszewski, K., Pacek, M., Wu, Y., Michalewski, H., and Miłoś, P · 2024
Later among the works it cites.
A survey on large language model based autonomous agents
Wang, L., Ma, C., Feng, X., Zhang, Z., Yang, H., Zhang, J., Chen, Z., Tang, J., Chen, X., Lin, Y., et al · 2024
Later among the works it cites.
Doremi: Optimizing data mixtures speeds up language model pretraining
Xie, S. M., Pham, H., Dong, X., Du, N., Liu, H., Lu, Y., Liang, P. S., Le, Q. V., Ma, T., and Yu, A. W · 2024
Later among the works it cites.
Temporal scaling law for large language models
Xiong, Y., Chen, X., Ye, X., Chen, H., Lin, Z., Lian, H., Niu, J., and Ding, G · 2024
Later among the works it cites.
Magpie: Alignment data synthesis from scratch by prompting aligned llms with nothing
Xu, Z., Jiang, F., Niu, L., Deng, Y., Poovendran, R., Choi, Y., and Lin, B. Y · 2024
Later among the works it cites.
Yang, A., Yang, B., Hui, B., Zheng, B., Yu, B., Zhou, C., Li, C., Li, C., Liu, D., Huang, F., Dong, G., Wei, H., Lin, H., Tang, J., Wang, J., Yang, J., Tu, J., Zhang, J., Ma, J., Xu, J., Zhou, J., Bai, J., He, J., Lin, J., Dang, K., Lu, K., Chen, K., Yang, K., Li, M., Xue, M., Ni, N., Zhang, P., Wang, P., Peng, R., Men, R., Gao, R., Lin, R., Wang, S., Bai, S., Tan, S., Zhu, T., Li, T., Liu, T., Ge, W., Deng, X., Zhou, X., Ren, X., Zhang, X., Wei, X., Ren, X., Fan, Y., Yao, Y., Zhang, Y., Wan, Y., Chu, Y., Liu, Y., Cui, Z., Zhang, Z., and Fan, Z · 2024
Later among the works it cites.
∞ \infty Bench: Extending long context evaluation beyond 100K tokens
Zhang, X., Chen, Y., Hu, S., Xu, Z., Chen, J., Hao, M., Han, X., Thai, Z., Wang, S., Liu, Z., and Sun, M · 2024
Later among the works it cites.
Longskywork: A training recipe for efficiently extending context length in large language models
Zhao, L., Wei, T., Zeng, L., Cheng, C., Yang, L., Cheng, P., Wang, L., Li, C., Wu, X., Zhu, B., Gan, Y., Hu, R., Yan, S., Fang, H., and Zhou, Y · 2024
Later among the works it cites.
From system 1 to system 2: A survey of reasoning large language models, 2025
Li, Z.-Z., Zhang, D., Zhang, M.-L., Zhang, J., Liu, Z., Yao, Y., Xu, H., Zheng, J., Wang, P.-J., Chen, X., Zhang, Y., Yin, F., Dong, J., Li, Z., Bi, B.-L., Mei, L.-R., Fang, J., Guo, Z., Song, L., and Liu, C.-L · 2025
Closest in time.