Fetching the paper…
Reading the bibliography…
Large language models (LLMs) are typically developed through large-scale pre-training followed by task-specific fine-tuning.
1908
Earlier work this paper cites.
2015
Earlier work this paper cites.
N. Tishby and N. Zaslavsky, “Deep learning and the information bottleneck principle,” in IEEE ITW , 2015, pp. 1–5
2015
Earlier work this paper cites.
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Adv. Neural Inf. Process. Syst. , vol. 30, 2017
2017
Earlier work this paper cites.
L. N. Smith, “Cyclical learning rates for training neural networks,” in IEEE WACV . IEEE, 2017, pp. 464–472
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
N. Shazeer and M. Stern, “Adafactor: Adaptive learning rates with sublinear memory cost,” in ICML , 2018, pp. 4596–4604
2018
Earlier work this paper cites.
P. Clark, I. Cowhey, O. Etzioni, T. Khot, A. Sabharwal, C. Schoenick, and O. Tafjord, “Think you have solved question answering? try arc, the ai2 reasoning challenge,” 2018
2018
Earlier work this paper cites.
C. B. Clement, M. Bierbaum, K. P. O’Keeffe, and A. A. Alemi, “On the use of arxiv as a dataset,” 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “Bert: Pre-training of deep bidirectional transformers for language understanding,” in NAACL-HLT , 2019, pp. 4171–4186
2019
Earlier work this paper cites.
R. Zellers, A. Holtzman, Y. Bisk, A. Farhadi, and Y. Choi, “Hellaswag: Can a machine really finish your sentence?” 2019
2019
Earlier work this paper cites.
K. Sakaguchi, R. L. Bras et al. , “Winogrande: An adversarial winograd schema challenge at scale,” 2019
2019
Earlier work this paper cites.
Y. Bisk, R. Zellers et al. , “Piqa: Reasoning about physical commonsense in natural language,” 2019
2019
Earlier work this paper cites.
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
T. Brown, B. Mann, N. Ryder, M. Subbiah et al. , “Language models are few-shot learners,” Adv. Neural Inf. Process. Syst. , vol. 33, pp. 1877–1901, 2020
2020
Earlier work this paper cites.
C. Raffel, N. Shazeer, A. Roberts, K. Lee et al. , “Exploring the limits of transfer learning with a unified text-to-text transformer,” J. Mach. Learn. Res , vol. 21, no. 140, pp. 1–67, 2020
2020
Earlier work this paper cites.
M. Tancik, P. Srinivasan et al. , “Fourier features let networks learn high frequency functions in low dimensional domains,” Adv. Neural Inf. Process. Syst. , vol. 33, pp. 7537–7547, 2020
2020
Earlier work this paper cites.
2021
Earlier work this paper cites.
allenai. (2021)
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt, “Measuring massive multitask language understanding,” 2021
2021
Earlier work this paper cites.
D. Hendrycks, C. Burns, S. Kadavath, A. Arora, S. Basart, E. Tang, D. Song, and J. Steinhardt, “Measuring mathematical problem solving with the math dataset,” 2021
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
J. Austin, A. Odena, M. Nye et al. , “Program synthesis with large language models,” 2021
2021
Earlier work this paper cites.
N. Goyal, C. Gao, V. Chaudhary, P.-J. Chen et al. , “The flores-101 evaluation benchmark for low-resource and multilingual machine translation,” 2021
2021
Earlier work this paper cites.
P. Izmailov, P. Kirichenko, N. Gruver, and A. G. Wilson, “On feature learning in the presence of spurious correlations,” Adv. Neural Inf. Process. Syst. , vol. 35, pp. 38 516–38 532, 2022
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
J. Gilmer, B. Ghorbani et al. , “A loss curvature perspective on training instabilities of deep learning models,” in ICLR , 2022
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
D. Kocetkov, R. Li, L. Ben Allal, J. Li et al. , “The stack: 3 tb of permissively licensed source code,” Preprint , 2022
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
S. Kreps, R. M. McCain, and M. Brundage, “All the news that’s fit to fabricate: Ai-generated text as a tool of media misinformation,” J EXP POLIT SCI. , vol. 9, no. 1, pp. 104–117, 2022
2022
Earlier work this paper cites.
M. Suzgun, N. Scales, N. Schärli et al. , “Challenging big-bench tasks and whether chain-of-thought can solve them,” 2022
2022
Earlier work this paper cites.
F. Cassano, J. Gouwar et al. , “Multipl-e: A scalable and extensible approach to benchmarking neural code generation,” 2022
2022
Earlier work this paper cites.
F. Shi, M. Suzgun, M. Freitag et al. , “Language models are multilingual chain-of-thought reasoners,” 2022
2022
Earlier work this paper cites.
2023
Earlier work this paper cites.
H. Jin, W. Wei, X. Wang, W. Zhang, and Y. Wu, “Rethinking learning rate tuning in the era of large language models,” in 2023 CogMI . IEEE, 2023, pp. 112–121
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
bloc97, “Ntk-aware scaled rope allows LLaMA models to have extended (8k+) context size without any fine-tuning and minimal perplexity degradation,” Jun 2023, reddit post, r/LocalLLaMA
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Cited alongside, same era.
K. Paster, M. D. Santos, Z. Azerbayev, and J. Ba, “Openwebmath: An open dataset of high-quality mathematical web text,” 2023
2023
Cited alongside, same era.
W. Lian, G. Wang, B. Goodson et al. , “Slimorca: An open dataset of gpt-4 augmented flan reasoning traces, with verification,” 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
L. Soldaini and K. Lo, “peS2o (Pretraining Efficiently on S2ORC) Dataset,” Allen Institute for AI, Tech. Rep., 2023, oDC-By, https://github.com/allenai/pes2o
L. Ben Allal, A. Lozhkov, G. Penedo, T. Wolf, and L. von Werra, “Smollm-corpus,” 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
S. Chaudhry and A. Sharma, “Data distribution-based curriculum learning,” IEEE Access , 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
A. Dubey, A. Jauhri et al. , “The llama 3 herd of models,” ArXiv , vol. abs/2407.21783, 2024
2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Z. Luo, C. Xu, P. Zhao et al. , “Wizardcoder: Empowering code large language models with evol-instruct,” 2023
2023
Cited alongside, same era.
W. Lian, B. Goodson, E. Pentland, A. Cook, C. Vong, and ”Teknium”, “Openorca: An open dataset of gpt augmented flan reasoning traces,” https://https://huggingface.co/datasets/Open-Orca/OpenOrca , 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
DeepSeek-AI, A. Liu et al. , “Deepseek-v3 technical report,” ArXiv , vol. abs/2412.19437, 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
I. Granite Team, “Granite 3.0 language models,” URL: https://github. com/ibm-granite/granite-3.0-language-models , 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
Q. A. Yang, B. Yang et al. , “Qwen2.5 technical report,” ArXiv , vol. abs/2412.15115, 2024
2024
Later among the works it cites.
Y. Wang, X. Ma et al. , “Mmlu-pro: A more robust and challenging multi-task language understanding benchmark,” 2024
2024
Later among the works it cites.
Y. Bai, X. Lv, J. Zhang et al. , “Longbench: A bilingual, multitask benchmark for long context understanding,” 2024
2024
Later among the works it cites.
C.-P. Hsieh, S. Sun, S. Kriman, S. Acharya et al. , “Ruler: What’s the real context size of your long-context language models?” 2024
2024
Later among the works it cites.
2025
Closest in time.
2025
Closest in time.
O. Wu, “Data optimization for llms: A survey,” Authorea Preprints , 2025
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
Y. Zhang, Y. Luo, Y. Yuan, and A. C.-C. Yao, “Autonomous data selection with zero-shot generative classifiers for mathematical texts,” ACL Findings , 2025
2025
Closest in time.
N. Z. Weingarten, Z. Yakhini, M. Butman, and R. Bustin, “The supervised information bottleneck,” Entropy , vol. 27, no. 5, p. 452, 2025
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
Qwen, :, A. Yang, B. Yang, B. Zhang, B. Hui et al. , “Qwen2.5 technical report,” 2025
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
A. Yang, B. Yu et al. , “Qwen2. 5-1m technical report,” arXiv preprint arXiv:2501.15383 , 2025
2025
Closest in time.
2025
Closest in time.
Meta AI, “Llama 4 Scout 17B-16E Instruct,” https://huggingface.co/meta-llama/Llama-4-Scout-17B-16E-Instruct , 2025, model card. Version effective 5 Apr 2025. Accessed 3 Jul 2025
2025
Closest in time.
2025
Closest in time.
E. Bakouch, C. M. Patiño et al. , “Smollm3: smol, multilingual, long-context reasoner,” Hugging Face Blog, July 2025
2025
Closest in time.