Fetching the paper…
Reading the bibliography…
High-quality supervised fine-tuning (SFT) data are crucial for eliciting strong capabilities from pretrained large language models (LLMs).
Does learning require memorization? a short tale about a long tail, 2021
Feldman, V · 1906
Earlier work this paper cites.
Advantage-weighted regression: Simple and scalable off-policy reinforcement learning, 2019
Peng, X. B., Kumar, A., Zhang, G., and Levine, S · 1910
Earlier work this paper cites.
Self-distillation amplifies regularization in hilbert space, 2020
Mobahi, H., Farajtabar, M., and Bartlett, P. L · 2002
Earlier work this paper cites.
Language models are few-shot learners, 2020
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2005
Earlier work this paper cites.
Elements of Information Theory
Cover, T. M. and Thomas, J. A · 2006
Earlier work this paper cites.
Evaluation of similarity-based explanations, 2021
Hanawa, K., Yokoi, S., Hara, S., and Inui, K · 2006
Earlier work this paper cites.
Critic regularized regression, 2021
Wang, Z., Novikov, A., Zolna, K., Springenberg, J. T., Reed, S., Shahriari, B., Siegel, N., Merel, J., Gulcehre, C., Heess, N., and de Freitas, N · 2006
Earlier work this paper cites.
Dataset condensation with gradient matching, 2021
Zhao, B., Mopuri, K. R., and Bilen, H · 2006
Earlier work this paper cites.
Better fine-tuning by reducing representational collapse, 2020
Aghajanyan, A., Shrivastava, A., Gupta, A., Goyal, N., Zettlemoyer, L., and Gupta, S · 2008
Earlier work this paper cites.
Measuring massive multitask language understanding, 2021a
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2009
Earlier work this paper cites.
A mathematical exploration of why language models help solve downstream tasks, 2021
Saunshi, N., Malladi, S., and Arora, S · 2010
Earlier work this paper cites.
On a connection between importance sampling and the likelihood ratio policy gradient
Tang, J. and Abbeel, P · 2010
Earlier work this paper cites.
Towards understanding ensemble, knowledge distillation and self-distillation in deep learning, 2023
Allen-Zhu, Z. and Li, Y · 2012
Earlier work this paper cites.
Information Geometry and Its Applications , volume 194 of Applied Mathematical Sciences
ichi Amari, S · 2016
Earlier work this paper cites.
Doubly robust off-policy value evaluation for reinforcement learning, 2016
Jiang, N. and Li, L · 2016
Earlier work this paper cites.
Off-policy deep reinforcement learning without exploration
Fujimoto, S., Meger, D., and Precup, D · 2018
Earlier work this paper cites.
Stabilizing off-policy q-learning via bootstrapping error reduction
Kumar, A., Fu, J., Soh, M., Tucker, G., and Levine, S · 2019
Earlier work this paper cites.
Program synthesis with large language models, 2021
Austin, J., Odena, A., Nye, M., Bosma, M., Michalewski, H., Dohan, D., Jiang, E., Cai, C., Terry, M., Le, Q., and Sutton, C · 2021
Earlier work this paper cites.
Evaluating large language models trained on code, 2021
Chen, M., Tworek, J., Jun, H., Yuan, Q., de Oliveira Pinto, H. P., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., Ray, A., Puri, R., Krueger, G., Petrov, M., Khlaaf, H., Sastry, G., Mishkin, P., Chan, B., Gray, S., Ryder, N., Pavlov, M., Power, A., Kaiser, L., Bavarian, M., Winter, C., Tillet, P., Such, F. P., Cummings, D., Plappert, M., Chantzis, F., Barnes, E., Herbert-Voss, A., Guss, W. H., Nichol, A., Paino, A., Tezak, N., Tang, J., Babuschkin, I., Balaji, S., Jain, S., Saunders, W., Hesse, C., Carr, A. N., Leike, J., Achiam, J., Misra, V., Morikawa, E., Radford, A., Knight, M., Brundage, M., Murati, M., Mayer, K., Welinder, P., McGrew, B., Amodei, D., McCandlish, S., Sutskever, I., and Zaremba, W · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems, 2021
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., Hesse, C., and Schulman, J · 2021
Earlier work this paper cites.
Evaluating distributional distortion in neural language modeling
LeBrun, B., Sordoni, A., and O’Donnell, T. J · 2021
Earlier work this paper cites.
Fine-tuning can distort pretrained features and underperform out-of-distribution, 2022
Kumar, A., Raghunathan, A., Jones, R., Ma, T., and Liang, P · 2022
Earlier work this paper cites.
Competition-level code generation with alphacode
Li, Y., Choi, D., Chung, J., Kushman, N., Schrittwieser, J., Leblond, R., Eccles, T., Keeling, J., Gimeno, F., Dal Lago, A., Hubert, T., Choy, P., de Masson d’Autume, C., Babuschkin, I., Chen, X., Huang, P.-S., Welbl, J., Gowal, S., Cherepanov, A., Molloy, J., Mankowitz, D. J., Sutherland Robson, E., Kohli, P., de Freitas, N., Kavukcuoglu, K., and Vinyals, O · 2022
Earlier work this paper cites.
Prioritized training on points that are learnable, worth learning, and not yet learnt, 2022
Mindermann, S., Brauner, J., Razzak, M., Sharma, M., Kirsch, A., Xu, W., Höltgen, B., Gomez, A. N., Morisot, A., Farquhar, S., and Gal, Y · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Earlier work this paper cites.
Learning to retrieve prompts for in-context learning, 2022
Rubin, O., Herzig, J., and Berant, J · 2022
Earlier work this paper cites.
Challenging big-bench tasks and whether chain-of-thought can solve them, 2022
Suzgun, M., Scales, N., Schärli, N., Gehrmann, S., Tay, Y., Chung, H. W., Chowdhery, A., Le, Q. V., Chi, E. H., Zhou, D., and Wei, J · 2022
Earlier work this paper cites.
Chain of thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., brian ichter, Xia, F., Chi, E. H., Le, Q. V., and Zhou, D · 2022
Earlier work this paper cites.
A theory for emergence of complex skills in language models
Arora, S. and Goyal, A · 2023
Earlier work this paper cites.
A general theoretical paradigm to understand learning from human preferences, 2023
Azar, M. G., Rowland, M., Piot, B., Guo, D., Calandriello, D., Valko, M., and Munos, R · 2023
Earlier work this paper cites.
Nepotistically trained generative-ai models collapse, 2023
Bohacek, M. and Farid, H · 2023
Earlier work this paper cites.
Large language models suffer from their own output: An analysis of the self-consuming training loop, 2023
Briesch, M., Sobania, D., and Rothlauf, F · 2023
Earlier work this paper cites.
Deft: Data efficient fine-tuning for large language models via unsupervised core-set selection
Das, D. and Khetan, V · 2023
Earlier work this paper cites.
Databricks dolly-15k, 2023
Databricks · 2023
Earlier work this paper cites.
Parameter-efficient fine-tuning of large-scale pre-trained language models
Ding, N., Qin, Y., Yang, G., Wei, F., Yang, Z., Su, Y., Hu, S., Chen, Y., Chan, C.-M., Chen, W., et al · 2023
Earlier work this paper cites.
Mods: Model-oriented data selection for instruction tuning
Du, Q., Zong, C., and Zhang, J · 2023
Earlier work this paper cites.
Citing: Large language models create curriculum for instruction tuning, 2023
Feng, T., Wang, Z., and Sun, J · 2023
Earlier work this paper cites.
Reinforced self-training (rest) for language modeling, 2023
Gulcehre, C., Paine, T. L., Srinivasan, S., Konyushkova, K., Weerts, L., Sharma, A., Siddhant, A., Ahern, A., Wang, M., Gu, C., Macherey, W., Doucet, A., Firat, O., and de Freitas, N · 2023
Earlier work this paper cites.
The curious decline of linguistic diversity: Training language models on synthetic text, 2023
Guo, Y., Shang, G., Vazirgiannis, M., and Clavel, C · 2023
Earlier work this paper cites.
Will large-scale generative models corrupt future datasets?
Hataya, R., Bao, H., and Arai, H · 2023
Earlier work this paper cites.
Preserving pre-trained features helps calibrate fine-tuned language models
He, G., Chen, J., and Zhu, J · 2023
Cited alongside, same era.
Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., de las Casas, D., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., Lavaud, L. R., Lachaux, M.-A., Stock, P., Scao, T. L., Lavril, T., Wang, T., Lacroix, T., and Sayed, W. E · 2023
Cited alongside, same era.
Openassistant conversations - democratizing large language model alignment
Köpf, A., Kilcher, Y., von Rütte, D., Anagnostidis, S., Tam, Z. R., Stevens, K., Barhoum, A., Nguyen, D., Stanley, O., Nagyfi, R., ES, S., Suri, S., Glushkov, D., Dantuluri, A., Maguire, A., Schuhmann, C., Nguyen, H., and Mattick, A · 2023
Cited alongside, same era.
Kung, P.-N., Yin, F., Wu, D., Chang, K.-W., and Peng, N · 2023
Cited alongside, same era.
Making language models better reasoners with step-aware verifier
Get more for less: Principled data selection for warming up fine-tuning in llms
Kang, F., Just, H. A., Sun, Y., Jahagirdar, H., Zhang, Y., Du, R., Sahu, A. K., and Jia, R · 2024
Later among the works it cites.
Understanding catastrophic forgetting in language models via implicit inference
Kotha, S., Springer, J. M., and Raghunathan, A · 2024
Later among the works it cites.
Tulu 3: Pushing frontiers in open language model post-training, 2024
Lambert, N., Morrison, J., Pyatkin, V., Huang, S., Ivison, H., Brahman, F., Miranda, L. J. V., Liu, A., Dziri, N., Lyu, S., Gu, Y., Malik, S., Graf, V., Hwang, J. D., Yang, J., Bras, R. L., Tafjord, O., Wilhelm, C., Soldaini, L., Smith, N. A., Wang, Y., Dasigi, P., and Hajishirzi, H · 2024
Later among the works it cites.
Instruction tuning with human curriculum, 2024
Lee, B. W., Cho, H., and Yoo, K. M · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Li, Y., Lin, Z., Zhang, S., Fu, Q., Chen, B., Lou, J.-G., and Chen, W · 2023
Cited alongside, same era.
Let’s verify step by step, 2023
Lightman, H., Kosaraju, V., Burda, Y., Edwards, H., Baker, B., Lee, T., Leike, J., Schulman, J., Sutskever, I., and Cobbe, K · 2023
Cited alongside, same era.
The flan collection: Designing data and methods for effective instruction tuning, 2023
Longpre, S., Hou, L., Vu, T., Webson, A., Chung, H. W., Tay, Y., Zhou, D., Le, Q. V., Zoph, B., Wei, J., and Roberts, A · 2023
Cited alongside, same era.
Deep learning on a data diet: Finding important examples early in training, 2023
Paul, M., Ganguli, S., and Dziugaite, G. K · 2023
Cited alongside, same era.
Direct preference optimization: Your language model is secretly a reward model, 2023
Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C. D., and Finn, C · 2023
Cited alongside, same era.
Offline reinforcement learning with on-policy q-function regularization, 2023
Shi, L., Dadashi, R., Chi, Y., Castro, P. S., and Geist, M · 2023
Cited alongside, same era.
The curse of recursion: Training on generated data makes models forget
Shumailov, I., Shumaylov, Z., Zhao, Y., Gal, Y., Papernot, N., and Anderson, R · 2023
Cited alongside, same era.
Principle-driven self-alignment of language models from scratch with minimal human supervision
Sun, Z., Shen, Y., Zhou, Q., Zhang, H., Chen, Z., Cox, D., Yang, Y., and Gan, C · 2023
Cited alongside, same era.
Li, M., Chen, L., Chen, J., He, S., Gu, J., and Zhou, T · 2024
Later among the works it cites.
From quantity to quality: Boosting LLM performance with self-guided data selection for instruction tuning
Li, M., Zhang, Y., Li, Z., Chen, J., Chen, L., Cheng, N., Wang, J., Zhou, T., and Xiao, J · 2024
Later among the works it cites.
Improve mathematical reasoning in language models by automated process supervision, 2024
Luo, L., Liu, Y., Liu, R., Phatale, S., Guo, M., Lara, H., Li, Y., Shu, L., Zhu, Y., Meng, L., Sun, J., and Rastogi, A · 2024
Later among the works it cites.
Mekala, D., Nguyen, A., and Shang, J · 2024
Later among the works it cites.
Aligning codellms with direct preference optimization, 2024
Miao, Y., Gao, B., Quan, S., Lin, J., Zan, D., Liu, J., Yang, J., Liu, T., and Deng, Z · 2024
Later among the works it cites.
Codestral-22b-v0.1
MistralAI · 2024
Later among the works it cites.
Mistral-small-instruct-2409
MistralAI · 2024
Later among the works it cites.
Codestral-22b-v0.1, 2024c
MistralAI · 2024
Later among the works it cites.
Does writing with language models reduce content diversity?
Padmakumar, V. and He, H · 2024
Later among the works it cites.
Scalebio: Scalable bilevel optimization for llm data reweighting, 2024
Pan, R., Zhang, J., Pan, X., Pi, R., Wang, X., and Zhang, T · 2024
Later among the works it cites.
Selectllm: Can llms select important instructions to annotate?
Parkar, R. S., Kim, J., Park, J. I., and Kang, D · 2024
Later among the works it cites.
Qin, Y., Yang, Y., Guo, P., Li, G., Shao, H., Shi, Y., Xu, Z., Gu, Y., Li, K., and Sun, X · 2024
Later among the works it cites.
Rl on incorrect synthetic data scales the efficiency of llm math reasoning by eight-fold, 2024
Setlur, A., Garg, S., Geng, X., Garg, N., Smith, V., and Kumar, A · 2024
Later among the works it cites.
Ai models collapse when trained on recursively generated data
Shumailov, I., Shumaylov, Z., Zhao, Y., Papernot, N., Anderson, R. J., and Gal, Y · 2024
Later among the works it cites.
Preference fine-tuning of llms should leverage suboptimal, on-policy data
Tajwar, F., Singh, A., Sharma, A., Rafailov, R., Schneider, J., Xie, T., Ermon, S., Finn, C., and Kumar, A · 2024
Later among the works it cites.
Curriculum learning with quality-driven data selection, 2024
Wu, B., Meng, F., and Chen, L · 2024
Later among the works it cites.
LESS: Selecting influential data for targeted instruction tuning
Xia, M., Malladi, S., Gururangan, S., Arora, S., and Chen, D · 2024
Later among the works it cites.
Iterative preference learning from human feedback: Bridging theory and practice for RLHF under KL-constraint
Xiong, W., Dong, H., Ye, C., Wang, Z., Zhong, H., Ji, H., Jiang, N., and Zhang, T · 2024
Later among the works it cites.
Compute-constrained data selection, 2024
Yin, J. O. and Rush, A. M · 2024
Later among the works it cites.
Metamath: Bootstrap your own mathematical questions for large language models
Yu, L., Jiang, W., Shi, H., YU, J., Liu, Z., Zhang, Y., Kwok, J., Li, Z., Weller, A., and Liu, W · 2024
Later among the works it cites.
Self-play fine-tuning of diffusion models for text-to-image generation
Yuan, H., Chen, Z., Ji, K., and Gu, Q · 2024
Later among the works it cites.
Automatic instruction evolving for large language models
Zeng, W., Xu, C., Zhao, Y., Lou, J.-G., and Chen, W · 2024
Later among the works it cites.
LMSYS-chat-1m: A large-scale real-world LLM conversation dataset
Zheng, L., Chiang, W.-L., Sheng, Y., Li, T., Zhuang, S., Wu, Z., Zhuang, Y., Li, Z., Lin, Z., Xing, E., Gonzalez, J. E., Stoica, I., and Zhang, H · 2024
Later among the works it cites.
Zhong, Y., Liu, S., Chen, J., Hu, J., Zhu, Y., Liu, X., Jin, X., and Zhang, H · 2024
Later among the works it cites.
Dai, Q., Zhang, D., Ma, J. W., and Peng, H · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025
DeepSeek-AI, Guo, D., Yang, D., Zhang, H., Song, J., Zhang, R., Xu, R., Zhu, Q., Ma, S., Wang, P., Bi, X., Zhang, X., Yu, X., Wu, Y., Wu, Z. F., Gou, Z., Shao, Z., Li, Z., Gao, Z., Liu, A., Xue, B., Wang, B., Wu, B., Feng, B., Lu, C., Zhao, C., Deng, C., Zhang, C., Ruan, C., Dai, D., Chen, D., Ji, D., Li, E., Lin, F., Dai, F., Luo, F., Hao, G., Chen, G., Li, G., Zhang, H., Bao, H., Xu, H., Wang, H., Ding, H., Xin, H., Gao, H., Qu, H., Li, H., Guo, J., Li, J., Wang, J., Chen, J., Yuan, J., Qiu, J., Li, J., Cai, J. L., Ni, J., Liang, J., Chen, J., Dong, K., Hu, K., Gao, K., Guan, K., Huang, K., Yu, K., Wang, L., Zhang, L., Zhao, L., Wang, L., Zhang, L., Xu, L., Xia, L., Zhang, M., Zhang, M., Tang, M., Li, M., Wang, M., Li, M., Tian, N., Huang, P., Zhang, P., Wang, Q., Chen, Q., Du, Q., Ge, R., Zhang, R., Pan, R., Wang, R., Chen, R. J., Jin, R. L., Chen, R., Lu, S., Zhou, S., Chen, S., Ye, S., Wang, S., Yu, S., Zhou, S., Pan, S., Li, S. S., Zhou, S., Wu, S., Ye, S., Yun, T., Pei, T., Sun, T., Wang, T., Zeng, W., Zhao, W., Liu, W., Liang, W., Gao, W., Yu, W., Zhang, W., Xiao, W. L., An, W., Liu, X., Wang, X., Chen, X., Nie, X., Cheng, X., Liu, X., Xie, X., Liu, X., Yang, X., Li, X., Su, X., Lin, X., Li, X. Q., Jin, X., Shen, X., Chen, X., Sun, X., Wang, X., Song, X., Zhou, X., Wang, X., Shan, X., Li, Y. K., Wang, Y. Q., Wei, Y. X., Zhang, Y., Xu, Y., Li, Y., Zhao, Y., Sun, Y., Wang, Y., Yu, Y., Zhang, Y., Shi, Y., Xiong, Y., He, Y., Piao, Y., Wang, Y., Tan, Y., Ma, Y., Liu, Y., Guo, Y., Ou, Y., Wang, Y., Gong, Y., Zou, Y., He, Y., Xiong, Y., Luo, Y., You, Y., Liu, Y., Zhou, Y., Zhu, Y. X., Xu, Y., Huang, Y., Li, Y., Zheng, Y., Zhu, Y., Ma, Y., Tang, Y., Zha, Y., Yan, Y., Ren, Z. Z., Ren, Z., Sha, Z., Fu, Z., Xu, Z., Xie, Z., Zhang, Z., Hao, Z., Ma, Z., Yan, Z., Wu, Z., Gu, Z., Zhu, Z., Liu, Z., Li, Z., Xie, Z., Song, Z., Pan, Z., Huang, Z., Xu, Z., Zhang, Z., and Zhang, Z · 2025
Closest in time.
Self-boosting large language models with synthetic preference data
Dong, Q., Dong, L., Zhang, X., Sui, Z., and Wei, F · 2025
Closest in time.
Open r1: A fully open reproduction of deepseek-r1, January 2025
Face, H · 2025
Closest in time.
Large-scale data selection for instruction tuning, 2025
Ivison, H., Zhang, M., Brahman, F., Koh, P. W., and Dasigi, P · 2025
Closest in time.
Nv-embed: Improved techniques for training llms as generalist embedding models, 2025
Lee, C., Roy, R., Xu, M., Raiman, J., Shoeybi, M., Catanzaro, B., and Ping, W · 2025
Closest in time.
Luo, Y., Yang, Z., Meng, F., Li, Y., Zhou, J., and Zhang, Y · 2025
Closest in time.
OLMo, T., Walsh, P., Soldaini, L., Groeneveld, D., Lo, K., Arora, S., Bhagia, A., Gu, Y., Huang, S., Jordan, M., Lambert, N., Schwenk, D., Tafjord, O., Anderson, T., Atkinson, D., Brahman, F., Clark, C., Dasigi, P., Dziri, N., Guerquin, M., Ivison, H., Koh, P. W., Liu, J., Malik, S., Merrill, W., Miranda, L. J. V., Morrison, J., Murray, T., Nam, C., Pyatkin, V., Rangapur, A., Schmitz, M., Skjonsberg, S., Wadden, D., Wilhelm, C., Wilson, M., Zettlemoyer, L., Farhadi, A., Smith, N. A., and Hajishirzi, H · 2025
Closest in time.
Open Thoughts, January 2025
Team, O. T · 2025
Closest in time.
Vicuna llm: An open-source chatbot developed by fine-tuning the llama model on user-shared conversations, achieving performance comparable to other advanced chatbots
Team, V. D · 2025
Closest in time.
Learning to reason under off-policy guidance, 2025
Yan, J., Li, Y., Hu, Z., Wang, Z., Cui, G., Qu, X., Cheng, Y., and Zhang, Y · 2025
Closest in time.