Fetching the paper…
Reading the bibliography…
In the past year, distillation has seen a renewed prominence in large language model (LLM) pretraining, exemplified by the Llama-3.2 and Gemma model families.
Improved knowledge distillation via teacher assistant, 2019
Seyed-Iman Mirzadeh, Mehrdad Farajtabar, Ang Li, Nir Levine, Akihiro Matsukawa, and Hassan Ghasemzadeh · 1902
Earlier work this paper cites.
Analysing mathematical reasoning abilities of neural models, 2019
David Saxton, Edward Grefenstette, Felix Hill, and Pushmeet Kohli · 1904
Earlier work this paper cites.
On the efficacy of knowledge distillation, 2019
Jang Hyun Cho and Bharath Hariharan · 1910
Earlier work this paper cites.
Self-distillation amplifies regularization in hilbert space, 2020
Hossein Mobahi, Mehrdad Farajtabar, and Peter L. Bartlett · 2002
Earlier work this paper cites.
Model compression
Cristian Buciluǎ, Rich Caruana, and Alexandru Niculescu-Mizil · 2006
Earlier work this paper cites.
Do deep nets really need to be deep?
Jimmy Ba and Rich Caruana · 2014
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean · 2015
Earlier work this paper cites.
Unifying distillation and privileged information
David Lopez-Paz, Léon Bottou, Bernhard Schölkopf, and Vladimir Vapnik · 2015
Earlier work this paper cites.
RACE: Large-scale ReAding comprehension dataset from examinations
Guokun Lai, Qizhe Xie, Hanxiao Liu, Yiming Yang, and Eduard Hovy · 2017
Earlier work this paper cites.
Wang Ling, Dani Yogatama, Chris Dyer, and Phil Blunsom · 2017
Earlier work this paper cites.
Born again neural networks
Tommaso Furlanello, Zachary Lipton, Michael Tschannen, Laurent Itti, and Anima Anandkumar · 2018
Earlier work this paper cites.
DROP: A reading comprehension benchmark requiring discrete reasoning over paragraphs
Dheeru Dua, Yizhong Wang, Pradeep Dasigi, Gabriel Stanovsky, Sameer Singh, and Matt Gardner · 2019
Earlier work this paper cites.
Towards understanding knowledge distillation
Mary Phuong and Christoph Lampert · 2019
Earlier work this paper cites.
The Pile: An 800gb dataset of diverse text for language modeling
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser, and Connor Leahy · 2020
Earlier work this paper cites.
Improved knowledge distillation via teacher assistant
Seyed Iman Mirzadeh, Mehrdad Farajtabar, Ang Li, Nir Levine, Akihiro Matsukawa, and Hassan Ghasemzadeh · 2020
Earlier work this paper cites.
Program synthesis with large language models
Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, et al · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems, 2021
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman · 2021
Earlier work this paper cites.
Knowledge distillation: A survey
Jianping Gou, Baosheng Yu, Stephen J Maybank, and Dacheng Tao · 2021
Earlier work this paper cites.
Dynamic knowledge distillation for pre-trained language models, 2021
Lei Li, Yankai Lin, Shuhuai Ren, Peng Li, Jie Zhou, and Xu Sun · 2021
Earlier work this paper cites.
A statistical perspective on distillation
Aditya K Menon, Ankit Singh Rawat, Sashank Reddi, Seungyeon Kim, and Sanjiv Kumar · 2021
Earlier work this paper cites.
Knowledge distillation: A good teacher is patient and consistent, 2022
Lucas Beyer, Xiaohua Zhai, Amélie Royer, Larisa Markeeva, Rohan Anil, and Alexander Kolesnikov · 2022
Cited alongside, same era.
Github code dataset, 2022
neogithub · 2022
Cited alongside, same era.
In-context learning and induction heads, 2022
Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Scott Johnston, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah · 2022
Cited alongside, same era.
Llemma: An open language model for mathematics, 2023
Zhangir Azerbayev, Hailey Schoelkopf, Keiran Paster, Marco Dos Santos, Stephen McAleer, Albert Q. Jiang, Jia Deng, Stella Biderman, and Sean Welleck · 2023
Cited alongside, same era.
Birth of a transformer: A memory viewpoint
Alberto Bietti, Vivien Cabannes, Diane Bouchacourt, Herve Jegou, and Leon Bottou · 2023
Rewarding progress: Scaling automated process verifiers for llm reasoning, 2024
Amrith Setlur, Chirag Nagpal, Adam Fisch, Xinyang Geng, Jacob Eisenstein, Rishabh Agarwal, Alekh Agarwal, Jonathan Berant, and Aviral Kumar · 2024
Later among the works it cites.
Scaling llm test-time compute optimally can be more effective than scaling model parameters, 2024
Charlie Snell, Jaehoon Lee, Kelvin Xu, and Aviral Kumar · 2024
Later among the works it cites.
Alphaevolve: A gemini-powered coding agent for designing advanced algorithms, 2025
AlphaEvolve · 2025
Closest in time.
Distillation scaling laws, 2025
Dan Busbridge, Amitis Shidani, Floris Weers, Jason Ramapuram, Etai Littwin, and Russ Webb · 2025
Closest in time.
Why knowledge distillation works in generative models: A minimal working explanation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Hunter Lightman, Vineet Kosaraju, Yura Burda, Harri Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe · 2023
Cited alongside, same era.
Knowledge distillation performs partial variance reduction
Mher Safaryan, Alexandra Peste, and Dan Alistarh · 2023
Cited alongside, same era.
Lifting the curse of capacity gap in distilling language models, 2023
Chen Zhang, Yang Yang, Jiahao Liu, Jingang Wang, Yunsen Xian, Benyou Wang, and Dawei Song · 2023
Cited alongside, same era.
On-policy distillation of language models: Learning from self-generated mistakes, 2024
Rishabh Agarwal, Nino Vieillard, Yongchao Zhou, Piotr Stanczyk, Sabela Ramos, Matthieu Geist, and Olivier Bachem · 2024
Cited alongside, same era.
Scaling test-time compute with open models, 2024
Edward Beeching, Lewis Tunstall, and Sasha Rush · 2024
Cited alongside, same era.
Alphamath almost zero: Process supervision without process, 2024
Guoxin Chen, Minpeng Liao, Chengxi Li, and Kai Fan · 2024
Cited alongside, same era.
Inference-aware fine-tuning for best-of-n sampling in large language models, 2024
Yinlam Chow, Guy Tennenholtz, Izzeddin Gur, Vincent Zhuang, Bo Dai, Sridhar Thiagarajan, Craig Boutilier, Rishabh Agarwal, Aviral Kumar, and Aleksandra Faust · 2024
Cited alongside, same era.
Sungmin Cha and Kyunghyun Cho · 2025
Closest in time.
Feng Chen, Allan Raventos, Nan Cheng, Surya Ganguli, and Shaul Druckmann · 2025
Closest in time.
Weight ensembling improves reasoning in language models, 2025
Xingyu Dang, Christina Baek, Kaiyue Wen, Zico Kolter, and Aditi Raghunathan · 2025
Closest in time.
Team Gemma, Aishwarya Kamath, Johan Ferret, Shreya Pathak, Nino Vieillard, Ramona Merhej, Sarah Perrin, Tatiana Matejovicova, Alexandre Ramé, Morgane Rivière, et al · 2025
Closest in time.
Multi-token prediction needs registers, 2025
Anastasios Gerontopoulos, Spyros Gidaris, and Nikos Komodakis · 2025
Closest in time.
Context-parametric inversion: Why instruction finetuning can worsen context reliance, 2025
Sachin Goyal, Christina Baek, J. Zico Kolter, and Aditi Raghunathan · 2025
Closest in time.
Miniplm: Knowledge distillation for pre-training language models, 2025
Yuxian Gu, Hao Zhou, Fandong Meng, Jie Zhou, and Minlie Huang · 2025
Closest in time.
Openthoughts: Data recipes for reasoning models, 2025
Etash Guha, Ryan Marten, et al · 2025
Closest in time.
Datacomp-lm: In search of the next generation of training sets for language models, 2025
Jeffrey Li, Alex Fang, et al · 2025
Closest in time.
Multi-agent verification: Scaling test-time compute with multiple verifiers, 2025
Shalev Lifshitz, Sheila A. McIlraith, and Yilun Du · 2025
Closest in time.
s1: Simple test-time scaling, 2025
Niklas Muennighoff, Zitong Yang, Weijia Shi, Xiang Lisa Li, Li Fei-Fei, Hannaneh Hajishirzi, Luke Zettlemoyer, Percy Liang, Emmanuel Candès, and Tatsunori Hashimoto · 2025
Closest in time.
Vaishnavh Nagarajan, Chen Henry Wu, Charles Ding, and Aditi Raghunathan · 2025
Closest in time.
Looking beyond the next token, 2025
Abitha Thankaraj, Yiding Jiang, J. Zico Kolter, and Yonatan Bisk · 2025
Closest in time.
An Yang, Anfeng Li, et al · 2025
Closest in time.
Naturalreasoning: Reasoning in the wild with 2.8m challenging questions, 2025
Weizhe Yuan, Jane Yu, Song Jiang, Karthik Padthe, Yang Li, Dong Wang, Ilia Kulikov, Kyunghyun Cho, Yuandong Tian, Jason E Weston, and Xian Li · 2025
Closest in time.
Yang Yue, Zhiqi Chen, Rui Lu, Andrew Zhao, Zhaokai Wang, Yang Yue, Shiji Song, and Gao Huang · 2025
Closest in time.