Fetching the paper…
Reading the bibliography…
Distillation has shown remarkable success in transferring knowledge from a Large Language Model (LLM) teacher to a student LLM.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter, 2020
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf · 1910
Earlier work this paper cites.
On information and sufficiency
Solomon Kullback and Richard A Leibler · 1951
Earlier work this paper cites.
Mathematical methods of organizing and planning production
Leonid V Kantorovich · 1960
Earlier work this paper cites.
On measures of entropy and information
Alfréd Rényi · 1961
Earlier work this paper cites.
Binary codes capable of correcting deletions, insertions, and reversals
VI Levenshtein · 1966
Earlier work this paper cites.
A general method applicable to the search for similarities in the amino acid sequence of two proteins
Saul B Needleman and Christian D Wunsch · 1970
Earlier work this paper cites.
From English To Foreign Languages: Transferring Pre-trained Language Models
Ke Tran · 2002
Earlier work this paper cites.
Utf-8, a transformation format of iso 10646
François Yergeau · 2003
Earlier work this paper cites.
Model compression
Cristian Buciluǎ, Rich Caruana, and Alexandru Niculescu-Mizil · 2006
Earlier work this paper cites.
How Good is Your Tokenizer? On the Monolingual Performance of Multilingual Language Models
Phillip Rust, Jonas Pfeiffer, Ivan Vulić, Sebastian Ruder, and Iryna Gurevych · 2012
Earlier work this paper cites.
Distilling the Knowledge in a Neural Network, March 2015
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Neural Machine Translation of Rare Words with Subword Units
Rico Sennrich, Barry Haddow, and Alexandra Birch · 2016
Earlier work this paper cites.
Gradnorm: Gradient normalization for adaptive loss balancing in deep multitask networks
Zhao Chen, Vijay Badrinarayanan, Chen-Yu Lee, and Andrew Rabinovich · 2018
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord · 2018
Earlier work this paper cites.
SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing
Taku Kudo and John Richardson · 2018
Earlier work this paper cites.
BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions
Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova · 2019
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
Language Models are Unsupervised Multitask Learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Earlier work this paper cites.
HellaSwag: Can a machine really finish your sentence?
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi · 2019
Earlier work this paper cites.
PIQA: Reasoning about Physical Commonsense in Natural Language
Yonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao, and Yejin Choi · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman · 2021
Earlier work this paper cites.
Measuring mathematical problem solving with the MATH dataset
Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt · 2021
Earlier work this paper cites.
Fast Vocabulary Transfer for Language Model Compression
Leonidas Gee, Andrea Zugarini, Leonardo Rigutini, and Paolo Torroni · 2022
Cited alongside, same era.
LoRA: Low-rank adaptation of large language models
Edward J Hu, yelong shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen · 2022
Cited alongside, same era.
Benjamin Minixhofer, Fabian Paischer, and Navid Rekabsaz · 2022
Cited alongside, same era.
Charformer: Fast Character Transformers via Gradient-based Subword Tokenization
Yi Tay, Vinh Q. Tran, Sebastian Ruder, Jai Gupta, Hyung Won Chung, Dara Bahri, Zhen Qin, Simon Baumgartner, Cong Yu, and Donald Metzler · 2022
Cited alongside, same era.
ByT5: Towards a token-free future with pre-trained byte-to-byte models
Linting Xue, Aditya Barua, Noah Constant, Rami Al-Rfou, Sharan Narang, Mihir Kale, Adam Roberts, and Colin Raffel · 2022
DistiLLM: Towards streamlined distillation for large language models
Jongwoo Ko, Sungnyun Kim, Tianyi Chen, and Se-Young Yun · 2024
Later among the works it cites.
Pack of llms: Model fusion at test-time via perplexity optimization, 2024
Costas Mavromatis, Petros Karypis, and George Karypis · 2024
Later among the works it cites.
Zero-shot tokenizer transfer, 2024
Benjamin Minixhofer, Edoardo Maria Ponti, and Ivan Vulić · 2024
Later among the works it cites.
Byte Latent Transformer: Patches Scale Better Than Tokens, December 2024
Artidoro Pagnoni, Ram Pasunuru, Pedro Rodriguez, John Nguyen, Benjamin Muller, Margaret Li, Chunting Zhou, Lili Yu, Jason Weston, Luke Zettlemoyer, Gargi Ghosh, Mike Lewis, Ari Holtzman, and Srinivasan Iyer · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
FOCUS: Effective Embedding Initialization for Monolingual Specialization of Multilingual Models
Konstantin Dobler and Gerard de Melo · 2023
Cited alongside, same era.
Specializing smaller language models towards multi-step reasoning
Yao Fu, Hao Peng, Litu Ou, Ashish Sabharwal, and Tushar Khot · 2023
Cited alongside, same era.
Efficient Transformers with Dynamic Token Pooling
Piotr Nawrot, Jan Chorowski, Adrian Lancucki, and Edoardo Maria Ponti · 2023
Cited alongside, same era.
CombLM: Adapting black-box language models through small fine-tuned models
Aitor Ormazabal, Mikel Artetxe, and Eneko Agirre · 2023
Cited alongside, same era.
Llama 2: Open Foundation and Fine-Tuned Chat Models, 2023
Hugo Touvron, Louis Martin, Kevin Stone, et al · 2023
Cited alongside, same era.
MEGABYTE: Predicting Million-byte Sequences with Multiscale Transformers
Lili Yu, Daniel Simig, Colin Flaherty, Armen Aghajanyan, Luke Zettlemoyer, and Mike Lewis · 2023
Cited alongside, same era.
Agieval: A human-centric benchmark for evaluating foundation models, 2023
Wanjun Zhong, Ruixiang Cui, Yiduo Guo, Yaobo Liang, Shuai Lu, Yanlin Wang, Amin Saied, Weizhu Chen, and Nan Duan · 2023
Cited alongside, same era.
Buu Phan, Brandon Amos, Itai Gat, Marton Havasi, Matthew Muckley, and Karen Ullrich · 2024
Later among the works it cites.
How to compute the probability of a word
Tiago Pimentel and Clara Meister · 2024
Later among the works it cites.
The non-local model merging problem: Permutation symmetries and variance collapse, 2024
Ekansh Sharma, Daniel M. Roy, and Gintare Karolina Dziugaite · 2024
Later among the works it cites.
Learning to decode collaboratively with multiple language models
Zejiang Shen, Hunter Lang, Bailin Wang, Yoon Kim, and David Sontag · 2024
Later among the works it cites.
Openmathinstruct-2: Accelerating ai for math with massive open-source instruction data
Shubham Toshniwal, Wei Du, Ivan Moshkov, Branislav Kisacanin, Alexan Ayrapetyan, and Igor Gitman · 2024
Later among the works it cites.
Greed is All You Need: An Evaluation of Tokenizer Inference Methods, 2024
Omri Uzan, Craig W. Schmidt, Chris Tanner, and Yuval Pinter · 2024
Later among the works it cites.
From Language Models over Tokens to Language Models over Characters, December 2024
Tim Vieira, Ben LeBrun, Mario Giulianelli, Juan Luis Gastaldi, Brian DuSell, John Terilla, Timothy J. O’Donnell, and Ryan Cotterell · 2024
Later among the works it cites.
Knowledge fusion of large language models
Fanqi Wan, Xinting Huang, Deng Cai, Xiaojun Quan, Wei Bi, and Shuming Shi · 2024
Later among the works it cites.
A survey on knowledge distillation of large language models, 2024
Xiaohan Xu, Ming Li, Chongyang Tao, Tao Shen, Reynold Cheng, Jinyang Li, Can Xu, Dacheng Tao, and Tianyi Zhou · 2024
Later among the works it cites.
Dual-Space Knowledge Distillation for Large Language Models
Songming Zhang, Xue Zhang, Zengkui Sun, Yufeng Chen, and Jinan Xu · 2024
Later among the works it cites.
Why do llms attend to the first token?, 2025
Federico Barbero, Álvaro Arroyo, Xiangming Gu, Christos Perivolaropoulos, Michael Bronstein, Petar Veličković, and Razvan Pascanu · 2025
Closest in time.
Towards cross-tokenizer distillation: the universal logit distillation loss for LLMs
Nicolas Boizard, Kevin El Haddad, Celine Hudelot, and Pierre Colombo · 2025
Closest in time.
Distillation Scaling Laws, February 2025
Dan Busbridge, Amitis Shidani, Floris Weers, Jason Ramapuram, Etai Littwin, and Russ Webb · 2025
Closest in time.
Gemma 3 technical report, 2025
Gemma Team, Aishwarya Kamath, Johan Ferret, Shreya Pathak, et al · 2025
Closest in time.
Tulu 3: Pushing frontiers in open language model post-training, 2025
Nathan Lambert, Jacob Morrison, Valentina Pyatkin, Shengyi Huang, et al · 2025
Closest in time.
Mistral small 3.1, March 2025
Mistral AI · 2025
Closest in time.
Qwen2.5 technical report, 2025
Qwen, :, An Yang, Baosong Yang, Beichen Zhang, et al · 2025
Closest in time.
On Teacher Hacking in Language Model Distillation, February 2025
Daniil Tiapkin, Daniele Calandriello, Johan Ferret, Sarah Perrin, Nino Vieillard, Alexandre Ramé, and Mathieu Blondel · 2025
Closest in time.