Fetching the paper…
Reading the bibliography…
Knowledge distillation leverages a teacher model to improve the training of a student model.
Efficient noise-tolerant learning from statistical queries
Michael Kearns · 1998
Earlier work this paper cites.
Speech & language processing
Dan Jurafsky · 2000
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Analysis of boolean functions
Ryan O’Donnell · 2014
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean · 2015
Earlier work this paper cites.
Sgdr: Stochastic gradient descent with warm restarts
Ilya Loshchilov and Frank Hutter · 2016
Earlier work this paper cites.
Exponentially vanishing sub-optimal local minima in multilayer neural networks
Daniel Soudry and Elad Hoffer · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Large scale distributed neural network training through online distillation
Rohan Anil, Gabriel Pereyra, Alexandre Passos, Róbert Ormándi, George E. Dahl, and Geoffrey E. Hinton · 2018
Earlier work this paper cites.
On the global convergence of gradient descent for over-parameterized models using optimal transport
Lénaïc Chizat and Francis R. Bach · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Theoretical insights into the optimization landscape of over-parameterized shallow neural networks
Mahdi Soltanolkotabi, Adel Javanmard, and Jason D Lee · 2018
Earlier work this paper cites.
On the efficacy of knowledge distillation
Jang Hyun Cho and Bharath Hariharan · 2019
Earlier work this paper cites.
Width provably matters in optimization for deep linear neural networks
Simon Du and Wei Hu · 2019
Earlier work this paper cites.
Linearized two-layers neural networks in high dimension
Behrooz Ghorbani, Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 2019
Earlier work this paper cites.
A structural probe for finding syntax in word representations
John Hewitt and Christopher D. Manning · 2019
Earlier work this paper cites.
Knowledge distillation via route constrained optimization
Xiao Jin, Baoyun Peng, Yichao Wu, Yu Liu, Jiaheng Liu, Ding Liang, Junjie Yan, and Xiaolin Hu · 2019
Earlier work this paper cites.
SGD on neural networks learns functions of increasing complexity
Dimitris Kalimeris, Gal Kaplun, Preetum Nakkiran, Benjamin L. Edelman, Tristan Yang, Boaz Barak, and Haofeng Zhang · 2019
Earlier work this paper cites.
Improved knowledge distillation via teacher assistant
Seyed Iman Mirzadeh, Mehrdad Farajtabar, Ang Li, Nir Levine, Akihiro Matsukawa, and H. Ghasemzadeh · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Earlier work this paper cites.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf · 2019
Earlier work this paper cites.
Bert rediscovers the classical nlp pipeline
Ian Tenney, Dipanjan Das, and Ellie Pavlick · 2019
Earlier work this paper cites.
Regularization matters: Generalization and optimization of neural nets vs their induced kernel
Colin Wei, Jason D Lee, Qiang Liu, and Tengyu Ma · 2019
Earlier work this paper cites.
Implicit bias of gradient descent for wide two-layer neural networks trained with the logistic loss
Lénaïc Chizat and Francis Bach · 2020
Earlier work this paper cites.
Infinite attention: Nngp and ntk for deep attention networks
Jiri Hron, Yasaman Bahri, Jascha Sohl-Dickstein, and Roman Novak · 2020
Earlier work this paper cites.
Self-distillation amplifies regularization in hilbert space
Hossein Mobahi, Mehrdad Farajtabar, and Peter Bartlett · 2020
Earlier work this paper cites.
Revisiting knowledge distillation via label smoothing regularization
Li Yuan, Francis EH Tay, Guilin Li, Tao Wang, and Jiashi Feng · 2020
Earlier work this paper cites.
Rethinking soft labels for knowledge distillation: A bias–variance tradeoff perspective
Helong Zhou, Liangchen Song, Jiajie Chen, Ye Zhou, Guoli Wang, Junsong Yuan, and Qian Zhang · 2020
Cited alongside, same era.
Knowledge distillation as semiparametric inference, 2021
Tri Dao, Govinda M Kamath, Vasilis Syrgkanis, and Lester Mackey · 2021
Cited alongside, same era.
Attention is not all you need: pure attention loses rank doubly exponentially with depth
Yihe Dong, Jean-Baptiste Cordonnier, and Andreas Loukas · 2021
Cited alongside, same era.
Dimension lower bounds for linear approaches to function approximation
Daniel Hsu · 2021
Cited alongside, same era.
Annealing knowledge distillation
A. Jafari, Mehdi Rezagholizadeh, Pranav Sharma, and A. Ghodsi · 2021
Cited alongside, same era.
Learning efficient vision transformers via fine-grained manifold distillation
Pareto frontiers in neural feature learning: Data, compute, width, and luck
Benjamin L Edelman, Surbhi Goel, Sham Kakade, Eran Malach, and Cyril Zhang · 2023
Later among the works it cites.
Textbooks are all you need
Suriya Gunasekar, Yi Zhang, Jyoti Aneja, Caio César Teodoro Mendes, Allie Del Giorno, Sivakanth Gopi, Mojan Javaheripi, Piero Kauffmann, Gustavo de Rosa, Olli Saarikivi, Adil Salim, Shital Shah, Harkirat Singh Behl, Xin Wang, Sébastien Bubeck, Ronen Eldan, Adam Tauman Kalai, Yin Tat Lee, and Yuanzhi Li · 2023
Later among the works it cites.
Supervised masked knowledge distillation for few-shot transformers
Han Lin, Guangxing Han, Jiawei Ma, Shiyuan Huang, Xudong Lin, and Shih-Fu Chang · 2023
Later among the works it cites.
Wizardmath: Empowering mathematical reasoning for large language models via reinforced evol-instruct
Haipeng Luo, Qingfeng Sun, Can Xu, Pu Zhao, Jianguang Lou, Chongyang Tao, Xiubo Geng, Qingwei Lin, Shifeng Chen, and Dongmei Zhang · 2023
Later among the works it cites.
Feature emergence via margin maximization: case studies in algebraic tasks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ding Jia, Kai Han, Yunhe Wang, Yehui Tang, Jianyuan Guo, Chao Zhang, and D. Tao · 2021
Cited alongside, same era.
A statistical perspective on distillation
Aditya K Menon, Ankit Singh Rawat, Sashank Reddi, Seungyeon Kim, and Sanjiv Kumar · 2021
Cited alongside, same era.
Follow your path: a progressive method for knowledge distillation, 2021
Wenxian Shi, Yuxuan Song, Hao Zhou, Bohan Li, and Lei Li · 2021
Cited alongside, same era.
Training data-efficient image transformers and distillation through attention
Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Herve Jegou · 2021
Cited alongside, same era.
The merged-staircase property: a necessary and nearly sufficient condition for sgd learning of sparse functions on two-layer neural networks
Emmanuel Abbe, Enric Boix-Adsera, and Theodor Misiakiewicz · 2022
Cited alongside, same era.
Hidden progress in deep learning: SGD learns parities near the computational limit
Boaz Barak, Benjamin L. Edelman, Surbhi Goel, Sham M. Kakade, Eran Malach, and Cyril Zhang · 2022
Cited alongside, same era.
Simplicity bias in transformers and their ability to learn sparse boolean functions
S. Bhattamishra, Arkil Patel, Varun Kanade, and Phil Blunsom · 2022
Cited alongside, same era.
Depen Morwani, Benjamin L. Edelman, Costin-Andrei Oncescu, Rosie Zhao, and Sham Kakade · 2023
Later among the works it cites.
Training dynamics of contextual n-grams in language models
Lucia Quirke, Lovis Heindrich, Wes Gurnee, and Neel Nanda · 2023
Later among the works it cites.
Alpaca: A strong, replicable instruction-following model
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B Hashimoto · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Later among the works it cites.
Metamath: Bootstrap your own mathematical questions for large language models
Longhui Yu, Weisen Jiang, Han Shi, Jincheng Yu, Zhengying Liu, Yu Zhang, James T Kwok, Zhenguo Li, Adrian Weller, and Weiyang Liu · 2023
Later among the works it cites.
Mammoth: Building math generalist models through hybrid instruction tuning
Xiang Yue, Xingwei Qu, Ge Zhang, Yao Fu, Wenhao Huang, Huan Sun, Yu Su, and Wenhu Chen · 2023
Later among the works it cites.
Do transformers parse while predicting the masked word?
Haoyu Zhao, A. Panigrahi, Rong Ge, and Sanjeev Arora · 2023
Later among the works it cites.
Judging llm-as-a-judge with mt-bench and chatbot arena, 2023
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric. P Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica · 2023
Later among the works it cites.
Provable advantage of curriculum learning on parity targets with mixed inputs
Emmanuel Abbe, Elisabetta Cornacchia, and Aryo Lotfi · 2024
Closest in time.
On-policy distillation of language models: Learning from self-generated mistakes
Rishabh Agarwal, Nino Vieillard, Yongchao Zhou, Piotr Stanczyk, Sabela Ramos Garea, Matthieu Geist, and Olivier Bachem · 2024
Closest in time.
In-context language learning: Architectures and algorithms
Ekin Akyürek, Bailin Wang, Yoon Kim, and Jacob Andreas · 2024
Closest in time.
A dynamical model of neural scaling laws
Blake Bordelon, Alexander B. Atanasov, and Cengiz Pehlevan · 2024
Closest in time.
SGD finds then tunes features in two-layer neural networks with near-optimal sample complexity: A case study in the XOR problem
Margalit Glasgow · 2024
Closest in time.
How do transformers fill in the blanks? a case study on matrix completion
Pulkit Gopalani, Ekdeep Singh Lubana, and Wei Hu · 2024
Closest in time.
Augmenting math word problems via iterative question composing
Haoxiong Liu and Andrew Chi-Chih Yao · 2024
Closest in time.
Best practices and lessons learned on synthetic data for language models
Ruibo Liu, Jerry Wei, Fangyu Liu, Chenglei Si, Yanzhe Zhang, Jinmeng Rao, Steven Zheng, Daiyi Peng, Diyi Yang, Denny Zhou, et al · 2024
Closest in time.
Orca-math: Unlocking the potential of slms in grade school math
Arindam Mitra, Hamed Khanpour, Corby Rosset, and Ahmed Awadallah · 2024
Closest in time.
On student-teacher deviations in distillation: does it pay to disobey?, 2024
Vaishnavh Nagarajan, Aditya Krishna Menon, Srinadh Bhojanapalli, Hossein Mobahi, and Sanjiv Kumar · 2024
Closest in time.
Understanding the gains from repeated self-distillation
Divyansh Pareek, Simon S. Du, and Sewoong Oh · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Machel Reid, Nikolay Savinov, Denis Teplyashin, Dmitry Lepikhin, Timothy Lillicrap, Jean-baptiste Alayrac, Radu Soricut, Angeliki Lazaridou, Orhan Firat, Julian Schrittwieser, et al · 2024
Closest in time.
Gemma: Open models based on gemini research and technology
Gemma Team, Thomas Mesnard, Cassidy Hardin, Robert Dadashi, Surya Bhupatiraju, Shreya Pathak, Laurent Sifre, Morgane Rivière, Mihir Sanjay Kale, Juliette Love, Pouya Tafti, Léonard Hussenot, Pier Giuseppe Sessa, Aakanksha Chowdhery, Adam Roberts, Aditya Barua, Alex Botev, Alex Castro-Ros, Ambrose Slone, Amélie Héliou, Andrea Tacchetti, Anna Bulanova, Antonia Paterson, Beth Tsai, Bobak Shahriari, Charline Le Lan, Christopher A. Choquette-Choo, Clément Crepy, Daniel Cer, Daphne Ippolito, David Reid, Elena Buchatskaya, Eric Ni, Eric Noland, Geng Yan, George Tucker, George-Christian Muraru, Grigory Rozhdestvenskiy, Henryk Michalewski, Ian Tenney, Ivan Grishchenko, Jacob Austin, James Keeling, Jane Labanowski, Jean-Baptiste Lespiau, Jeff Stanway, Jenny Brennan, Jeremy Chen, Johan Ferret, Justin Chiu, Justin Mao-Jones, Katherine Lee, Kathy Yu, Katie Millican, Lars Lowe Sjoesund, Lisa Lee, Lucas Dixon, Machel Reid, Maciej Mikuła, Mateo Wirth, Michael Sharman, Nikolai Chinaev, Nithum Thain, Olivier Bachem, Oscar Chang, Oscar Wahltinez, Paige Bailey, Paul Michel, Petko Yotov, Rahma Chaabouni, Ramona Comanescu, Reena Jana, Rohan Anil, Ross McIlroy, Ruibo Liu, Ryan Mullins, Samuel L Smith, Sebastian Borgeaud, Sertan Girgin, Sholto Douglas, Shree Pandya, Siamak Shakeri, Soham De, Ted Klimenko, Tom Hennigan, Vlad Feinberg, Wojciech Stokowiec, Yu hui Chen, Zafarali Ahmed, Zhitao Gong, Tris Warkentin, Ludovic Peran, Minh Giang, Clément Farabet, Oriol Vinyals, Jeff Dean, Koray Kavukcuoglu, Demis Hassabis, Zoubin Ghahramani, Douglas Eck, Joelle Barral, Fernando Pereira, Eli Collins, Armand Joulin, Noah Fiedel, Evan Senter, Alek Andreev, and Kathleen Kenealy · 2024
Closest in time.
Knowledge distillation based on transformed teacher matching
Kaixiang Zheng and En-Hui Yang · 2024
Closest in time.