Fetching the paper…
Reading the bibliography…
Distillation is the task of replacing a complicated machine learning model with a simpler model that approximates the original [BCNM06,HVD15].
On the uniform convergence of relative frequencies of events to their probabilities
VN Vapnik and A Ya Chervonenkis · 1971
Earlier work this paper cites.
The method of ordered risk minimization, i
Vladimir Vapnik and A Ya Chervonenkis · 1974
Earlier work this paper cites.
Classification and regression trees
L Breiman, JH Friedman, R Olshen, and CJ Stone · 1984
Earlier work this paper cites.
A theory of the learnable
Leslie G Valiant · 1984
Earlier work this paper cites.
Computational limitations on learning from examples
Leonard Pitt and Leslie G Valiant · 1988
Earlier work this paper cites.
Learnability and the vapnik-chervonenkis dimension
Anselm Blumer, Andrzej Ehrenfeucht, David Haussler, and Manfred K Warmuth · 1989
Earlier work this paper cites.
Learning decision trees from random examples
Andrzej Ehrenfeucht and David Haussler · 1989
Earlier work this paper cites.
Word association norms, mutual information, and lexicography
Kenneth Church and Patrick Hanks · 1990
Earlier work this paper cites.
Learning decision trees using the fourier spectrum
Eyal Kushilevitz and Yishay Mansour · 1991
Earlier work this paper cites.
Rule induction through integrated symbolic and subsymbolic processing
Clayton McMillan, Michael C Mozer, and Paul Smolensky · 1991
Earlier work this paper cites.
Rule induction in a neural network through integrated symbolic and subsymbolic processing
Clayton McMillan · 1992
Earlier work this paper cites.
Extracting provably correct rules from artificial neural networks
Sebastian Thrun · 1993
Earlier work this paper cites.
Weakly learning dnf and characterizing statistical query learning using fourier analysis
Avrim Blum, Merrick Furst, Jeffrey Jackson, Michael Kearns, Yishay Mansour, and Steven Rudich · 1994
Earlier work this paper cites.
Using sampling and queries to extract rules from trained neural networks
Mark W Craven and Jude W Shavlik · 1994
Earlier work this paper cites.
An introduction to computational learning theory
Michael J Kearns and Umesh Vazirani · 1994
Earlier work this paper cites.
Survey and critique of techniques for extracting rules from trained artificial neural networks
Robert Andrews, Joachim Diederich, and Alan B Tickle · 1995
Earlier work this paper cites.
Extracting tree-structured representations of trained networks
Mark Craven and Jude Shavlik · 1995
Earlier work this paper cites.
The nature of statistical learning theory, 1995
Vladimir N Vapnik · 1995
Earlier work this paper cites.
Born again trees
Leo Breiman and Nong Shang · 1996
Earlier work this paper cites.
Adaptive versus nonadaptive attribute-efficient learning
Peter Damaschke · 1998
Earlier work this paper cites.
Computational aspects of parallel attribute-efficient learning
Peter Damaschke · 1998
Earlier work this paper cites.
WordNet: An electronic lexical database
Christiane Fellbaum · 1998
Earlier work this paper cites.
Neural network learning: Theoretical foundations
Martin Anthony and Peter L Bartlett · 1999
Earlier work this paper cites.
Exact learning when irrelevant variables abound
David Guijarro, Vıctor Lavın, and Vijay Raghavan · 1999
Earlier work this paper cites.
Decision tree approximations of boolean functions
Dinesh Mehta and Vijay Raghavan · 2002
Earlier work this paper cites.
Optimal two-stage algorithms for group testing problems
Annalisa De Bonis, Leszek Gasieniec, and Ugo Vaccaro · 2005
Earlier work this paper cites.
Model compression
Cristian Buciluǎ, Rich Caruana, and Alexandru Niculescu-Mizil · 2006
Earlier work this paper cites.
Seeing the forest through the trees: Learning a comprehensible model from an ensemble
Anneleen Van Assche and Hendrik Blockeel · 2007
Earlier work this paper cites.
Structure compilation: trading structure for features
Percy Liang, Hal Daumé III, and Dan Klein · 2008
Earlier work this paper cites.
Learnability, stability and uniform convergence
Shai Shalev-Shwartz, Ohad Shamir, Nathan Srebro, and Karthik Sridharan · 2010
Earlier work this paper cites.
Learning a classifier when the labeling is known
Shalev Ben-David and Shai Ben-David · 2011
Earlier work this paper cites.
One tree to explain them all
Ulf Johansson, Cecilia Sönströd, and Tuve Löfström · 2011
Earlier work this paper cites.
Access to unlabeled data can speed up prediction time
Ruth Urner, Shai Shalev-Shwartz, and Shai Ben-David · 2011
Earlier work this paper cites.
Linguistic regularities in continuous space word representations
Tomáš Mikolov, Wen-tau Yih, and Geoffrey Zweig · 2013
Earlier work this paper cites.
Convex optimization: Algorithms and complexity
Sébastien Bubeck · 2015
Earlier work this paper cites.
Deep learning with limited numerical precision
Suyog Gupta, Ankur Agrawal, Kailash Gopalakrishnan, and Pritish Narayanan · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean · 2015
Earlier work this paper cites.
Understanding intermediate layers using linear classifier probes
Guillaume Alain and Yoshua Bengio · 2016
Earlier work this paper cites.
A latent variable model approach to pmi-based word embeddings
Sanjeev Arora, Yuanzhi Li, Yingyu Liang, Tengyu Ma, and Andrej Risteski · 2016
Cited alongside, same era.
Exact learning of juntas from membership queries
Nader H Bshouty and Areej Costa · 2016
Cited alongside, same era.
Man is to computer programmer as woman is to homemaker? debiasing word embeddings
Tolga Bolukbasi, Kai-Wei Chang, James Y Zou, Venkatesh Saligrama, and Adam T Kalai · 2016
Cited alongside, same era.
” why should i trust you?” explaining the predictions of any classifier
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin · 2016
Cited alongside, same era.
Does string-based neural mt learn source syntax?
Xing Shi, Inkit Padhi, and Kevin Knight · 2016
Cited alongside, same era.
Hidden progress in deep learning: Sgd learns parities near the computational limit
Boaz Barak, Benjamin Edelman, Surbhi Goel, Sham Kakade, Eran Malach, and Cyril Zhang · 2022
Later among the works it cites.
High-dimensional asymptotics of feature learning: How one gradient step improves the representation
Jimmy Ba, Murat A Erdogdu, Taiji Suzuki, Zhichao Wang, Denny Wu, and Greg Yang · 2022
Later among the works it cites.
Properly learning decision trees in almost polynomial time
Guy Blanc, Jane Lange, Mingda Qiao, and Li-Yang Tan · 2022
Later among the works it cites.
Gulp: a prediction-based metric between representations
Enric Boix-Adsera, Hannah Lawrence, George Stepaniants, and Philippe Rigollet · 2022
Later among the works it cites.
Interpretable by design: Learning predictors by composing interpretable queries
Aditya Chattopadhyay, Stewart Slocum, Benjamin D Haeffele, Rene Vidal, and Donald Geman · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yichen Zhou and Giles Hooker · 2016
Cited alongside, same era.
A learning problem that is independent of the set theory zfc axioms
Shai Ben-David, Pavel Hrubes, Shay Moran, Amir Shpilka, and Amir Yehudayoff · 2017
Cited alongside, same era.
Interpretability via model extraction
Osbert Bastani, Carolyn Kim, and Hamsa Bastani · 2017
Cited alongside, same era.
Distilling a neural network into a soft decision tree
Nicholas Frosst and Geoffrey Hinton · 2017
Cited alongside, same era.
A genetic algorithm for interpretable model extraction from decision tree ensembles
Gilles Vandewiele, Kiani Lannoye, Olivier Janssens, Femke Ongenae, Filip De Turck, and Sofie Van Hoecke · 2017
Cited alongside, same era.
To prune, or not to prune: exploring the efficacy of pruning for model compression
Michael Zhu and Suyog Gupta · 2017
Cited alongside, same era.
What you can cram into a single vector: Probing sentence embeddings for linguistic properties
Alexis Conneau, German Kruszewski, Guillaume Lample, Loïc Barrault, and Marco Baroni · 2018
Cited alongside, same era.
Neural networks can learn representations with gradient descent
Alexandru Damian, Jason Lee, and Mahdi Soltanolkotabi · 2022
Later among the works it cites.
Nelson Elhage, Tristan Hume, Catherine Olsson, Nicholas Schiefer, Tom Henighan, Shauna Kravec, Zac Hatfield-Dodds, Robert Lasenby, Dawn Drain, Carol Chen, et al · 2022
Later among the works it cites.
A survey of quantization methods for efficient neural network inference
Amir Gholami, Sehoon Kim, Zhen Dong, Zhewei Yao, Michael W Mahoney, and Kurt Keutzer · 2022
Later among the works it cites.
Emergent world representations: Exploring a sequence model trained on a synthetic task
Kenneth Li, Aspen K Hopkins, David Bau, Fernanda Viégas, Hanspeter Pfister, and Martin Wattenberg · 2022
Later among the works it cites.
Neural networks efficiently learn low-dimensional representations with sgd
Alireza Mousavi-Hosseini, Sejun Park, Manuela Girotti, Ioannis Mitliagkas, and Murat A Erdogdu · 2022
Later among the works it cites.
Minh Pham, Minsu Cho, Ameya Joshi, and Chinmay Hegde · 2022
Later among the works it cites.
Interpretable machine learning: Fundamental principles and 10 grand challenges
Cynthia Rudin, Chaofan Chen, Zhi Chen, Haiyang Huang, Lesia Semenova, and Chudi Zhong · 2022
Later among the works it cites.
Linear adversarial concept erasure
Shauli Ravfogel, Michael Twiton, Yoav Goldberg, and Ryan D Cotterell · 2022
Later among the works it cites.
A generic approach for reproducible model distillation
Yunzhe Zhou, Peiru Xu, and Giles Hooker · 2022
Later among the works it cites.
Sgd learning on neural networks: leap complexity and saddle-to-saddle dynamics
Emmanuel Abbe, Enric Boix-Adsera, and Theodor Misiakiewicz · 2023
Later among the works it cites.
Physics of language models: Part 1, context-free grammar
Zeyuan Allen-Zhu and Yuanzhi Li · 2023
Later among the works it cites.
Physics of language models: Part 3.1, knowledge storage and extraction
Zeyuan Allen-Zhu and Yuanzhi Li · 2023
Later among the works it cites.
Sgd with large step sizes learns sparse features
Maksym Andriushchenko, Aditya Vardhan Varre, Loucas Pillaud-Vivien, and Nicolas Flammarion · 2023
Later among the works it cites.
On learning gaussian multi-index models with gradient flow
Alberto Bietti, Joan Bruna, and Loucas Pillaud-Vivien · 2023
Later among the works it cites.
Transformers learn through gradual rank increase
Enric Boix-Adsera, Etai Littwin, Emmanuel Abbe, Samy Bengio, and Joshua Susskind · 2023
Later among the works it cites.
Leace: Perfect linear concept erasure in closed form
Nora Belrose, David Schneider-Joseph, Shauli Ravfogel, Ryan Cotterell, Edward Raff, and Stella Biderman · 2023
Later among the works it cites.
Learning two-layer neural networks, one (giant) step at a time
Yatin Dandi, Florent Krzakala, Bruno Loureiro, Luca Pesce, and Ludovic Stephan · 2023
Later among the works it cites.
Pareto frontiers in neural feature learning: Data, compute, width, and luck
Benjamin L Edelman, Surbhi Goel, Sham Kakade, Eran Malach, and Cyril Zhang · 2023
Later among the works it cites.
Inspecting and editing knowledge representations in language models
Evan Hernandez, Belinda Z Li, and Jacob Andreas · 2023
Later among the works it cites.
Properly learning decision trees with queries is np-hard
Caleb Koch, Carmen Strassle, and Li-Yang Tan · 2023
Later among the works it cites.
Superpolynomial lower bounds for decision tree learning and testing
Caleb Koch, Carmen Strassle, and Li-Yang Tan · 2023
Later among the works it cites.
Laughing hyena distillery: Extracting compact recurrences from convolutions
Stefano Massaroli, Michael Poli, Daniel Y Fu, Hermann Kumbong, Rom N Parnichkun, Aman Timalsina, David W Romero, Quinn McIntyre, Beidi Chen, Atri Rudra, et al · 2023
Later among the works it cites.
Samuel Marks and Max Tegmark · 2023
Later among the works it cites.
Emergent linear representations in world models of self-supervised sequence models
Neel Nanda, Andrew Lee, and Martin Wattenberg · 2023
Later among the works it cites.
The linear representation hypothesis and the geometry of large language models
Kiho Park, Yo Joong Choe, and Victor Veitch · 2023
Later among the works it cites.
Log-linear guardedness and its implications
Shauli Ravfogel, Yoav Goldberg, and Ryan Cotterell · 2023
Later among the works it cites.
The truth is in there: Improving reasoning in language models with layer-selective rank reduction
Pratyusha Sharma, Jordan T Ash, and Dipendra Misra · 2023
Later among the works it cites.
Linear representations of sentiment in large language models
Curt Tigges, Oskar John Hollinsworth, Atticus Geiger, and Neel Nanda · 2023
Later among the works it cites.
Concept algebra for score-based conditional model
Zihao Wang, Lin Gui, Jeffrey Negrea, and Victor Veitch · 2023
Later among the works it cites.
Transformers are uninterpretable with myopic methods: a case study with bounded dyck grammars
Kaiyue Wen, Yuchen Li, Bingbin Liu, and Andrej Risteski · 2023
Later among the works it cites.
Personal communication, 2024
Fan Chen and Alexander Rakhlin · 2024
Closest in time.
A survey on knowledge distillation of large language models
Xiaohan Xu, Ming Li, Chongyang Tao, Tao Shen, Reynold Cheng, Jinyang Li, Can Xu, Dacheng Tao, and Tianyi Zhou · 2024
Closest in time.