Fetching the paper…
Reading the bibliography…
Knowledge distillation (KD) is a simple and successful method to transfer knowledge from a teacher to a student model solely based on functional activity.
Functional Regularisation for Continual Learning with Gaussian Processes
Michalis K Titsias, Jonathan Schwarz, Alexander G. de G. Matthews, Razvan Pascanu, and Yee Whye Teh · 1901
Earlier work this paper cites.
Benchmarking neural network robustness to common corruptions and perturbations, mar 2019
Dan Hendrycks and Thomas Dietterich · 1903
Earlier work this paper cites.
MNIST-C: A Robustness Benchmark for Computer Vision
Norman Mu and Justin Gilmer · 1906
Earlier work this paper cites.
Xinyu Zhang, Qiang Wang, Jian Zhang, and Zhao Zhong · 1912
Earlier work this paper cites.
A simple way to make neural networks robust against diverse image corruptions
Evgenia Rusak, Lukas Schott, Roland S. Zimmermann, Julian Bitterwolf, Oliver Bringmann, Matthias Bethge, and Wieland Brendel · 2001
Earlier work this paper cites.
Hieu Pham, Zihang Dai, Qizhe Xie, and Quoc V. Le · 2003
Earlier work this paper cites.
Model compression
Cristian Bucilǎ, Rich Caruana, and Alexandra Niculescu-Mizil · 2006
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2010
Earlier work this paper cites.
Very Deep VAEs Generalize Autoregressive Models and Can Outperform Them on Images
Rewon Child · 2011
Earlier work this paper cites.
The mnist database of handwritten digit images for machine learning research
Li Deng · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Training data-efficient image transformers & distillation through attention
Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Hervé Jégou · 2012
Earlier work this paper cites.
Knowledge Distillation Thrives on Data Augmentation
Huan Wang, Suhas Lohit, Michael Jones, and Yun Fu · 2012
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization
Diederik P. Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P. Kingma and Max Welling · 2014
Earlier work this paper cites.
Distilling the Knowledge in a Neural Network
Geoffrey Hinton and Jeff Dean · 2015
Earlier work this paper cites.
Max Jaderberg, Karen Simonyan, Andrew Zisserman, and Koray Kavukcuoglu · 2015
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli · 2015
Earlier work this paper cites.
Deep Residual Learning for Image Recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Representational distance learning for deep neural networks
Patrick McClure and Nikolaus Kriegeskorte · 2016
Earlier work this paper cites.
Rethinking the inception architecture for computer vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna · 2016
Earlier work this paper cites.
A rotation and a translation suffice: Fooling cnns with simple transformations
Logan Engstrom, Brandon Tran, Dimitris Tsipras, Ludwig Schmidt, and Aleksander Madry · 2017
Cited alongside, same era.
Sgdr: Stochastic gradient descent with warm restarts, 2017
Ilya Loshchilov and Frank Hutter · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Paying more attention to attention: Improving the performance of convolutional neural networks via attention transfer
Sergey Zagoruyko and Nikos Komodakis · 2017
Cited alongside, same era.
mixup: Beyond Empirical Risk Minimization
Hongyi Zhang, Moustapha Cisse, Yann N. Dauphin, and David Lopez-Paz · 2017
Cited alongside, same era.
Maximum-entropy adversarial data augmentation for improved generalization and robustness
Long Zhao, Ting Liu, Xi Peng, and Dimitris Metaxas · 2020
Later among the works it cites.
Knowledge distillation: A good teacher is patient and consistent
Lucas Beyer, Xiaohua Zhai, Amélie Royer, Larisa Markeeva, Rohan Anil, and Alexander Kolesnikov · 2021
Later among the works it cites.
PoseAug: A Differentiable Pose Augmentation Framework for 3D Human Pose Estimation
Kehong Gong, Jianfeng Zhang, and Jiashi Feng · 2021
Later among the works it cites.
The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution Generalization
Dan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath, Frank Wang, Evan Dorundo, Rahul Desai, Tyler Zhu, Samyak Parajuli, Mike Guo, Dawn Song, Jacob Steinhardt, and Justin Gilmer · 2021
Later among the works it cites.
Trivialaugment: Tuning-free yet state-of-the-art data augmentation
Samuel G Müller and Frank Hutter · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zhun Zhong, Liang Zheng, Guoliang Kang, Shaozi Li, and Yi Yang · 2017
Cited alongside, same era.
Deep learning using rectified linear units (relu)
Abien Fred Agarap · 2018
Cited alongside, same era.
AutoAugment: Learning Augmentation Policies from Data
Ekin D. Cubuk, Barret Zoph, Dandelion Mane, Vijay Vasudevan, and Quoc V. Le · 2018
Cited alongside, same era.
Robert Geirhos, Claudio Michaelis, Felix A. Wichmann, Patricia Rubisch, Matthias Bethge, and Wieland Brendel · 2018
Cited alongside, same era.
Generalizing to Unseen Domains via Adversarial Data Augmentation
Riccardo Volpi, John Duchi, Hongseok Namkoong, Vittorio Murino, Ozan Sener, and Silvio Savarese · 2018
Cited alongside, same era.
ADA: Adversarial data augmentation for object detection
Sima Behpour, Kris M. Kitani, and Brian D. Ziebart · 2019
Cited alongside, same era.
Measuring and regularizing networks in function space
Ari S Benjamin, David Rolnick, and Konrad P Kording · 2019
Cited alongside, same era.
Later among the works it cites.
MATE-KD: Masked adversarial text, a companion to knowledge distillation
Ahmad Rashid, Vasileios Lioutas, and Mehdi Rezagholizadeh · 2021
Later among the works it cites.
MLP-Mixer: An all-MLP Architecture for Vision
Ilya Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai, Thomas Unterthiner, Jessica Yung, Andreas Steiner, Daniel Keysers, Jakob Uszkoreit, Mario Lucic, and Alexey Dosovitskiy · 2021
Later among the works it cites.
Data Augmentation Generative Adversarial Networks
Anthreas Antoniou, Amos Storkey, and Harrison Edwards · 2022
Later among the works it cites.
DearKD: Data-Efficient Early Knowledge Distillation for Vision Transformers
Xianing Chen, Qiong Cao, Yujie Zhong, Jing Zhang, Shenghua Gao, and Dacheng Tao · 2022
Later among the works it cites.
8-bit optimizers via block-wise quantization
Tim Dettmers, Mike Lewis, Sam Shleifer, and Luke Zettlemoyer · 2022
Later among the works it cites.
CILDA: Contrastive Data Augmentation using Intermediate Layer Knowledge Distillation
Md Akmal Haidar, Mehdi Rezagholizadeh, Abbas Ghaddar, Khalil Bibi, Philippe Langlais, and Pascal Poupart · 2022
Later among the works it cites.
Autoregressive image generation using residual quantization
Doyup Lee, Chiheon Kim, Saehoon Kim, Minsu Cho, and Wook-Shin Han · 2022
Later among the works it cites.
Can Functional Transfer Methods Capture Simple Inductive Biases?
Arne Nix, Suhas Shrinivasan, Edgar Y Walker, and Fabian Sinz · 2022
Later among the works it cites.
Hierarchical Text-Conditional Image Generation with CLIP Latents, April 2022
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen · 2022
Later among the works it cites.
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer · 2022
Later among the works it cites.
Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding, May 2022
Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, S. Sara Mahdavi, Rapha Gontijo Lopes, Tim Salimans, Jonathan Ho, David J. Fleet, and Mohammad Norouzi · 2022
Later among the works it cites.
Adversarial Data Augmentation for Task-Specific Knowledge Distillation of Pre-trained Transformers
Minjia Zhang, Niranjan Uma Naresh, and Yuxiong He · 2022
Later among the works it cites.
Leveling Down in Computer Vision: Pareto Inefficiencies in Fair Deep Classifiers
Dominik Zietlow, Michael Lohaus, Guha Balakrishnan, Matthäus Kleindessner, Francesco Locatello, Bernhard Schölkopf, and Chris Russell · 2022
Later among the works it cites.
Symbolic discovery of optimization algorithms, 2023
Xiangning Chen, Chen Liang, Da Huang, Esteban Real, Kaiyuan Wang, Yao Liu, Hieu Pham, Xuanyi Dong, Thang Luong, Cho-Jui Hsieh, Yifeng Lu, and Quoc V. Le · 2023
Closest in time.
Scaling vision transformers to 22 billion parameters, 2023
Mostafa Dehghani, Josip Djolonga, Basil Mustafa, Piotr Padlewski, Jonathan Heek, Justin Gilmer, Andreas Steiner, Mathilde Caron, Robert Geirhos, Ibrahim Alabdulmohsin, Rodolphe Jenatton, Lucas Beyer, Michael Tschannen, Anurag Arnab, Xiao Wang, Carlos Riquelme, Matthias Minderer, Joan Puigcerver, Utku Evci, Manoj Kumar, Sjoerd van Steenkiste, Gamaleldin F. Elsayed, Aravindh Mahendran, Fisher Yu, Avital Oliver, Fantine Huot, Jasmijn Bastings, Mark Patrick Collier, Alexey Gritsenko, Vighnesh Birodkar, Cristina Vasconcelos, Yi Tay, Thomas Mensink, Alexander Kolesnikov, Filip Pavetić, Dustin Tran, Thomas Kipf, Mario Lučić, Xiaohua Zhai, Daniel Keysers, Jeremiah Harmsen, and Neil Houlsby · 2023
Closest in time.
Dinov2: Learning robust visual features without supervision, 2023
Maxime Oquab, Timothée Darcet, Théo Moutakanni, Huy Vo, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, Mahmoud Assran, Nicolas Ballas, Wojciech Galuba, Russell Howes, Po-Yao Huang, Shang-Wen Li, Ishan Misra, Michael Rabbat, Vasu Sharma, Gabriel Synnaeve, Hu Xu, Hervé Jegou, Julien Mairal, Patrick Labatut, Armand Joulin, and Piotr Bojanowski · 2023
Closest in time.