Fetching the paper…
Reading the bibliography…
The rapid expansion of foundation pre-trained models and their fine-tuned counterparts has significantly contributed to the advancement of machine learning.
Extracting symbolic rules from trained neural network ensembles
Z.-H. Zhou, Y. Jiang, and S. Chen · 2003
Earlier work this paper cites.
Nec4.5: Neural ensemble based C4.5
Z.-H. Zhou and Y. Jiang · 2004
Earlier work this paper cites.
Scikit-learn: Machine learning in Python
F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay · 2011
Earlier work this paper cites.
From N to N+1: multiclass transfer incremental learning
I. Kuzborskij, F. Orabona, and B. Caputo · 2013
Earlier work this paper cites.
Distance-based image classification: Generalizing to new classes at near-zero cost
T. Mensink, J. Verbeek, F. Perronnin, and G. Csurka · 2013
Earlier work this paper cites.
How transferable are features in deep neural networks?
J. Yosinski, J. Clune, Y. Bengio, and H. Lipson · 2014
Earlier work this paper cites.
Distilling the knowledge in a neural network
G. E. Hinton, O. Vinyals, and J. Dean · 2015
Earlier work this paper cites.
Fitnets: Hints for thin deep nets
A. Romero, N. Ballas, S. E. Kahou, A. Chassang, C. Gatta, and Y. Bengio · 2015
Earlier work this paper cites.
Tensorflow: a system for large-scale machine learning
M. Abadi, P. Barham, J. Chen, Z. Chen, A. Davis, J. Dean, M. Devin, S. Ghemawat, G. Irving, and M. Isard · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Like what you like: Knowledge distill via neuron selectivity transfer
Z. Huang and N. Wang · 2017
Earlier work this paper cites.
Fast rates by transferring from auxiliary hypotheses
I. Kuzborskij and F. Orabona · 2017
Earlier work this paper cites.
A gift from knowledge distillation: Fast optimization, network minimization and transfer learning
J. Yim, D. Joo, J.-H. Bae, and J. Kim · 2017
Earlier work this paper cites.
Preference based adaptation for learning objectives
Y.-X. Ding and Z.-H. Zhou · 2018
Earlier work this paper cites.
Explicit inductive bias for transfer learning with convolutional networks
X. Li, Y. Grandvalet, and F. Davoine · 2018
Earlier work this paper cites.
Spectral normalization for generative adversarial networks
T. Miyato, T. Kataoka, M. Koyama, and Y. Yoshida · 2018
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
S. Benoit, D. Zachary, C. Soumith, G. Sam, P. Adam, M. Francisco, L. Adam, C. Gregory, L. Zeming, Y. Edward, D. Alban, T. Alykhan, K. Andreas, B. James, A. Luca, R. Martin, G. Natalia, C. Sasank, K. Trevor, F. Lu, and B. Junjie · 2019
Earlier work this paper cites.
Catastrophic forgetting meets negative transfer: Batch spectral shrinkage for safe transfer learning
X. Chen, S. Wang, B. Fu, M. Long, and J. Wang · 2019
Cited alongside, same era.
BERT: pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2019
Cited alongside, same era.
Parameter-efficient transfer learning for NLP
N. Houlsby, A. Giurgiu, S. Jastrzebski, B. Morrone, Q. de Laroussilhe, A. Gesmundo, M. Attariyan, and S. Gelly · 2019
Cited alongside, same era.
Delta: Deep learning transfer using feature map with attention for convolutional networks
X. Li, H. Xiong, H. Wang, Y. Rao, L. Liu, and J. Huan · 2019
Cited alongside, same era.
Relational knowledge distillation
W. Park, D. Kim, Y. Lu, and M. Cho · 2019
Cited alongside, same era.
Similarity-preserving knowledge distillation
F. Tung and G. Mori · 2019
Heterogeneous few-shot model rectification with semantic mapping
H.-J. Ye, D.-C. Zhan, Y. Jiang, and Z.-H. Zhou · 2021
Later among the works it cites.
Head2toe: Utilizing intermediate representations for better transfer learning
U. Evci, V. Dumoulin, H. Larochelle, and M. C. Mozer · 2022
Later among the works it cites.
Lora: Low-rank adaptation of large language models
E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen · 2022
Later among the works it cites.
Visual prompt tuning
M. Jia, L. Tang, B.-C. Chen, C. Cardie, S. J. Belongie, B. Hariharan, and S.-N. Lim · 2022
Later among the works it cites.
Convolutional bypasses are better vision transformer adapters
S. Jie and Z.-H. Deng · 2022
Later among the works it cites.
Scaling & shifting your features: A new baseline for efficient model tuning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Simpleshot: Revisiting nearest-neighbor classification for few-shot learning
Y. Wang, W.-L. Chao, K. Q. Weinberger, and L. van der Maaten · 2019
Cited alongside, same era.
Heterogeneous model reuse via optimizing multiparty multiclass margin
X.-Z. Wu, S. Liu, and Z.-H. Zhou · 2019
Cited alongside, same era.
Do we really need to access the source data? source hypothesis transfer for unsupervised domain adaptation
J. Liang, D. Hu, and J. Feng · 2020
Cited alongside, same era.
Adapterhub: A framework for adapting transformers
J. Pfeiffer, A. Rücklé, C. Poth, A. Kamath, I. Vulić, S. Ruder, K. Cho, and I. Gurevych · 2020
Cited alongside, same era.
Model fusion via optimal transport
S. P. Singh and M. Jaggi · 2020
Cited alongside, same era.
Contrastive representation distillation
Y. Tian, D. Krishnan, and P. Isola · 2020
Cited alongside, same era.
D. Lian, D. Zhou, J. Feng, and X. Wang · 2022
Later among the works it cites.
Peft: State-of-the-art parameter-efficient fine-tuning methods
S. Mangrulkar, S. Gugger, L. Debut, Y. Belkada, and S. Paul · 2022
Later among the works it cites.
Merging models with fisher-weighted averaging
M. Matena and C. Raffel · 2022
Later among the works it cites.
Bitfit: Simple parameter-efficient fine-tuning for transformer-based masked language-models
E. B. Zaken, Y. Goldberg, and S. Ravfogel · 2022
Later among the works it cites.
Git re-basin: Merging models modulo permutation symmetries
S. K. Ainsworth, J. Hayase, and S. S. Srinivasa · 2023
Closest in time.
Llama efficient tuning
hiyouga · 2023
Closest in time.
Fact: Factor-tuning for lightweight adaptation on vision transformer
S. Jie and Z.-H. Deng · 2023
Closest in time.
REPAIR: renormalizing permuted activations for interpolation repair
K. Jordan, H. Sedghi, O. Saukh, R. Entezari, and B. Neyshabur · 2023
Closest in time.
Zipit! merging models from different tasks without training
G. Stoica, D. Bolya, J. Bjorner, T. Hearn, and J. Hoffman · 2023
Closest in time.
Visual query tuning: Towards effective usage of intermediate representations for parameter and memory efficient transfer learning
C.-H. Tu, Z. Mai, and W.-L. Chao · 2023
Closest in time.
Generalized knowledge distillation via relationship matching
H. Ye, S. Lu, and D. Zhan · 2023
Closest in time.