Fetching the paper…
Reading the bibliography…
Adapters are a parameter-efficient alternative to fine-tuning, which augment a frozen base network to learn new tasks.
Optimal brain damage
LeCun, Y., Denker, J., and Solla, S · 1989
Earlier work this paper cites.
Second order derivatives for network pruning: Optimal brain surgeon
Hassibi, B. and Stork, D. G · 1992
Earlier work this paper cites.
Quantization
Gray, R. and Neuhoff, D · 1998
Earlier work this paper cites.
Asirra: A captcha that exploits interest-aligned manual image categorization
Elson, J., Douceur, J. J., Howell, J., and Saul, J · 2007
Earlier work this paper cites.
Automated flower classification over a large number of classes
Nilsback, M.-E. and Zisserman, A · 2008
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A · 2009
Earlier work this paper cites.
One weird trick for parallelizing convolutional neural networks
Krizhevsky, A · 2014
Earlier work this paper cites.
On pixel-wise explanations for non-linear classifier decisions by layer-wise relevance propagation
Bach, S., Binder, A., Montavon, G., Klauschen, F., Müller, K.-R., and Samek, W · 2015
Earlier work this paper cites.
Learning both weights and connections for efficient neural networks
Han, S., Pool, J., Tran, J., and Dally, W. J · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
Hinton, G. E., Vinyals, O., and Dean, J · 2015
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., Berg, A. C., and Fei-Fei, L · 2015
Earlier work this paper cites.
Communication-efficient learning of deep networks from decentralized data
McMahan, H. B., Moore, E., Ramage, D., Hampson, S., and y Arcas, B. A · 2016
Earlier work this paper cites.
Gpu kernels for block-sparse weights
Gray, S., Radford, A., and Kingma, D. P · 2017
Earlier work this paper cites.
Pruning filters for efficient convnets
Li, H., Kadav, A., Durdanovic, I., Samet, H., and Graf, H. P · 2017
Earlier work this paper cites.
Pruning convolutional neural networks for resource efficient inference
Molchanov, P., Tyree, S., Karras, T., Aila, T., and Kautz, J · 2017
Earlier work this paper cites.
Learning multiple visual domains with residual adapters
Rebuffi, S.-A., Bilen, H., and Vedaldi, A · 2017
Earlier work this paper cites.
meProp: Sparsified back propagation for accelerated deep learning with reduced overfitting
Sun, X., Ren, X., Ma, S., and Wang, H · 2017
Cited alongside, same era.
Channel pruning based on mean gradient for accelerating convolutional neural networks
Liu, C. and Wu, H · 2018
Cited alongside, same era.
Piggyback: Adapting a single network to multiple tasks by learning to mask weights
Mallya, A., Davis, D., and Lazebnik, S · 2018
Cited alongside, same era.
Efficient parametrization of multi-domain deep neural networks
Rebuffi, S.-A., Bilen, H., and Vedaldi, A · 2018
Cited alongside, same era.
Improving efficiency in convolutional neural network with multilinear filters
Tran, D. T., Iosifidis, A., and Gabbouj, M · 2018
Cited alongside, same era.
To prune, or not to prune: Exploring the efficacy of pruning for model compression
Block pruning for faster transformers
Lagunas, F., Charlaix, E., Sanh, V., and Rush, A · 2021
Later among the works it cites.
Prefix-tuning: Optimizing continuous prompts for generation
Li, X. L. and Liang, P · 2021
Later among the works it cites.
Pruning and quantization for deep neural network acceleration: A survey
Liang, T., Glossner, J., Wang, L., Shi, S., and Zhang, X · 2021
Later among the works it cites.
Compacter: Efficient low-rank hypercomplex adapter layers
Mahabadi, R. K., Henderson, J., and Ruder, S · 2021
Later among the works it cites.
AdapterFusion: Non-destructive task composition for transfer learning
Pfeiffer, J., Kamath, A., Rücklé, A., Cho, K., and Gurevych, I · 2021
Later among the works it cites.
AdapterDrop: On the efficiency of adapters in transformers
Rücklé, A., Geigle, G., Glockner, M., Beck, T., Pfeiffer, J., Reimers, N., and Gurevych, I · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zhu, M. and Gupta, S · 2018
Cited alongside, same era.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Frankle, J. and Carbin, M · 2019
Cited alongside, same era.
Compressing by learning in a low-rank and sparse decomposition form
Guo, K., Xie, X., Xu, X., and Xing, X · 2019
Cited alongside, same era.
Parameter-efficient transfer learning for NLP
Houlsby, N., Giurgiu, A., Jastrzebski, S., Morrone, B., De Laroussilhe, Q., Gesmundo, A., Attariyan, M., and Gelly, S · 2019
Cited alongside, same era.
Layer-Wise Relevance Propagation: An Overview , pp. 193–209
Montavon, G., Binder, A., Lapuschkin, S., Samek, W., and Müller, K.-R · 2019
Cited alongside, same era.
BERT and PALs: Projected attention layers for efficient adaptation in multi-task learning
Stickland, A. C. and Murray, I · 2019
Cited alongside, same era.
EfficientNet: Rethinking model scaling for convolutional neural networks
Tan, M. and Le, Q · 2019
Cited alongside, same era.
Later among the works it cites.
Pruning by explaining: A novel criterion for deep neural network pruning
Yeom, S.-K., Seegerer, P., Lapuschkin, S., Binder, A., Wiedemann, S., Müller, K.-R., and Samek, W · 2021
Later among the works it cites.
Counter-interference adapter for multilingual machine translation
Zhu, Y., Feng, J., Zhao, C., Wang, M., and Li, L · 2021
Later among the works it cites.
Towards a unified view of parameter-efficient transfer learning
He, J., Zhou, C., Ma, X., Berg-Kirkpatrick, T., and Neubig, G · 2022
Closest in time.
Continual 3d convolutional neural networks for real-time processing of videos
Hedegaard, L. and Iosifidis, A · 2022
Closest in time.
DatasetOps, 2022
Hedegaard, L., Oleksiienko, I., and Legaard, C. M · 2022
Closest in time.
Training compute-optimal large language models
Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., Casas, D. d. L., Hendricks, L. A., Welbl, J., Clark, A., Hennigan, T., Noland, E., Millican, K., Driessche, G. v. d., Damoc, B., Guy, A., Osindero, S., Simonyan, K., Elsen, E., Rae, J. W., Vinyals, O., and Sifre, L · 2022
Closest in time.
LoRA: Low-rank adaptation of large language models
Hu, E. J., yelong shen, Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2022
Closest in time.
SOSP: Efficiently capturing global correlations by second-order structured pruning
Nonnenmacher, M., Pfeil, T., Steinwart, I., and Reeb, D · 2022
Closest in time.
Continual Transformers: Redundancy-free attention for online inference
Hedegaard, L., Bakhtiarnia, A., and Iosifidis, A · 2023
Closest in time.