Fetching the paper…
Reading the bibliography…
Training large transformer models from scratch for a target task requires lots of data and is computationally demanding.
Catastrophic forgetting in connectionist networks
French, R. M · 1999
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Glorot, X. and Bengio, Y · 2010
Earlier work this paper cites.
Learning both weights and connections for efficient neural network
Han, S., Pool, J., Tran, J., and Dally, W · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Channel pruning for accelerating very deep neural networks
He, Y., Zhang, X., and Sun, J · 2017
Earlier work this paper cites.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Frankle, J. and Carbin, M · 2018
Earlier work this paper cites.
Once-for-all: Train one network and specialize it for efficient deployment
Cai, H., Gan, C., Wang, T., Zhang, Z., and Han, S · 2019
Earlier work this paper cites.
Transformer-xl: Attentive language models beyond a fixed-length context
Dai, Z., Yang, Z., Yang, Y., Carbonell, J., Le, Q. V., and Salakhutdinov, R · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
Earlier work this paper cites.
Bert rediscovers the classical nlp pipeline
Tenney, I., Das, D., and Pavlick, E · 2019
Earlier work this paper cites.
What is the state of neural network pruning?
Blalock, D., Gonzalez Ortiz, J. J., Frankle, J., and Guttag, J · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al · 2020
Cited alongside, same era.
The pile: An 800gb dataset of diverse text for language modeling
Gao, L., Biderman, S., Black, S., Golding, L., Hoppe, T., Foster, C., Phang, J., He, H., Thite, A., Nabeshima, N., et al · 2020
Cited alongside, same era.
Weight distillation: Transferring the knowledge in neural network parameters
Lin, Y., Li, Y., Wang, Z., Li, B., Du, Q., Xiao, T., and Zhu, J · 2020
Cited alongside, same era.
The right tool for the job: Matching model and instance complexities
A survey on vision transformer
Han, K., Wang, Y., Chen, H., Chen, X., Guo, J., Liu, Z., Tang, Y., Xiao, A., Xu, C., Xu, Y., et al · 2022
Later among the works it cites.
Knowledge distillation via the target-aware transformer
Lin, S., Xie, H., Wang, B., Yu, K., Chang, X., Liang, X., and Wang, G · 2022
Later among the works it cites.
Self-evolving vision transformer for chest x-ray diagnosis through knowledge distillation
Park, S., Kim, G., Oh, Y., Seo, J. B., Lee, S. M., Kim, J. H., Moon, S., Lim, J.-K., Park, C. M., and Ye, J. C · 2022
Later among the works it cites.
Jump to conclusions: Short-cutting transformers with linear transformations
Din, A. Y., Karidi, T., Choshen, L., and Geva, M · 2023
Closest in time.
Reinforce data, multiply impact: Improved model accuracy and robustness with dataset reinforcement
Faghri, F., Pouransari, H., Mehta, S., Farajtabar, M., Farhadi, A., Rastegari, M., and Tuzel, O · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Schwartz, R., Stanovsky, G., Swayamdipta, S., Dodge, J., and Smith, N. A · 2020
Cited alongside, same era.
Bignas: Scaling up neural architecture search with big single-stage models
Yu, J., Jin, P., Liu, H., Bender, G., Kindermans, P.-J., Tan, M., Huang, T., Song, X., Pang, R., and Le, Q · 2020
Cited alongside, same era.
Knowledge distillation: A survey
Gou, J., Yu, B., Maybank, S. J., and Tao, D · 2021
Cited alongside, same era.
Mediators in determining what processing bert performs first
Slobodkin, A., Choshen, L., and Abend, O · 2021
Cited alongside, same era.
Gpt3. int8 (): 8-bit matrix multiplication for transformers at scale
Dettmers, T., Lewis, M., Belkada, Y., and Zettlemoyer, L · 2022
Cited alongside, same era.
Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space
Geva, M., Caciularu, A., Wang, K. R., and Goldberg, Y · 2022
Cited alongside, same era.
Alphanet: Improved training of supernets with alpha-divergence
Wang, D., Gong, C., Li, M., Liu, Q., and Chandra, V
Cited in the paper.
Attentivenas: Improving neural architecture search via attentive sampling
Wang, D., Li, M., Gong, C., and Chandra, V
Cited in the paper.
Name of the model checkpoint
HuggingFace · 2023
Closest in time.
Deja vu: Contextual sparsity for efficient llms at inference time
Liu, Z., Wang, J., Dao, T., Zhou, T., Yuan, B., Song, Z., Shrivastava, A., Zhang, C., Tian, Y., Re, C., et al · 2023
Closest in time.
Relu strikes back: Exploiting activation sparsity in large language models, 2023
Mirzadeh, I., Alizadeh, K., Mehta, S., Mundo, C. C. D., Tuzel, O., Samei, G., Rastegari, M., and Farajtabar, M · 2023
Closest in time.
A survey of controllable text generation using transformer-based pre-trained language models
Zhang, H., Song, H., Li, S., Zhou, M., and Song, D · 2023
Closest in time.
A survey on efficient training of transformers
Zhuang, B., Liu, J., Pan, Z., He, H., Weng, Y., and Shen, C · 2023
Closest in time.