Fetching the paper…
Reading the bibliography…
Language models only really need to use an exponential fraction of their neurons for individual inferences.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Radford, A., Narasimhan, K., Salimans, T., Sutskever, I., et al · 2018
Earlier work this paper cites.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R · 2018
Earlier work this paper cites.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
Sanh, V., Debut, L., Chaumond, J., and Wolf, T · 2019
Earlier work this paper cites.
Patient knowledge distillation for bert model compression
Sun, S., Cheng, Y., Gan, Z., and Liu, J · 2019
Cited alongside, same era.
Well-read students learn better: On the importance of pre-training compact models
Turc, I., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Cited alongside, same era.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Cited alongside, same era.
Xtremedistiltransformers: Task transfer for task-agnostic distillation
Mukherjee, S., Awadallah, A. H., and Gao, J · 2021
Cited alongside, same era.
Anil, R., Dai, A. M., Firat, O., Johnson, M., Lepikhin, D., Passos, A., Shakeri, S., Taropa, E., Bailey, P., Chen, Z., et al · 2023
Closest in time.
Belcak, P. and Wattenhofer, R · 2023
Closest in time.
Cramming: Training a language model on a single gpu in one day
Geiping, J. and Goldstein, T · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…