Fetching the paper…
Reading the bibliography…
Layer-wise distillation is a powerful tool to compress large models (i.e.
SemEval-2017 task 1: Semantic textual similarity multilingual and crosslingual focused evaluation
Cer, D., Diab, M., Agirre, E., Lopez-Gazpio, I., and Specia, L · 2001
Earlier work this paper cites.
Automatically constructing a corpus of sentential paraphrases
Dolan, W. B. and Brockett, C · 2005
Earlier work this paper cites.
The second PASCAL recognising textual entailment challenge
Bar-Haim, R., Dagan, I., Dolan, B., Ferro, L., and Giampiccolo, D · 2006
Earlier work this paper cites.
The pascal recognising textual entailment challenge
Dagan, I., Glickman, O., and Magnini, B · 2006
Earlier work this paper cites.
The third PASCAL recognizing textual entailment challenge
Giampiccolo, D., Magnini, B., Dagan, I., and Dolan, B · 2007
Earlier work this paper cites.
The fifth pascal recognizing textual entailment challenge
Bentivogli, L., Dagan, I., Dang, H. T., Giampiccolo, D., and Magnini, B · 2009
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Socher, R., Perelygin, A., Wu, J., Chuang, J., Manning, C. D., Ng, A., and Potts, C · 2013
Earlier work this paper cites.
Distilling the knowledge in a neural network
Hinton, G., Vinyals, O., and Dean, J · 2015
Earlier work this paper cites.
Fitnets: Hints for thin deep nets
Romero, A., Ballas, N., Kahou, S. E., Chassang, A., Gatta, C., and Bengio, Y · 2015
Earlier work this paper cites.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Zhu, Y., Kiros, R., Zemel, R. S., Salakhutdinov, R., Urtasun, R., Torralba, A., and Fidler, S · 2015
Earlier work this paper cites.
Gaussian error linear units (gelus)
Hendrycks, D. and Gimpel, K · 2016
Earlier work this paper cites.
The LAMBADA dataset: Word prediction requiring a broad discourse context
Paperno, D., Kruszewski, G., Lazaridou, A., Pham, N. Q., Bernardi, R., Pezzelle, S., Baroni, M., Boleda, G., and Fernández, R · 2016
Earlier work this paper cites.
SQuAD: 100,000+ questions for machine comprehension of text
Rajpurkar, P., Zhang, J., Lopyrev, K., and Liang, P · 2016
Earlier work this paper cites.
SQuAD: 100,000+ questions for machine comprehension of text
Rajpurkar, P., Zhang, J., Lopyrev, K., and Liang, P · 2016
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Sennrich, R., Haddow, B., and Birch, A · 2016
Earlier work this paper cites.
Pointer sentinel mixture models
Merity, S., Xiong, C., Bradbury, J., and Socher, R · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Earlier work this paper cites.
Know what you don’t know: Unanswerable questions for SQuAD
Rajpurkar, P., Jia, R., and Liang, P · 2018
Earlier work this paper cites.
A simple method for commonsense reasoning
Trinh, T. H. and Le, Q. V · 2018
Cited alongside, same era.
A broad-coverage challenge corpus for sentence understanding through inference
Williams, A., Nangia, N., and Bowman, S · 2018
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Cited alongside, same era.
Openwebtext corpus, 2019
Gokaslan, A., Cohen, V., Pavlick, E., and Tellex, S · 2019
Cited alongside, same era.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
TinyBERT: Distilling BERT for natural language understanding
Jiao, X., Yin, Y., Shang, L., Jiang, X., Chen, X., Li, L., Wang, F., and Liu, Q · 2020
Later among the works it cites.
Mixkd: Towards efficient distillation of large-scale language models
Liang, K. J., Hao, W., Shen, D., Zhou, Y., Chen, W., Chen, C., and Carin, L · 2020
Later among the works it cites.
Contrastive distillation on intermediate representations for language model compression
Sun, S., Gan, Z., Fang, Y., Cheng, Y., Wang, S., and Liu, J · 2020
Later among the works it cites.
MobileBERT: a compact task-agnostic BERT for resource-limited devices
Sun, Z., Yu, H., Song, X., Liu, R., Yang, Y., and Zhou, D · 2020
Later among the works it cites.
Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers
Wang, W., Wei, F., Dong, L., Bao, H., Yang, N., and Zhou, M · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2019
Cited alongside, same era.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
Sanh, V., Debut, L., Chaumond, J., and Wolf, T · 2019
Cited alongside, same era.
Patient knowledge distillation for BERT model compression
Sun, S., Cheng, Y., Gan, Z., and Liu, J · 2019
Cited alongside, same era.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R · 2019
Cited alongside, same era.
Neural network acceptability judgments
Warstadt, A., Singh, A., and Bowman, S. R · 2019
Cited alongside, same era.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Cited alongside, same era.
BERT-of-theseus: Compressing BERT by progressive module replacing
Xu, C., Zhou, W., Ge, T., Wei, F., and Zhou, M · 2020
Later among the works it cites.
Multitask prompted training enables zero-shot task generalization
Sanh, V., Webson, A., Raffel, C., Bach, S. H., Sutawika, L., Alyafeai, Z., Chaffin, A., Stiegler, A., Scao, T. L., Raja, A., et al · 2021
Later among the works it cites.
Follow your path: a progressive method for knowledge distillation
Shi, W., Song, Y., Zhou, H., Li, B., and Li, L · 2021
Later among the works it cites.
MiniLMv2: Multi-head self-attention relation distillation for compressing pretrained transformers
Wang, W., Bao, H., Huang, S., Dong, L., and Wei, F · 2021
Later among the works it cites.
Moebert: from bert to mixture-of-experts via importance-guided adaptation
Zuo, S., Zhang, Q., Liang, C., He, P., Zhao, T., and Chen, W · 2022
Closest in time.
DeBERTav3: Improving deBERTa using ELECTRA-style pre-training with gradient-disentangled embedding sharing
He, P., Gao, J., and Chen, W · 2023
Closest in time.
Hsieh, C.-Y., Li, C.-L., Yeh, C.-K., Nakhost, H., Fujii, Y., Ratner, A., Krishna, R., Lee, C.-Y., and Pfister, T · 2023
Closest in time.
Lion: Adversarial distillation of closed-source large language model
Jiang, Y., Chan, C., Chen, M., and Wang, W · 2023
Closest in time.
Homodistil: Homotopic task-agnostic distillation of pre-trained transformers
Liang, C., Jiang, H., Li, Z., Tang, X., Yin, B., and Zhao, T · 2023
Closest in time.
The flan collection: Designing data and methods for effective instruction tuning
Longpre, S., Hou, L., Vu, T., Webson, A., Chung, H. W., Tay, Y., Zhou, D., Le, Q. V., Zoph, B., Wei, J., et al · 2023
Closest in time.
Peng, B., Li, C., He, P., Galley, M., and Gao, J · 2023
Closest in time.
Stanford alpaca: An instruction-following llama model
Taori, R., Gulrajani, I., Zhang, T., Dubois, Y., Li, X., Guestrin, C., Liang, P., and Hashimoto, T. B · 2023
Closest in time.