Fetching the paper…
Reading the bibliography…
Knowledge distillation (KD) is widely used for compressing a teacher model to a smaller student model, reducing its inference cost and memory footprint while preserving model capabilities.
A unified bias-variance decomposition and its applications
Pedro, D · 2000
Earlier work this paper cites.
On the effectiveness of the skew divergence for statistical language analysis
Lee, L · 2001
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Papineni, K., Roukos, S., Ward, T., and Zhu, W.-J · 2002
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Lin, C.-Y · 2004
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Distilling the knowledge in a neural network
Hinton, G. E., Vinyals, O., and Dean, J · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Earlier work this paper cites.
Sequence-level knowledge distillation
Kim, Y. and Rush, A. M · 2016
Earlier work this paper cites.
Overview of the IWSLT 2017 evaluation campaign
Cettolo, M., Federico, M., Bentivogli, L., Niehues, J., Stüker, S., Sudoh, K., Yoshino, K., and Federmann, C · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2017
Earlier work this paper cites.
Get to the point: Summarization with pointer-generator networks
See, A., Liu, P. J., and Manning, C. D · 2017
Earlier work this paper cites.
Don’t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization
Narayan, S., Cohen, S. B., and Lapata, M · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Radford, A., Narasimhan, K., Salimans, T., Sutskever, I., et al · 2018
Earlier work this paper cites.
Samsum corpus: A human-annotated dialogue dataset for abstractive summarization
Gliwa, B., Mochol, I., Biesek, M., and Wawer, A · 2019
Earlier work this paper cites.
Openwebtext corpus, 2019
Gokaslan, A., Cohen, V., Pavlick, E., and Tellex, S · 2019
Earlier work this paper cites.
Stabilizing off-policy q-learning via bootstrapping error reduction
Kumar, A., Fu, J., Soh, M., Tucker, G., and Levine, S · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
Earlier work this paper cites.
Practical and consistent estimation of f-divergences
Rubenstein, P., Bousquet, O., Djolonga, J., Riquelme, C., and Tolstikhin, I. O · 2019
Cited alongside, same era.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
Sanh, V., Debut, L., Chaumond, J., and Wolf, T · 2019
Cited alongside, same era.
Patient knowledge distillation for BERT model compression
Sun, S., Cheng, Y., Gan, Z., and Liu, J · 2019
Cited alongside, same era.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Cited alongside, same era.
Revisiting fundamentals of experience replay
Fedus, W., Ramachandran, P., Agarwal, R., Bengio, Y., Larochelle, H., Rowland, M., and Dabney, W · 2020
Cited alongside, same era.
Autoregressive knowledge distillation through imitation learning
Anil, R., Dai, A. M., Firat, O., Johnson, M., Lepikhin, D., Passos, A., Shakeri, S., Taropa, E., Bailey, P., Chen, Z., et al · 2023
Later among the works it cites.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, March 2023
Chiang, W.-L., Li, Z., Lin, Z., Sheng, Y., Wu, Z., Zhang, H., Zheng, L., Zhuang, S., Zhuang, Y., Gonzalez, J. E., Stoica, I., and Xing, E. P · 2023
Later among the works it cites.
Free dolly: Introducing the world’s first truly open instruction-tuned llm, 2023
Conover, M., Hayes, M., Mathur, A., Xie, J., Wan, J., Shah, S., Ghodsi, A., Wendell, P., Zaharia, M., and Xin, R · 2023
Later among the works it cites.
Openllama: An open reproduction of llama, May 2023
Geng, X. and Liu, H · 2023
Later among the works it cites.
Unnatural instructions: Tuning language models with (almost) no human labor
Honovich, O., Scialom, T., Levy, O., and Schick, T · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lin, A., Wohlwend, J., Chen, H., and Lei, T · 2020
Cited alongside, same era.
Improved knowledge distillation via teacher assistant
Mirzadeh, S. I., Farajtabar, M., Li, A., Levine, N., Matsukawa, A., and Ghasemzadeh, H · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Cited alongside, same era.
Divergence frontiers for generative models: Sample complexity, quantization effects, and frontier integrals
Liu, L., Pillutla, K., Welleck, S., Oh, S., Choi, Y., and Harchaoui, Z · 2021
Cited alongside, same era.
mT5: A massively multilingual pre-trained text-to-text transformer
Xue, L., Constant, N., Roberts, A., Kale, M., Al-Rfou, R., Siddhant, A., Barua, A., and Raffel, C · 2021
Cited alongside, same era.
Why exposure bias matters: An imitation learning perspective of error accumulation in language generation
Arora, K., El Asri, L., Bahuleyan, H., and Cheung, J · 2022
Cited alongside, same era.
LoRA: Low-rank adaptation of large language models
Hu, E. J., yelong shen, Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2022
Cited alongside, same era.
Hsieh, C.-Y., Li, C.-L., Yeh, C.-K., Nakhost, H., Fujii, Y., Ratner, A., Krishna, R., Lee, C.-Y., and Pfister, T · 2023
Later among the works it cites.
Tailoring language generation models under total variation distance
Ji, H., Ke, P., Hu, Z., Zhang, R., and Huang, M · 2023
Later among the works it cites.
Plastic: Improving input and label plasticity for sample efficient reinforcement learning
Lee, H., Cho, H., Kim, H., Gwak, D., Kim, J., Choo, J., Yun, S.-Y., and Yun, C · 2023
Later among the works it cites.
Gpt-4 technical report, 2023
OpenAI · 2023
Later among the works it cites.
Instruction tuning with gpt-4, 2023
Peng, B., Li, C., He, P., Galley, M., and Gao, J · 2023
Later among the works it cites.
Stanford alpaca: An instruction-following llama model
Taori, R., Gulrajani, I., Zhang, T., Dubois, Y., Li, X., Guestrin, C., Liang, P., and Hashimoto, T. B · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Later among the works it cites.
Self-instruct: Aligning language models with self-generated instructions
Wang, Y., Kordi, Y., Mishra, S., Liu, A., Smith, N. A., Khashabi, D., and Hajishirzi, H · 2023
Later among the works it cites.
f-divergence minimization for sequence-level knowledge distillation
Wen, Y., Li, Z., Du, W., and Mou, L · 2023
Later among the works it cites.
Judging LLM-as-a-judge with MT-bench and chatbot arena
Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E., Zhang, H., Gonzalez, J. E., and Stoica, I · 2023
Later among the works it cites.
On-policy distillation of language models: Learning from self-generated mistakes
Agarwal, R., Vieillard, N., Stanczyk, P., Ramos, S., Geist, M., and Bachem, O · 2024
Closest in time.
MiniLLM: Knowledge distillation of large language models
Gu, Y., Dong, L., Wei, F., and Huang, M · 2024
Closest in time.