Fetching the paper…
Reading the bibliography…
Post-training of language models (LMs) increasingly relies on the following two stages: (i) knowledge distillation, where the LM is trained to imitate a larger teacher LM, and (ii) reinforcement learning from human feedback (RLHF), where the LM is aligned by optimizing a reward model.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter, 2020
Sanh, V., Debut, L., Chaumond, J., and Wolf, T · 1910
Earlier work this paper cites.
‘improving ratings’: audit in the british university system
Strathern, M · 1997
Earlier work this paper cites.
Model compression
Buciluǎ, C., Caruana, R., and Niculescu-Mizil, A · 2006
Earlier work this paper cites.
Findings of the 2014 workshop on statistical machine translation
Bojar, O., Buck, C., Federmann, C., Haddow, B., Koehn, P., Leveling, J., Monz, C., Pecina, P., Post, M., Saint-Amand, H., Soricut, R., Specia, L., and Tamchyna, A · 2014
Earlier work this paper cites.
Distilling the knowledge in a neural network, 2015
Hinton, G., Vinyals, O., and Dean, J · 2015
Earlier work this paper cites.
Concrete problems in ai safety
Amodei, D., Olah, C., Steinhardt, J., Christiano, P., Schulman, J., and Mané, D · 2016
Earlier work this paper cites.
Sequence-level knowledge distillation
Kim, Y. and Rush, A. M · 2016
Earlier work this paper cites.
Don‘t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization
Narayan, S., Cohen, S. B., and Lapata, M · 2018
Earlier work this paper cites.
Tinybert: Distilling bert for natural language understanding
Jiao, X., Yin, Y., Shang, L., Jiang, X., Chen, X., Li, L., Wang, F., and Liu, Q · 2019
Earlier work this paper cites.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Earlier work this paper cites.
Autoregressive knowledge distillation through imitation learning
Lin, A., Wohlwend, J., Chen, H., and Lei, T · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Earlier work this paper cites.
Learning to summarize with human feedback
Stiennon, N., Ouyang, L., Wu, J., Ziegler, D., Lowe, R., Voss, C., Radford, A., Amodei, D., and Christiano, P. F · 2020
Earlier work this paper cites.
Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers
Wang, W., Wei, F., Dong, L., Bao, H., Yang, N., and Zhou, M · 2020
Cited alongside, same era.
Unsolved problems in ml safety
Hendrycks, D., Carlini, N., Schulman, J., and Steinhardt, J · 2021
Cited alongside, same era.
A statistical perspective on distillation
Menon, A. K., Rawat, A. S., Reddi, S., Kim, S., and Kumar, S · 2021
Cited alongside, same era.
Finetuned language models are zero-shot learners
Wei, J., Bosma, M., Zhao, V., Guu, K., Yu, A. W., Lester, B., Du, N., Dai, A. M., and Le, Q. V · 2021
Cited alongside, same era.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Bai, Y., Jones, A., Ndousse, K., Askell, A., Chen, A., DasSarma, N., Drain, D., Fort, S., Ganguli, D., Henighan, T., et al · 2022
Cited alongside, same era.
Do not blindly imitate the teacher: Using perturbed loss for knowledge distillation
Zhang, R., Shen, J., Liu, T., Liu, J., Bendersky, M., Najork, M., and Zhang, C · 2023
Later among the works it cites.
On-policy distillation of language models: Learning from self-generated mistakes
Agarwal, R., Vieillard, N., Zhou, Y., Stanczyk, P., Garea, S. R., Geist, M., and Bachem, O · 2024
Later among the works it cites.
Odin: Disentangled reward mitigates hacking in rlhf
Chen, L., Zhu, C., Soselia, D., Chen, J., Zhou, T., Goldstein, T., Huang, H., Shoeybi, M., and Catanzaro, B · 2024
Later among the works it cites.
Scaling instruction-finetuned language models
Chung, H. W., Hou, L., Longpre, S., Zoph, B., Tay, Y., Fedus, W., Li, Y., Wang, X., Dehghani, M., Brahma, S., et al · 2024
Later among the works it cites.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cross-task generalization via natural language crowdsourcing instructions
Mishra, S., Khashabi, D., Baral, C., and Hajishirzi, H · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Cited alongside, same era.
The effects of reward misspecification: Mapping and mitigating misaligned models
Pan, A., Bhatia, K., and Steinhardt, J · 2022
Cited alongside, same era.
Scaling up models and data with t5x
Roberts, A., Chung, H. W., Levskaya, A., Mishra, G., Bradbury, J., Andor, D., Narang, S., Lester, B., Gaffney, C., Mohiuddin, A., Hawthorne, C., Lewkowycz, A., Salcianu, A., van Zee, M., Austin, J., Goodman, S., Soares, L. B., Hu, H., Tsvyashchenko, S., Chowdhery, A., Bastings, J., Bulian, J., Garcia, X., Ni, J., Chen, A., Kenealy, K., Clark, J. H., Lee, S., Garrette, D., Lee-Thorp, J., Raffel, C., Shazeer, N., Ritter, M., Bosma, M., Passos, A., Maitin-Shepard, J., Fiedel, N., Omernick, M., Saeta, B., Sepassi, R., Spiridonov, A., Newlan, J., and Gesmundo, A · 2022
Cited alongside, same era.
Super-naturalinstructions:generalization via declarative instructions on 1600+ tasks
Wang, Y., Mishra, S., Alipoormolabashi, P., Kordi, Y., Mirzaei, A., Arunkumar, A., Ashok, A., Dhanasekaran, A. S., Naik, A., Stap, D., et al · 2022
Cited alongside, same era.
Scaling laws for reward model overoptimization
Gao, L., Schulman, J., and Hilton, J · 2023
Cited alongside, same era.
The flan collection: Designing data and methods for effective instruction tuning
Longpre, S., Hou, L., Vu, T., Webson, A., Chung, H. W., Tay, Y., Zhou, D., Le, Q. V., Zoph, B., Wei, J., et al · 2023
Cited alongside, same era.
Georgiev, P., Lei, V. I., Burnell, R., Bai, L., Gulati, A., Tanzer, G., Vincent, D., Pan, Z., Wang, S., et al · 2024
Later among the works it cites.
MiniLLM: Knowledge distillation of large language models
Gu, Y., Dong, L., Wei, F., and Huang, M · 2024
Later among the works it cites.
Promptkd: Distilling student-friendly knowledge for generative language models via prompt tuning
Kim, G., Jang, D., and Yang, E · 2024
Later among the works it cites.
DistiLLM: Towards streamlined distillation for large language models
Ko, J., Kim, S., Chen, T., and Yun, S.-Y · 2024
Later among the works it cites.
Scaling laws for reward model overoptimization in direct alignment algorithms
Rafailov, R., Chittepu, Y., Park, R., Sikchi, H., Hejna, J., Knox, B., Finn, C., and Niekum, S · 2024
Later among the works it cites.
Warm: On the benefits of weight averaged reward models
Ramé, A., Vieillard, N., Hussenot, L., Dadashi, R., Cideron, G., Bachem, O., and Ferret, J · 2024
Later among the works it cites.
Gemma 2: Improving open language models at a practical size
Riviere, M., Pathak, S., Sessa, P. G., Hardin, C., Bhupatiraju, S., Hussenot, L., Mesnard, T., Shahriari, B., Ramé, A., et al · 2024
Later among the works it cites.
Language models learn to mislead humans via rlhf
Wen, J., Zhong, R., Khan, A., Perez, E., Steinhardt, J., Huang, M., Bowman, S. R., He, H., and Feng, S · 2024
Later among the works it cites.
Reward hacking in reinforcement learning
Weng, L · 2024
Later among the works it cites.