Fetching the paper…
Reading the bibliography…
Instruction fine-tuning of large language models (LLMs) is a powerful method for improving task-specific performance, but it can inadvertently lead to a phenomenon where models generate harmful responses when faced with malicious prompts.
Ouyang, Long, Wu, Jeff, Jiang, Xu, Almeida, Diogo, Wainwright, Carroll L., Mishkin, Pamela, Zhang, Chong, Agarwal, Sandhini, Slama, Katarina, Ray, Alex, Schulman, John, Hilton, Jacob, Kelton, Fraser, Miller, Luke, Simens, Maddie, Askell, Amanda, Welinder, Peter, Christiano, Paul, Leike, Jan, and Lowe, Ryan. (2024). Training language models to follow instructions with human feedback. Proceedings of the 36th International Conference on Neural Information Processing Systems (NIPS ’22)
2011
Earlier work this paper cites.
Christiano, P.F., Leike, J., Brown, T., Martic, M., Legg, S., Amodei, D.: Deep Reinforcement Learning from Human Preferences. In: Advances in Neural Information Processing Systems (NeurIPS), vol. 30, 2017. https://proceedings.neurips.cc/paper_files/paper/2017/file/d5e2c0adad503c91f91df240d0cd4e49-Paper.pdf
2017
Earlier work this paper cites.
Houlsby, Neil, Giurgiu, Andrei, Jastrzebski, Stanislaw, Morrone, Bruna, de Laroussilhe, Quentin, Gesmundo, Andrea, Attariyan, Mona, and Gelly, Sylvain. (2019). Parameter-Efficient Transfer Learning for NLP. In Kamalika Chaudhuri and Ruslan Salakhutdinov (Eds.), Proceedings of the 36th International Conference on Machine Learning (ICML 2019)
2019
Earlier work this paper cites.
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J.D., Dhariwal, P., Neelakantan, A., et al.: Language Models are Few-Shot Learners. In: NeurIPS 2020, pp. 1877–1901. https://proceedings.neurips.cc/paper/2020/file/1457c0d6bfcb4967418bfb8ac142f64a-Paper.pdf
2020
Earlier work this paper cites.
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., Steinhardt, J.: Measuring Massive Multitask Language Understanding. ICLR 2021. https://openreview.net/forum?id=d7KBjmI3GmQ
2021
Earlier work this paper cites.
Hu, E.J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., Chen, W.: LoRA: Low-rank adaptation of large language models. In: International Conference on Learning Representations (ICLR) (2022). https://openreview.net/forum?id=nZeVKeeFYf9
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
Mangrulkar, S., Gugger, S., Debut, L., Belkada, Y., Paul, S., Bossan, B.: PEFT: State-of-the-art Parameter-Efficient Fine-Tuning methods. https://github.com/huggingface/peft , last accessed 2023/10/25
2023
Earlier work this paper cites.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Wei, J., Li, X., Lei, G., Zhang, Y. (2023). Wei et al. Respond to "Safer, More Precise Management of Osteoarthritis Pain". American Journal of Epidemiology
2023
Cited alongside, same era.
Zheng, Lianmin, Chiang, Wei-Lin, Sheng, Ying, Zhuang, Siyuan, Wu, Zhanghao, Zhuang, Yonghao, Lin, Zi, Li, Zhuohan, Li, Dacheng, Xing, Eric, Zhang, Hao, Gonzalez, Joseph E., and Stoica, Ion. (2023). Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. Thirty-seventh Conference on Neural Information Processing Systems Datasets and Benchmarks Track
Bianchi, F., Suzgun, M., Attanasio, G., Röttger, P., Jurafsky, D., Hashimoto, T., Zou, J.: Safety-Tuned LLaMAs: Lessons from Improving the Safety of Large Language Models that Follow Instructions. In: ICLR 2024. https://openreview.net/forum?id=gT5hALch9z
2024
Closest in time.
2024
Closest in time.
Röttger, P., Kirk, H., Vidgen, B., Attanasio, G., Bianchi, F., Hovy, D.: XSTest: Identifying Exaggerated Safety Behaviours in LLMs. NAACL-HLT 2024, pp. 5377–5400. https://aclanthology.org/2024.naacl-long.301
2024
Closest in time.
2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2023
Cited alongside, same era.
2024
Cited alongside, same era.
Jain, S., Kirk, R., Lubana, E.S., Dick, R.P., Tanaka, H., Rocktäschel, T., Grefenstette, E., Krueger, D.: Mechanistically analyzing the effects of fine-tuning on procedurally defined tasks. In: The Twelfth International Conference on Learning Representations (ICLR) (2024). https://openreview.net/forum?id=A0HKeKl4Nl
2024
Cited alongside, same era.
Qi, X., Zeng, Y., Xie, T., Chen, P.-Y., Jia, R., Mittal, P., Henderson, P.: Fine-tuning aligned language models compromises safety, even when users do not intend to! In: The Twelfth International Conference on Learning Representations (ICLR) (2024). https://openreview.net/forum?id=hTEGyKf0dZ
2024
Cited alongside, same era.
Zhan, Q., Fang, R., Bindu, R., Gupta, A., Hashimoto, T., Kang, D.: Removing RLHF Protections in GPT-4 via Fine-Tuning. In: Proceedings of the 2024 NAACL-HLT, vol. 2, pp. 681–687, Mexico City. https://aclanthology.org/2024.naacl-short.59
2024
Cited alongside, same era.
Zhang, J., Chen, S., Liu, J., He, J.: Composing parameter-efficient modules with arithmetic operations. In: Proceedings of the 37th International Conference on Neural Information Processing Systems (NIPS ’23), Red Hook, NY, USA, Curran Associates Inc., 2024, Article No. 552, pp. 1–22
2024
Closest in time.
Wu, X., Huang, S., Wei, F.: Mixture of LoRA Experts. In: The Twelfth International Conference on Learning Representations, 2024. https://openreview.net/forum?id=uWvKBCYh4S
2024
Closest in time.