Fetching the paper…
Reading the bibliography…
Recent studies show that Large Language Models (LLMs) with safety alignment can be jail-broken by fine-tuning on a dataset mixed with harmful data.
Proximal alternating minimization and projection methods for nonconvex problems: An approach based on the kurdyka-łojasiewicz inequality
Attouch, H., Bolte, J., Redont, P., and Soubeyran, A · 2010
Earlier work this paper cites.
Policy shaping: Integrating human feedback with reinforcement learning
Griffith, S., Subramanian, K., Scholz, J., Isbell, C. L., and Thomaz, A. L · 2013
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Socher, R., Perelygin, A., Wu, J., Chuang, J., Manning, C. D., Ng, A. Y., and Potts, C · 2013
Earlier work this paper cites.
Character-level convolutional networks for text classification
Zhang, X., Zhao, J., and LeCun, Y · 2015
Earlier work this paper cites.
Linear convergence of gradient and proximal-gradient methods under the polyak-łojasiewicz condition
Karimi, H., Nutini, J., and Schmidt, M · 2016
Earlier work this paper cites.
Douglas–rachford splitting for nonconvex optimization with application to nonconvex feasibility problems
Li, G. and Pong, T. K · 2016
Earlier work this paper cites.
Overcoming catastrophic forgetting in neural networks
Kirkpatrick, J., Pascanu, R., Rabinowitz, N., Veness, J., Desjardins, G., Rusu, A. A., Milan, K., Quan, J., Ramalho, T., Grabska-Barwinska, A., et al · 2017
Earlier work this paper cites.
Federated optimization in heterogeneous networks
Li, T., Sahu, A. K., Zaheer, M., Sanjabi, M., Talwalkar, A., and Smith, V · 2018
Earlier work this paper cites.
An algorithmic framework of variable metric over-relaxed hybrid proximal extra-gradient method
Shen, L., Sun, P., Wang, Y., Liu, W., and Zhang, T · 2018
Earlier work this paper cites.
On the convergence of fedavg on non-iid data
Li, X., Huang, K., Yang, W., Wang, S., and Zhang, Z · 2019
Earlier work this paper cites.
Meta-learning with implicit gradients
Rajeswaran, A., Finn, C., Kakade, S. M., and Levine, S · 2019
Earlier work this paper cites.
Ye, S., Feng, X., Zhang, T., Ma, X., Lin, S., Li, Z., Xu, K., Wen, W., Liu, S., Tang, J., et al · 2019
Earlier work this paper cites.
Efficient meta learning via minibatch proximal update
Zhou, P., Yuan, X., Xu, H., Yan, S., and Feng, J · 2019
Earlier work this paper cites.
Low-rank compression of neural nets: Learning the rank of each layer
Idelbayev, Y. and Carreira-Perpinán, M. A · 2020
Earlier work this paper cites.
Map inference via ℓ 2 \ell_{2} -sphere linear program reformulation
Wu, B., Shen, L., Zhang, T., and Ghanem, B · 2020
Earlier work this paper cites.
Federated learning based on dynamic regularization
Acar, D. A. E., Zhao, Y., Navarro, R. M., Mattina, M., Whatmough, P. N., and Saligrama, V · 2021
Earlier work this paper cites.
Proximal gradient descent-ascent: Variable convergence under k { \{ \ \backslash L } \} geometry
Chen, Z., Zhou, Y., Xu, T., and Liang, Y · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., et al · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2021
Earlier work this paper cites.
Defending against backdoors in federated learning with robust learning rate
Ozdayi, M. S., Kantarcioglu, M., and Gel, Y. R · 2021
Earlier work this paper cites.
Fedcm: Federated learning with client-level momentum
Xu, J., Wang, S., Wang, L., and Yao, A. C.-C · 2021
Earlier work this paper cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Bai, Y., Jones, A., Ndousse, K., Askell, A., Chen, A., DasSarma, N., Drain, D., Fort, S., Ganguli, D., Henighan, T., et al · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Earlier work this paper cites.
Opt: Open pre-trained transformer language models
Zhang, S., Roller, S., Goyal, N., Artetxe, M., Chen, M., Chen, S., Dewan, C., Diab, M., Li, X., Lin, X. V., et al · 2022
Cited alongside, same era.
Bianchi, F., Suzgun, M., Attanasio, G., Röttger, P., Jurafsky, D., Hashimoto, T., and Zou, J · 2023
Cited alongside, same era.
Jailbreaking black box large language models in twenty queries
Chao, P., Robey, A., Dobriban, E., Hassani, H., Pappas, G. J., and Wong, E · 2023
Cited alongside, same era.
Safe rlhf: Safe reinforcement learning from human feedback
Dai, J., Pan, X., Sun, R., Ji, J., Xu, X., Liu, M., Wang, Y., and Yang, Y · 2023
Cited alongside, same era.
Learning and forgetting unsafe examples in large language models
Zhao, J., Deng, Z., Madras, D., Zou, J., and Ren, M · 2023
Later among the works it cites.
Representation engineering: A top-down approach to ai transparency
Zou, A., Phan, L., Chen, S., Campbell, J., Guo, P., Ren, R., Pan, A., Yin, X., Mazeika, M., Dombrowski, A.-K., et al · 2023
Later among the works it cites.
Are aligned neural networks adversarially aligned?
Carlini, N., Nasr, M., Choquette-Choo, C. A., Jagielski, M., Gao, I., Koh, P. W. W., Ippolito, D., Tramer, F., and Schmidt, L · 2024
Closest in time.
Chen, C., Huang, B., Li, Z., Chen, Z., Lai, S., Xu, X., Gu, J.-C., Gu, J., Yao, H., Xiao, C., et al · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dong, H., Xiong, W., Goyal, D., Pan, R., Diao, S., Zhang, J., Shum, K., and Zhang, T · 2023
Cited alongside, same era.
Large language model-powered smart contract vulnerability detection: New perspectives
Hu, S., Huang, T., İlhan, F., Tekin, S. F., and Liu, L · 2023
Cited alongside, same era.
Bert4eth: A pre-trained transformer for ethereum fraud detection
Hu, S., Zhang, Z., Luo, B., Lu, S., He, B., and Liu, L · 2023
Cited alongside, same era.
Fusion of global and local knowledge for personalized federated learning
Huang, T., Shen, L., Sun, Y., Lin, W., and Tao, D · 2023
Cited alongside, same era.
Beavertails: Towards improved safety alignment of llm via a human-preference dataset
Ji, J., Liu, M., Dai, J., Pan, X., Zhang, C., Bian, C., Sun, R., Wang, Y., and Yang, Y · 2023
Cited alongside, same era.
Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., Casas, D. d. l., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., et al · 2023
Cited alongside, same era.
Lora fine-tuning efficiently undoes safety training in llama 2-chat 70b
Lermen, S., Rogers-Smith, C., and Ladish, J · 2023
Cited alongside, same era.
Alpacaeval: An automatic evaluator of instruction-following models
Li, X., Zhang, T., Dubois, Y., Taori, R., Gulrajani, I., Guestrin, C., Liang, P., and Hashimoto, T. B · 2023
Cited alongside, same era.
Choi, H. K., Du, X., and Li, Y · 2024
Closest in time.
Towards secure tuning: Mitigating security risks arising from benign instruction fine-tuning
Du, Y., Zhao, S., Cao, J., Ma, M., Zhao, D., Fan, F., Liu, T., and Qin, B · 2024
Closest in time.
Mimicking user data: On mitigating fine-tuning risks in closed large language models
Eiras, F., Petrov, A., Torr, P. H., Kumar, M. P., and Bibi, A · 2024
Closest in time.
Mitigating forgetting in llm supervised fine-tuning and preference learning
Fernando, H., Shen, H., Ram, P., Zhou, Y., Samulowitz, H., Baracaldo, N., and Chen, T · 2024
Closest in time.
Covert malicious finetuning: Challenges in safeguarding llm adaptation
Halawi, D., Wei, A., Wallace, E., Wang, T. T., Haghtalab, N., and Steinhardt, J · 2024
Closest in time.
What’s in your" safe" data?: Identifying benign data that breaks safety
He, L., Xia, M., and Henderson, P · 2024
Closest in time.
Safe lora: the silver lining of reducing safety risks when fine-tuning large language models
Hsu, C.-Y., Tsai, Y.-L., Lin, C.-H., Chen, P.-Y., Yu, C.-M., and Huang, C.-Y · 2024
Closest in time.
Zipzap: Efficient training of language models for large-scale fraud detection on blockchain
Hu, S., Huang, T., Chow, K.-H., Wei, W., Wu, Y., and Liu, L · 2024
Closest in time.
No two devils alike: Unveiling distinct mechanisms of fine-tuning attacks
Leong, C. T., Cheng, Y., Xu, K., Wang, J., Wang, H., and Li, W · 2024
Closest in time.
Keeping llms aligned after fine-tuning: The crucial role of prompt templates
Lyu, K., Zhao, H., Gu, X., Yu, D., Goyal, A., and Arora, S · 2024
Closest in time.
Navigating the safety landscape: Measuring risks in finetuning large language models
Peng, S., Chen, P.-Y., Hull, M., and Chau, D. H · 2024
Closest in time.
Seal: Safety-enhanced aligned llm fine-tuning via bilevel data selection
Shen, H., Chen, P.-Y., Das, P., and Chen, T · 2024
Closest in time.
Tamper-resistant safeguards for open-weight llms
Tamirisa, R., Bharathi, B., Phan, L., Zhou, A., Gatti, A., Suresh, T., Lin, M., Wang, J., Wang, R., Arel, R., et al · 2024
Closest in time.
Mitigating fine-tuning jailbreak attack with backdoor enhanced alignment
Wang, J., Li, J., Li, Y., Qi, X., Chen, M., Hu, J., Li, Y., Li, B., and Xiao, C · 2024
Closest in time.
Emerging safety attack and defense in federated instruction tuning of large language models
Ye, R., Chai, J., Liu, X., Yang, Y., Wang, Y., and Chen, S · 2024
Closest in time.
On the vulnerability of safety alignment in open-access llms
Yi, J., Ye, R., Chen, Q., Zhu, B., Chen, S., Lian, D., Sun, G., Xie, X., and Wu, F · 2024
Closest in time.
Locking down the finetuned llms safety
Zhu, M., Yang, L., Wei, Y., Zhang, N., and Zhang, Y · 2024
Closest in time.
Safety fine-tuning at (almost) no cost: A baseline for vision large language models
Zong, Y., Bohdal, O., Yu, T., Yang, Y., and Hospedales, T · 2024
Closest in time.