Fetching the paper…
Reading the bibliography…
Rapid advances in the capabilities of large language models (LLMs) have raised widespread concerns regarding their potential for malicious use.
On information and sufficiency
S. Kullback and R. A. Leibler · 1951
Earlier work this paper cites.
Ethical and philosophical consideration of the dual-use dilemma in the biological sciences
S. Miller and M. J. Selgelid · 2007
Earlier work this paper cites.
ADADELTA: an adaptive learning rate method
M. D. Zeiler · 2012
Earlier work this paper cites.
Estimating or propagating gradients through stochastic neurons for conditional computation, 2013
Y. Bengio, N. Léonard, and A. Courville · 2013
Earlier work this paper cites.
SGDR: stochastic gradient descent with restarts
I. Loshchilov and F. Hutter · 2016
Earlier work this paper cites.
Pointer sentinel mixture models, 2016
S. Merity, C. Xiong, J. Bradbury, and R. Socher · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
P. F. Christiano, J. Leike, T. Brown, M. Martic, S. Legg, and D. Amodei · 2017
Earlier work this paper cites.
Model-agnostic meta-learning for fast adaptation of deep networks, 2017
C. Finn, P. Abbeel, and S. Levine · 2017
Earlier work this paper cites.
Snapshot ensembles: Train 1, get m for free, 2017
G. Huang, Y. Li, G. Pleiss, Z. Liu, J. E. Hopcroft, and K. Q. Weinberger · 2017
Earlier work this paper cites.
Adam: A method for stochastic optimization, 2017
D. P. Kingma and J. Ba · 2017
Earlier work this paper cites.
Understanding black-box predictions via influence functions
P. W. Koh and P. Liang · 2017
Earlier work this paper cites.
On first-order meta-learning algorithms, 2018
A. Nichol, J. Achiam, and J. Schulman · 2018
Earlier work this paper cites.
The woman worked as a babysitter: On biases in language generation
E. Sheng, K.-W. Chang, P. Natarajan, and N. Peng · 2019
Earlier work this paper cites.
The pile: An 800gb dataset of diverse text for language modeling, 2020
L. Gao, S. Biderman, S. Black, L. Golding, T. Hoppe, C. Foster, J. Phang, H. He, A. Thite, N. Nabeshima, S. Presser, and C. Leahy · 2020
Earlier work this paper cites.
Forgetting outside the box: Scrubbing deep networks of information accessible from input-output observations
A. Golatkar, A. Achille, and S. Soatto · 2020
Earlier work this paper cites.
The radicalization risks of gpt-3 and advanced neural language models, 2020
K. McGuffie and A. Newhouse · 2020
Earlier work this paper cites.
Zero: Memory optimizations toward training trillion parameter models
S. Rajbhandari, J. Rasley, O. Ruwase, and Y. He · 2020
Earlier work this paper cites.
Measuring massive multitask language understanding, 2021
D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models, 2021
E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen · 2021
Earlier work this paper cites.
ZeRO-Offload: Democratizing Billion-Scale model training
J. Ren, S. Rajbhandari, R. Y. Aminabadi, O. Ruwase, S. Yang, M. Zhang, D. Li, and Y. He · 2021
Earlier work this paper cites.
If influence functions are the answer, then what is the question?
J. Bae, N. Ng, A. Lo, M. Ghassemi, and R. B. Grosse · 2022
Earlier work this paper cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Y. Bai, A. Jones, K. Ndousse, A. Askell, A. Chen, N. DasSarma, D. Drain, S. Fort, D. Ganguli, T. Henighan, et al · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al · 2022
Cited alongside, same era.
Who’s harry potter? approximate unlearning in llms
R. Eldan and M. Russinovich · 2023
Cited alongside, same era.
Self-destructing models: Increasing the costs of harmful dual uses of foundation models, 2023
P. Henderson, E. Mitchell, C. D. Manning, D. Jurafsky, and C. Finn · 2023
Cited alongside, same era.
Llama guard: Llm-based input-output safeguard for human-ai conversations, 2023
H. Inan, K. Upasani, J. Chi, R. Rungta, K. Iyer, Y. Mao, M. Tontchev, Q. Hu, B. Fuller, D. Testuggine, and M. Khabsa · 2023
Unlearning bias in language models by partitioning gradients
C. Yu, S. Jeoung, A. Kasi, P. Yu, and H. Ji · 2023
Later among the works it cites.
Removing rlhf protections in gpt-4 via fine-tuning, 2023
Q. Zhan, R. Fang, R. Bindu, A. Gupta, T. Hashimoto, and D. Kang · 2023
Later among the works it cites.
Pytorch fsdp: Experiences on scaling fully sharded data parallel, 2023
Y. Zhao, A. Gu, R. Varma, L. Luo, C.-C. Huang, M. Xu, L. Wright, H. Shojanazeri, M. Ott, S. Shleifer, A. Desmaison, C. Balioglu, P. Damania, B. Nguyen, G. Chauhan, Y. Hao, A. Mathews, and S. Li · 2023
Later among the works it cites.
Llama 3 model card
AI@Meta · 2024
Closest in time.
Leace: Perfect linear concept erasure in closed form
N. Belrose, D. Schneider-Joseph, S. Ravfogel, R. Cotterell, E. Raff, and S. Biderman · 2024
Closest in time.
Ulma: Unified language model alignment with human demonstration and point-wise preference, 2024
T. Cai, X. Song, J. Jiang, F. Teng, J. Gu, and G. Zhang · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Baseline defenses for adversarial attacks against aligned language models, 2023
N. Jain, A. Schwarzschild, Y. Wen, G. Somepalli, J. Kirchenbauer, P. yeh Chiang, M. Goldblum, A. Saha, J. Geiping, and T. Goldstein · 2023
Cited alongside, same era.
Camel: Communicative agents for "mind" exploration of large language model society, 2023
G. Li, H. A. A. K. Hammoud, H. Itani, D. Khizbullin, and B. Ghanem · 2023
Cited alongside, same era.
Autodan: Generating stealthy jailbreak prompts on aligned large language models
X. Liu, N. Xu, M. Chen, and C. Xiao · 2023
Cited alongside, same era.
Gpt-4 technical report, 2023
OpenAI · 2023
Cited alongside, same era.
Fine-tuning aligned language models compromises safety, even when users do not intend to!, 2023
X. Qi, Y. Zeng, T. Xie, P.-Y. Chen, R. Jia, P. Mittal, and P. Henderson · 2023
Cited alongside, same era.
Direct preference optimization: Your language model is secretly a reward model
R. Rafailov, A. Sharma, E. Mitchell, S. Ermon, C. D. Manning, and C. Finn · 2023
Cited alongside, same era.
Smoothllm: Defending large language models against jailbreaking attacks, 2023
A. Robey, E. Wong, H. Hassani, and G. J. Pappas · 2023
Cited alongside, same era.
Closest in time.
Ctftime writeups archive
CTFtime · 2024
Closest in time.
The road less scheduled, 2024
A. Defazio, Xingyu, Yang, H. Mehta, K. Mishchenko, A. Khaled, and A. Cutkosky · 2024
Closest in time.
Sophon: Non-fine-tunable learning to restrain task transferability for pre-trained models, 2024
J. Deng, S. Pang, Y. Chen, L. Xia, Y. Bai, H. Weng, and W. Xu · 2024
Closest in time.
The wmdp benchmark: Measuring and reducing malicious use with unlearning, 2024
N. Li, A. Pan, A. Gopal, S. Yue, D. Berrios, A. Gatti, J. D. Li, A.-K. Dombrowski, S. Goel, L. Phan, G. Mukobi, N. Helm-Burger, R. Lababidi, L. Justen, A. B. Liu, M. Chen, I. Barrass, O. Zhang, X. Zhu, R. Tamirisa, B. Bharathi, A. Khoja, Z. Zhao, A. Herbert-Voss, C. B. Breuer, A. Zou, M. Mazeika, Z. Wang, P. Oswal, W. Liu, A. A. Hunt, J. Tienken-Harder, K. Y. Shih, K. Talley, J. Guan, R. Kaplan, I. Steneker, D. Campbell, B. Jokubaitis, A. Levinson, J. Wang, W. Qian, K. K. Karmakar, S. Basart, S. Fitz, M. Levine, P. Kumaraguru, U. Tupakula, V. Varadharajan, Y. Shoshitaishvili, J. Ba, K. M. Esvelt, A. Wang, and D. Hendrycks · 2024
Closest in time.
Rethinking machine unlearning for large language models
S. Liu, Y. Yao, J. Jia, S. Casper, N. Baracaldo, P. Hase, X. Xu, Y. Yao, H. Li, K. R. Varshney, et al · 2024
Closest in time.
The llama 3 herd of models, 2024
Llama Team, AI @ Meta · 2024
Closest in time.
Eight methods to evaluate robust unlearning in llms
A. Lynch, P. Guo, A. Ewart, S. Casper, and D. Hadfield-Menell · 2024
Closest in time.
Harmbench: A standardized evaluation framework for automated red teaming and robust refusal
M. Mazeika, L. Phan, X. Yin, A. Zou, Z. Wang, N. Mu, E. Sakhaee, N. Li, S. Basart, B. Li, D. Forsyth, and D. Hendrycks · 2024
Closest in time.
On evaluating the durability of safeguards for open-weight llms, 2024
X. Qi, B. Wei, N. Carlini, Y. Huang, T. Xie, L. He, M. Jagielski, M. Nasr, P. Mittal, and P. Henderson · 2024
Closest in time.
A. Sheshadri, A. Ewart, P. Guo, A. Lynch, C. Wu, V. Hebbar, H. Sleight, A. C. Stickland, E. Perez, D. Hadfield-Menell, and S. Casper · 2024
Closest in time.
Magpie: Alignment data synthesis from scratch by prompting aligned llms with nothing, 2024
Z. Xu, F. Jiang, L. Niu, Y. Deng, R. Poovendran, Y. Choi, and B. Y. Lin · 2024
Closest in time.
Rigorllm: Resilient guardrails for large language models against undesired content, 2024
Z. Yuan, Z. Xiong, Y. Zeng, N. Yu, R. Jia, D. Song, and B. Li · 2024
Closest in time.
Robust prompt optimization for defending language models against jailbreaking attacks, 2024
A. Zhou, B. Li, and H. Wang · 2024
Closest in time.
Improving alignment and robustness with short circuiting
A. Zou, L. Phan, J. Wang, D. Duenas, M. Lin, M. Andriushchenko, R. Wang, Z. Kolter, M. Fredrikson, and D. Hendrycks · 2024
Closest in time.