Fetching the paper…
Reading the bibliography…
Safety alignment is indispensable for Large Language Models (LLMs) to defend threats from malicious instructions.
Analysis of a complex of statistical variables into principal components
Hotelling, H. 1933 · 1933
Earlier work this paper cites.
Pointer Sentinel Mixture Models
Merity, S.; Xiong, C.; Bradbury, J.; and Socher, R. 2017 · 2017
Earlier work this paper cites.
Don’t Give Me the Details, Just the Summary! Topic-Aware Convolutional Neural Networks for Extreme Summarization
Narayan, S.; Cohen, S. B.; and Lapata, M. 2018 · 2018
Earlier work this paper cites.
Measuring massive multitask language understanding
Hendrycks, D.; Burns, C.; Basart, S.; Zou, A.; Mazeika, M.; Song, D.; and Steinhardt, J. 2020 · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C.; Shazeer, N.; Roberts, A.; Lee, K.; Narang, S.; Matena, M.; Zhou, Y.; Li, W.; and Liu, P. J. 2020 · 2020
Earlier work this paper cites.
Transformer Feed-Forward Layers Build Predictions by Promoting Concepts in the Vocabulary Space
Geva, M.; Caciularu, A.; Wang, K.; and Goldberg, Y. 2022 · 2022
Earlier work this paper cites.
TruthfulQA: Measuring How Models Mimic Human Falsehoods
Lin, S.; Hilton, J.; and Evans, O. 2022 · 2022
Earlier work this paper cites.
Achiam, J.; Adler, S.; Agarwal, S.; Ahmad, L.; Akkaya, I.; Aleman, F. L.; Almeida, D.; Altenschmidt, J.; Altman, S.; Anadkat, S.; et al. 2023 · 2023
Earlier work this paper cites.
Toxicity in chatgpt: Analyzing persona-assigned language models
Deshpande, A.; Murahari, V.; Rajpurohit, T.; Kalyan, A.; and Narasimhan, K. 2023 · 2023
Earlier work this paper cites.
Llama guard: Llm-based input-output safeguard for human-ai conversations
Inan, H.; Upasani, K.; Chi, J.; Rungta, R.; Iyer, K.; Mao, Y.; Tontchev, M.; Hu, Q.; Fuller, B.; Testuggine, D.; et al. 2023 · 2023
Earlier work this paper cites.
Pretraining language models with human preferences
Korbak, T.; Shi, K.; Chen, A.; Bhalerao, R. V.; Buckley, C.; Phang, J.; Bowman, S. R.; and Perez, E. 2023 · 2023
Earlier work this paper cites.
Inference-time intervention: eliciting truthful answers from a language model
Li, K.; Patel, O.; Viégas, F.; Pfister, H.; and Wattenberg, M. 2023 · 2023
Earlier work this paper cites.
A holistic approach to undesired content detection in the real world
Markov, T.; Zhang, C.; Agarwal, S.; Nekoul, F. E.; Lee, T.; Adler, S.; Jiang, A.; and Weng, L. 2023 · 2023
Earlier work this paper cites.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H.; Martin, L.; Stone, K.; Albert, P.; Almahairi, A.; Babaei, Y.; Bashlykov, N.; Batra, S.; Bhargava, P.; Bhosale, S.; et al. 2023 · 2023
Cited alongside, same era.
Activation addition: Steering language models without optimization
Turner, A.; Thiergart, L.; Udell, D.; Leech, G.; Mini, U.; and MacDiarmid, M. 2023 · 2023
Cited alongside, same era.
Introducing the next generation of Claude
Anthropic. 2024 · 2024
Cited alongside, same era.
Mitigating Exaggerated Safety in Large Language Models
Bhalani, R.; and Ray, R. 2024 · 2024
Cited alongside, same era.
Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow Instructions
Bianchi, F.; Suzgun, M.; Attanasio, G.; Rottger, P.; Jurafsky, D.; Hashimoto, T.; and Zou, J. 2024 · 2024
Steering Llama 2 via Contrastive Activation Addition
Rimsky, N.; Gabrieli, N.; Schulz, J.; Tong, M.; Hubinger, E.; and Turner, A. 2024 · 2024
Closest in time.
XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models
Röttger, P.; Kirk, H.; Vidgen, B.; Attanasio, G.; Bianchi, F.; and Hovy, D. 2024 · 2024
Closest in time.
Navigating the OverKill in Large Language Models
Shi, C.; Wang, X.; Ge, Q.; Gao, S.; Yang, X.; Gui, T.; Zhang, Q.; Huang, X.; Zhao, X.; and Lin, D. 2024 · 2024
Closest in time.
Trustllm: Trustworthiness in large language models
Sun, L.; Huang, Y.; Wang, H.; Wu, S.; Zhang, Q.; Gao, C.; Huang, Y.; Lyu, W.; Zhang, Y.; Li, X.; et al. 2024 · 2024
Closest in time.
Introducing Qwen1.5
Team, Q. 2024 · 2024
Closest in time.
The Art of Defending: A Systematic Evaluation and Analysis of LLM Defense Strategies on Safety and Over-Defensiveness
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
SCANS: Mitigating the exaggerated safety for llms via safety-conscious activation steering
Cao, Z.; Yang, Y.; and Zhao, H. 2024 · 2024
Cited alongside, same era.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
Chiang, W.-L.; Li, Z.; Lin, Z.; Sheng, Y.; Wu, Z.; Zhang, H.; Zheng, L.; Zhuang, S.; Zhuang, Y.; Gonzalez, J. E.; et al. 2023 · 2024
Cited alongside, same era.
OR-Bench: An Over-Refusal Benchmark for Large Language Models
Cui, J.; Chiang, W.-L.; Stoica, I.; and Hsieh, C.-J. 2024 · 2024
Cited alongside, same era.
Catastrophic Jailbreak of Open-source LLMs via Exploiting Generation
Huang, Y.; Gupta, S.; Xia, M.; Li, K.; and Chen, D. 2024 · 2024
Cited alongside, same era.
Perspective API
Jigsaw., G. 2017 · 2024
Cited alongside, same era.
Style Vectors for Steering Generative Large Language Models
Konen, K.; Jentzsch, S.; Diallo, D.; Schütt, P.; Bensch, O.; El Baff, R.; Opitz, D.; and Hecking, T. 2024 · 2024
Cited alongside, same era.
Open the Pandora’s Box of LLMs: Jailbreaking LLMs through Representation Engineering
Li, T.; Zheng, X.; and Huang, X. 2024 · 2024
Cited alongside, same era.
Varshney, N.; Dolin, P.; Seth, A.; and Baral, C. 2024 · 2024
Closest in time.
Wang, T.; Jiao, X.; He, Y.; Chen, Z.; Zhu, Y.; Chu, X.; Gao, J.; Wang, Y.; and Ma, L. 2024 · 2024
Closest in time.
GradSafe: Detecting Jailbreak Prompts for LLMs via Safety-Critical Gradient Analysis
Xie, Y.; Fang, M.; Pi, R.; and Gong, N. 2024 · 2024
Closest in time.
SafeDecoding: Defending against Jailbreak Attacks via Safety-Aware Decoding
Xu, Z.; Jiang, F.; Niu, L.; Jia, J.; Lin, B. Y.; and Poovendran, R. 2024 · 2024
Closest in time.
Zhao, W.; Hu, Y.; Li, Z.; Deng, Y.; Zhao, Y.; Qin, B.; and Chua, T.-S. 2024 · 2024
Closest in time.
On prompt-driven safeguarding for large language models
Zheng, C.; Yin, F.; Zhou, H.; Meng, F.; Zhou, J.; Chang, K.-W.; Huang, M.; and Peng, N. 2024 · 2024
Closest in time.
ROSE Doesn’t Do That: Boosting the Safety of Instruction-Tuned Large Language Models with Reverse Prompt Contrastive Decoding
Zhong, Q.; Ding, L.; Liu, J.; Du, B.; and Tao, D. 2024 · 2024
Closest in time.