Fetching the paper…
Reading the bibliography…
Safety-aligned language models often exhibit fragile and imbalanced safety mechanisms, increasing the likelihood of generating unsafe content.
Plug and Play Language Models: A Simple Approach to Controlled Text Generation
Dathathri, S.; Madotto, A.; Lan, J.; Hung, J.; Frank, E.; Molino, P.; Yosinski, J.; and Liu, R. 2020 · 2020
Earlier work this paper cites.
A safety assessment model based on belief rule base with new optimization method
Feng, Z.; Zhou, Z.; Hu, C.; Ban, X.; and Hu, G. 2020 · 2020
Earlier work this paper cites.
RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models
Gehman, S.; Gururangan, S.; Sap, M.; Choi, Y.; and Smith, N. A. 2020 · 2020
Earlier work this paper cites.
Measuring Massive Multitask Language Understanding
Hendrycks, D.; Burns, C.; Basart, S.; Zou, A.; Mazeika, M.; Song, D.; and Steinhardt, J. 2021 · 2021
Earlier work this paper cites.
Ethical and social risks of harm from Language Models
Weidinger, L.; Mellor, J.; Rauh, M.; Griffin, C.; Uesato, J.; Huang, P.-S.; Cheng, M.; Glaese, M.; Balle, B.; Kasirzadeh, A.; Kenton, Z.; Brown, S.; Hawkins, W.; Stepleton, T.; Biles, C.; Birhane, A.; Haas, J.; Rimell, L.; Hendricks, L. A.; Isaac, W.; Legassick, S.; Irving, G.; and Gabriel, I. 2021 · 2021
Earlier work this paper cites.
FUDGE: Controlled Text Generation With Future Discriminators
Yang, K.; and Klein, D. 2021 · 2021
Earlier work this paper cites.
Constitutional AI: Harmlessness from AI Feedback
Bai, Y.; Kadavath, S.; Kundu, S.; Askell, A.; Kernion, J.; Jones, A.; Chen, A.; Goldie, A.; Mirhoseini, A.; McKinnon, C.; Chen, C.; Olsson, C.; Olah, C.; Hernandez, D.; Drain, D.; Ganguli, D.; Li, D.; Tran-Johnson, E.; Perez, E.; Kerr, J.; Mueller, J.; Ladish, J.; Landau, J.; Ndousse, K.; Lukosuite, K.; Lovitt, L.; Sellitto, M.; Elhage, N.; Schiefer, N.; Mercado, N.; DasSarma, N.; Lasenby, R.; Larson, R.; Ringer, S.; Johnston, S.; Kravec, S.; Showk, S. E.; Fort, S.; Lanham, T.; Telleen-Lawton, T.; Conerly, T.; Henighan, T.; Hume, T.; Bowman, S. R.; Hatfield-Dodds, Z.; Mann, B.; Amodei, D.; Joseph, N.; McCandlish, S.; Brown, T.; and Kaplan, J. 2022 · 2022
Earlier work this paper cites.
TruthfulQA: Measuring How Models Mimic Human Falsehoods
Lin, S.; Hilton, J.; and Evans, O. 2022 · 2022
Earlier work this paper cites.
NeuroLogic A*esque Decoding: Constrained Text Generation with Lookahead Heuristics
Lu, X.; Welleck, S.; West, P.; Jiang, L.; Kasai, J.; Khashabi, D.; Le Bras, R.; Qin, L.; Yu, Y.; Zellers, R.; Smith, N. A.; and Choi, Y. 2022 · 2022
Earlier work this paper cites.
Locating and Editing Factual Associations in GPT
Meng, K.; Bau, D.; Andonian, A.; and Belinkov, Y. 2022 · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Ouyang, L.; Wu, J.; Jiang, X.; Almeida, D.; Wainwright, C. L.; Mishkin, P.; Zhang, C.; Agarwal, S.; Slama, K.; Ray, A.; Schulman, J.; Hilton, J.; Kelton, F.; Miller, L.; Simens, M.; Askell, A.; Welinder, P.; Christiano, P.; Leike, J.; and Lowe, R. 2022 · 2022
Earlier work this paper cites.
COLD Decoding: Energy-based Constrained Text Generation with Langevin Dynamics
Qin, L.; Welleck, S.; Khashabi, D.; and Choi, Y. 2022 · 2022
Earlier work this paper cites.
Extracting Latent Steering Vectors from Pretrained Language Models
Subramani, N.; Suresh, N.; and Peters, M. E. 2022 · 2022
Earlier work this paper cites.
Accelerating Large Language Model Decoding with Speculative Sampling
Chen, C.; Borgeaud, S.; Irving, G.; Lespiau, J.-B.; Sifre, L.; and Jumper, J. 2023 · 2023
Earlier work this paper cites.
Controlled Text Generation via Language Model Arithmetic
Dekoninck, J.; Fischer, M.; Beurer-Kellner, L.; and Vechev, M. T. 2023 · 2023
Earlier work this paper cites.
Detoxifying Text with MaRCo: Controllable Revision with Experts and Anti-Experts
Hallinan, S.; Liu, A.; Choi, Y.; and Sap, M. 2023 · 2023
Cited alongside, same era.
Inspecting and Editing Knowledge Representations in Language Models
Hernandez, E.; Li, B. Z.; and Andreas, J. 2023 · 2023
Cited alongside, same era.
Baseline Defenses for Adversarial Attacks Against Aligned Language Models
Jain, N.; Schwarzschild, A.; Wen, Y.; Somepalli, G.; Kirchenbauer, J.; yeh Chiang, P.; Goldblum, M.; Saha, A.; Geiping, J.; and Goldstein, T. 2023 · 2023
Cited alongside, same era.
Jiang, A. Q.; Sablayrolles, A.; Mensch, A.; Bamford, C.; Chaplot, D. S.; de las Casas, D.; Bressand, F.; Lengyel, G.; Lample, G.; Saulnier, L.; Lavaud, L. R.; Lachaux, M.-A.; Stock, P.; Scao, T. L.; Lavril, T.; Wang, T.; Lacroix, T.; and Sayed, W. E. 2023 · 2023
Cited alongside, same era.
Critic-Guided Decoding for Controlled Text Generation
Controlled Text Generation via Language Model Arithmetic
Dekoninck, J.; Fischer, M.; Beurer-Kellner, L.; and Vechev, M. 2024 · 2024
Closest in time.
MASTERKEY: Automated Jailbreaking of Large Language Model Chatbots
Deng, G.; Liu, Y.; Li, Y.; Wang, K.; Zhang, Y.; Li, Z.; Wang, H.; Zhang, T.; and Liu, Y. 2024 · 2024
Closest in time.
Sowing the Wind, Reaping the Whirlwind: The Impact of Editing Language Models
Hazra, R.; Layek, S.; Banerjee, S.; and Poria, S. 2024 · 2024
Closest in time.
DeAL: Decoding-time Alignment for Large Language Models
Huang, J. Y.; Sengupta, S.; Bonadiman, D.; an Lai, Y.; Gupta, A.; Pappas, N.; Mansour, S.; Kirchhoff, K.; and Roth, D. 2024 · 2024
Closest in time.
DeepInception: Hypnotize Large Language Model to Be Jailbreaker
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kim, M.; Lee, H.; Yoo, K. M.; Park, J.; Lee, H.; and Jung, K. 2023 · 2023
Cited alongside, same era.
Holistic Evaluation of Language Models
Liang, P.; Bommasani, R.; Lee, T.; Tsipras, D.; Soylu, D.; Yasunaga, M.; Zhang, Y.; Narayanan, D.; Wu, Y.; Kumar, A.; Newman, B.; Yuan, B.; Yan, B.; Zhang, C.; Cosgrove, C. A.; Manning, C. D.; Re, C.; Acosta-Navas, D.; Hudson, D. A.; Zelikman, E.; Durmus, E.; Ladhak, F.; Rong, F.; Ren, H.; Yao, H.; WANG, J.; Santhanam, K.; Orr, L.; Zheng, L.; Yuksekgonul, M.; Suzgun, M.; Kim, N.; Guha, N.; Chatterji, N. S.; Khattab, O.; Henderson, P.; Huang, Q.; Chi, R. A.; Xie, S. M.; Santurkar, S.; Ganguli, S.; Hashimoto, T.; Icard, T.; Zhang, T.; Chaudhary, V.; Wang, W.; Li, X.; Mai, Y.; Zhang, Y.; and Koreeda, Y. 2023 · 2023
Cited alongside, same era.
PREADD: Prefix-Adaptive Decoding for Controlled Text Generation
Pei, J.; Yang, K.; and Klein, D. 2023 · 2023
Cited alongside, same era.
GEDI: GEnerative and DIscriminative Training for Self-Supervised Learning
Sansone, E.; and Manhaeve, R. 2023 · 2023
Cited alongside, same era.
On Second Thought, Let’s Not Think Step by Step! Bias and Toxicity in Zero-Shot Reasoning
Shaikh, O.; Zhang, H.; Held, W.; Bernstein, M.; and Yang, D. 2023 · 2023
Cited alongside, same era.
Llama 2: Open Foundation and Fine-Tuned Chat Models
Touvron, H.; Martin, L.; Stone, K.; Albert, P.; Almahairi, A.; Babaei, Y.; Bashlykov, N.; Batra, S.; Bhargava, P.; Bhosale, S.; Bikel, D.; Blecher, L.; Ferrer, C. C.; Chen, M.; Cucurull, G.; Esiobu, D.; Fernandes, J.; Fu, J.; Fu, W.; Fuller, B.; Gao, C.; Goswami, V.; Goyal, N.; Hartshorn, A.; Hosseini, S.; Hou, R.; Inan, H.; Kardas, M.; Kerkez, V.; Khabsa, M.; Kloumann, I.; Korenev, A.; Koura, P. S.; Lachaux, M.-A.; Lavril, T.; Lee, J.; Liskovich, D.; Lu, Y.; Mao, Y.; Martinet, X.; Mihaylov, T.; Mishra, P.; Molybog, I.; Nie, Y.; Poulton, A.; Reizenstein, J.; Rungta, R.; Saladi, K.; Schelten, A.; Silva, R.; Smith, E. M.; Subramanian, R.; Tan, X. E.; Tang, B.; Taylor, R.; Williams, A.; Kuan, J. X.; Xu, P.; Yan, Z.; Zarov, I.; Zhang, Y.; Fan, A.; Kambadur, M.; Narang, S.; Rodriguez, A.; Stojnic, R.; Edunov, S.; and Scialom, T. 2023 · 2023
Cited alongside, same era.
Faithfulness-Aware Decoding Strategies for Abstractive Summarization
Wan, D.; Liu, M.; McKeown, K.; Dreyer, M.; and Bansal, M. 2023 · 2023
Cited alongside, same era.
Aligning Large Language Models with Human: A Survey
Wang, Y.; Zhong, W.; Li, L.; Mi, F.; Zeng, X.; Huang, W.; Shang, L.; Jiang, X.; and Liu, Q. 2023 · 2023
Cited alongside, same era.
Li, X.; Zhou, Z.; Zhu, J.; Yao, J.; Liu, T.; and Han, B. 2024 · 2024
Closest in time.
Is This the Subspace You Are Looking for? An Interpretability Illusion for Subspace Activation Patching
Makelov, A.; Lange, G.; Geiger, A.; and Nanda, N. 2024 · 2024
Closest in time.
Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Qi, X.; Zeng, Y.; Xie, T.; Chen, P.-Y.; Jia, R.; Mittal, P.; and Henderson, P. 2024 · 2024
Closest in time.
XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models
Röttger, P.; Kirk, H. R.; Vidgen, B.; Attanasio, G.; Bianchi, F.; and Hovy, D. 2024 · 2024
Closest in time.
Stay on Topic with Classifier-Free Guidance
Sanchez, G.; Spangher, A.; Fan, H.; Levi, E.; Ammanamanchi, P. S.; and Biderman, S. 2024 · 2024
Closest in time.
Navigating the OverKill in Large Language Models
Shi, C.; Wang, X.; Ge, Q.; Gao, S.; Yang, X.; Gui, T.; Zhang, Q.; Huang, X.; Zhao, X.; and Lin, D. 2024 · 2024
Closest in time.
Function Vectors in Large Language Models
Todd, E.; Li, M. L.; Sharma, A. S.; Mueller, A.; Wallace, B. C.; and Bau, D. 2024 · 2024
Closest in time.
SafeDecoding: Defending against Jailbreak Attacks via Safety-Aware Decoding
Xu, Z.; Jiang, F.; Niu, L.; Jia, J.; Lin, B. Y.; and Poovendran, R. 2024 · 2024
Closest in time.
GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts
Yu, J.; Lin, X.; Yu, Z.; and Xing, X. 2024 · 2024
Closest in time.
Towards Best Practices of Activation Patching in Language Models: Metrics and Methods
Zhang, F.; and Nanda, N. 2024 · 2024
Closest in time.