Fetching the paper…
Reading the bibliography…
Refusal behavior in large language models (LLMs) enables them to decline responding to harmful, unethical, or inappropriate prompts, ensuring alignment with ethical standards.
K. Pearson, “On lines and planes of closest fit to systems of points in space,” Philosophical Magazine , vol. 2, no. 11, pp. 559–572, 1901
1901
Earlier work this paper cites.
R. Blair, “The amygdala and ventromedial prefrontal cortex in morality and psychopathy,” Trends in Cognitive Sciences , vol. 11, no. 9, pp. 387–392, 2007, opinion article
2007
Earlier work this paper cites.
A. Glenn, A. Raine, and R. Schug, “The neural correlates of moral decision-making in psychopathy,” Molecular Psychiatry , vol. 14, pp. 5–6, 2009, published 19 December 2008. [Online]. Available: https://doi.org/10.1038/mp.2008.104
2008
Earlier work this paper cites.
L. van der Maaten and G. Hinton, “Visualizing data using t-SNE,” Journal of Machine Learning Research , vol. 9, pp. 2579–2605, 2008
2008
Earlier work this paper cites.
S. Kim and D. Lee, “Prefrontal cortex and impulsive decision making,” Biological Psychiatry , vol. 69, no. 12, pp. 1140–1146, 2011, epub 2010 Aug 21
2010
Earlier work this paper cites.
2015
Earlier work this paper cites.
G. Goh, “Decoding the thought vector,” https://gabgoh.github.io/ThoughtVectors/, 2016
2016
Earlier work this paper cites.
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
C. Olah, N. Cammarata, L. Schubert, G. Goh, M. Petrov, and S. Carter, “Zoom in: An introduction to circuits,” Distill , 2020, https://distill.pub/2020/circuits/zoom-in
2020
Earlier work this paper cites.
A. Schilling, A. Maier, R. Gerum, C. Metzner, and P. Krauss, “Quantifying the separability of data classes in neural networks,” Neural Networks , vol. 139, pp. 278–293, July 2021, epub 2021 Apr 5
2021
Earlier work this paper cites.
2022
Cited alongside, same era.
N. Nanda and J. Bloom, “Transformerlens,” https://github.com/TransformerLensOrg/TransformerLens, 2022
2022
Cited alongside, same era.
T. Bricken, A. Templeton, J. Batson, B. Chen, A. Jermyn, T. Conerly, N. Turner, C. Anil, C. Denison, A. Askell, R. Lasenby, Y. Wu, S. Kravec, N. Schiefer, T. Maxwell, N. Joseph, Z. Hatfield-Dodds, A. Tamkin, K. Nguyen, B. McLean, J. E. Burke, T. Hume, S. Carter, T. Henighan, and C. Olah, “Towards monosemanticity: Decomposing language models with dictionary learning,” Transformer Circuits Thread , 2023, https://transformer-circuits.pub/2023/monosemantic-features/index.html
2023
Cited alongside, same era.
2024
Later among the works it cites.
2024
Later among the works it cites.
M. Labonne, “harmless_alpaca,” https://huggingface.co/datasets/mlabonne/harmless_alpaca, 2024. [Online]. Available: https://huggingface.co/datasets/mlabonne/harmless_alpaca
2024
Later among the works it cites.
M. Labonne, “harmful_behaviors,” https://huggingface.co/datasets/mlabonne/harmful_behaviors, 2024. [Online]. Available: https://huggingface.co/datasets/mlabonne/harmful_behaviors
2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2023
Cited alongside, same era.
R. Taori, I. Gulrajani, T. Zhang, Y. Dubois, X. Li, C. Guestrin, P. Liang, and T. B. Hashimoto, “Stanford alpaca: An instruction-following llama model,” https://github.com/tatsu-lab/stanford_alpaca, 2023. [Online]. Available: https://github.com/tatsu-lab/stanford_alpaca
2023
Cited alongside, same era.
A. Zou, Z. Wang, J. Z. Kolter, and M. Fredrikson, “Universal and transferable adversarial attacks on aligned language models,” 2023
2023
Cited alongside, same era.
Meta, “Model cards and prompt formats - llama 3.2,” https://www.llama.com/docs/model-cards-and-prompt-formats/llama3_2, 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Later among the works it cites.
M. Labonne, “Uncensor any llm with abliteration,” Hugging Face Community Blog , June 2024, published June 13, 2024. [Online]. Available: https://huggingface.co/blog/mlabonne/abliteration
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
N. Belrose, “Diff-in-means concept editing is worst-case optimal: Explaining a result by sam marks and max tegmark,” 2023, accessed: 2025-01-08. [Online]. Available: https://blog.eleuther.ai/diff-in-means/
2025
Closest in time.
F. Hildebrandt, “Refusal-llms,” 2025, accessed: 2025-01-11. [Online]. Available: https://github.com/FabianHildebrandt/Refusal-LLMs
2025
Closest in time.