Fetching the paper…
Reading the bibliography…
Safety alignment mechanisms in large language models prevent responses to harmful queries through learned refusal behavior, yet these same mechanisms impede legitimate research applications including cognitive modeling, adversarial testing, and security analysis.
Mikolov, T., Sutskever, I., Chen, K., Corrado, G. S., & Dean, J. “Distributed representations of words and phrases and their compositionality.” NeurIPS (2013)
2013
Earlier work this paper cites.
Zellers, R., et al. “HellaSwag: Can a Machine Really Finish Your Sentence?” ACL (2019)
2019
Earlier work this paper cites.
Hendrycks, D., et al. “Measuring Massive Multitask Language Understanding.” ICLR (2021)
2021
Earlier work this paper cites.
Cobbe, K., et al. “Training Verifiers to Solve Math Word Problems.” arXiv:2110.14168 (2021)
2021
Earlier work this paper cites.
Ouyang, L., et al. “Training language models to follow instructions with human feedback.” NeurIPS (2022)
2022
Earlier work this paper cites.
Bai, Y., et al. “Constitutional AI: Harmlessness from AI feedback.” arXiv:2212.08073 (2022)
2022
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
Rafailov, R., et al. “Direct preference optimization: Your language model is secretly a reward model.” NeurIPS (2023)
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Cited alongside, same era.
Tunstall, L., et al. “Zephyr: Direct Distillation of LM Alignment.” arXiv:2310.16944 (2023)
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2024
Cited alongside, same era.
Labonne, M. “Uncensor any LLM with abliteration.” HuggingFace Blog (2024)
Jain, S., et al. “What makes and breaks safety fine-tuning? A mechanistic study.” NeurIPS (2024)
2024
Later among the works it cites.
Zou, A., et al. “Improving alignment robustness with circuit breakers.” arXiv:2406.04313 (2024)
2024
Later among the works it cites.
Human-CentricAI. “LLM-Refusal-Classifier: RoBERTa-based refusal/disclaimer classifier.” Hugging Face Hub (2025). https://huggingface.co/Human-CentricAI/LLM-Refusal-Classifier
2025
Closest in time.
2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2024
Cited alongside, same era.
Lai, J. “Projected Abliteration: Norm-preserving biprojected abliteration.” HuggingFace Blog (2024). https://huggingface.co/blog/grimjim/projected-abliteration
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
p-e-w. “Heretic: Fully automatic censorship removal for language models.” GitHub. https://github.com/p-e-w/heretic
Cited in the paper.
FailSpy. “abliterator: Ablate features in transformer-based LLMs.” GitHub. https://github.com/FailSpy/abliterator
Cited in the paper.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.