Fetching the paper…

Mitigating harm in language models with conditional-likelihood filtration · Around