Fetching the paper…
Reading the bibliography…
Diffusion models, while powerful, can inadvertently generate harmful or undesirable content, raising significant ethical and safety concerns.
Sparse coding with an overcomplete basis set: A strategy employed by v1?
Olshausen, B. A. and Field, D. J · 1997
Earlier work this paper cites.
Makhzani, A. and Frey, B. J · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P · 2014
Earlier work this paper cites.
Towards making systems forget with machine unlearning
Cao, Y. and Yang, J · 2015
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
Sohl-Dickstein, J., Weiss, E., Maheswaranathan, N., and Ganguli, S · 2015
Earlier work this paper cites.
Going deeper with convolutions
Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S., Anguelov, D., Erhan, D., Vanhoucke, V., and Rabinovich, A · 2015
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
Heusel, M., Ramsauer, H., Unterthiner, T., Nessler, B., and Hochreiter, S · 2017
Earlier work this paper cites.
Denoising diffusion probabilistic models
Ho, J., Jain, A., and Abbeel, P · 2020
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Earlier work this paper cites.
Red-teaming the stable diffusion safety filter
Rando, J., Paleka, D., Lindner, D., Heim, L., and Tramèr, F · 2022
Earlier work this paper cites.
High-resolution image synthesis with latent diffusion models
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B · 2022
Earlier work this paper cites.
LAION-5b: An open large-scale dataset for training next generation image-text models
Schuhmann, C., Beaumont, R., Vencu, R., Gordon, C. W., Wightman, R., Cherti, M., Coombes, T., Katta, A., Mullis, C., Wortsman, M., Schramowski, P., Kundurthy, S. R., Crowson, K., Schmidt, L., Kaczmarczyk, R., and Jitsev, J · 2022
Earlier work this paper cites.
Localizing and editing knowledge in text-to-image generative models
Basu, S., Zhao, N., Morariu, V. I., Feizi, S., and Manjunatha, V · 2023
Earlier work this paper cites.
Towards monosemanticity: Decomposing language models with dictionary learning
Bricken, T., Templeton, A., Batson, J., Chen, B., Jermyn, A., Conerly, T., Turner, N., Anil, C., Denison, C., Askell, A., Lasenby, R., Wu, Y., Kravec, S., Schiefer, N., Maxwell, T., Joseph, N., Hatfield-Dodds, Z., Tamkin, A., Nguyen, K., McLean, B., Burke, J. E., Hume, T., Carter, S., Henighan, T., and Olah, C · 2023
Earlier work this paper cites.
Erasing concepts from diffusion models
Gandikota, R., Materzynska, J., Fiotto-Kaufman, J., and Bau, D · 2023
Earlier work this paper cites.
Ablating concepts in text-to-image diffusion models
Kumari, N., Zhang, B., Wang, S.-Y., Shechtman, E., Zhang, R., and Zhu, J.-Y · 2023
Earlier work this paper cites.
Diffusion models already have a semantic latent space
Kwon, M., Jeong, J., and Uh, Y · 2023
Earlier work this paper cites.
Understanding the latent space of diffusion models through the lens of riemannian geometry
Park, Y.-H., Kwon, M., Choi, J., Jo, J., and Uh, Y · 2023
Earlier work this paper cites.
Safe latent diffusion: Mitigating inappropriate degeneration in diffusion models
Schramowski, P., Brack, M., Deiseroth, B., and Kersting, K · 2023
Cited alongside, same era.
What the daam: Interpreting stable diffusion using cross attention
Tang, R., Liu, L., Pandey, A., Jiang, Z., Yang, G., Kumar, K., Stenetorp, P., Lin, J., and Türe, F · 2023
Cited alongside, same era.
An x-ray is worth 15 features: Sparse autoencoders for interpretable radiology report generation
Abdulaal, A., Fry, H., Montaña-Brown, N., Ijishakin, A., Gao, J., Hyland, S., Alexander, D. C., and Castro, D. C · 2024
Cited alongside, same era.
Andersen v. stability ai ltd., 2024
Andersen · 2024
Cited alongside, same era.
On mechanistic knowledge localization in text-to-image generative models
Basu, S., Rezaei, K., Kattakinda, P., Morariu, V. I., Zhao, N., Rossi, R. A., Manjunatha, V., and Feizi, S · 2024
Cited alongside, same era.
H-space sparse autoencoders
Ijishakin, A., Ang, M. L., Baljer, L., Tan, D. C. H., Fry, H. L., Abdulaal, A., Lynch, A., and Cole, J. H · 2024
Later among the works it cites.
Interpreting attention layer outputs with sparse autoencoders
Kissane, C., Krzyzanowski, R., Bloom, J. I., Conmy, A., and Nanda, N · 2024
Later among the works it cites.
Get what you want, not what you don’t: Image content suppression for text-to-image diffusion models
Li, S., van de Weijer, J., taihang Hu, Khan, F., Hou, Q., Wang, Y., and jian Yang · 2024
Later among the works it cites.
Mace: Mass concept erasure in diffusion models
Lu, S., Wang, Z., Li, L., Liu, Y., and Kong, A. W.-K · 2024
Later among the works it cites.
One-dimensional adapter to rule them all: Concepts diffusion models and erasing applications
Lyu, M., Yang, Y., Hong, H., Chen, H., Jin, X., He, Y., Xue, H., Han, J., and Ding, G · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bussmann, B., Leask, P., and Nanda, N · 2024
Cited alongside, same era.
Case study: Interpreting, manipulating, and controlling clip with sparse autoencoders, 2024
Daujotas, G · 2024
Cited alongside, same era.
Salun: Empowering machine unlearning via gradient-based weight saliency in both image classification and generation
Fan, C., Liu, J., Zhang, Y., Wong, E., Wei, D., and Liu, S · 2024
Cited alongside, same era.
Applying sparse autoencoders to unlearn knowledge in language models
Farrell, E., Lau, Y.-T., and Conmy, A · 2024
Cited alongside, same era.
Towards multimodal interpretability: Learning sparse interpretable features in vision transformers, 2024
Fry, H · 2024
Cited alongside, same era.
Unified concept editing in diffusion models
Gandikota, R., Orgad, H., Belinkov, Y., Materzyńska, J., and Bau, D · 2024
Cited alongside, same era.
Reliable and efficient concept erasure of text-to-image diffusion models
Gong, C., Chen, K., Wei, Z., Chen, J., and Jiang, Y.-G · 2024
Cited alongside, same era.
Paulo, G., Mallen, A., Juang, C., and Belrose, N · 2024
Later among the works it cites.
Unpacking sdxl turbo: Interpreting text-to-image models with sparse autoencoders
Surkov, V., Wendler, C., Terekhov, M., Deschenaux, J., West, R., and Gulcehre, C · 2024
Later among the works it cites.
Scaling monosemanticity: Extracting interpretable features from claude 3 sonnet
Templeton, A., Conerly, T., Marcus, J., Lindsey, J., Bricken, T., Chen, B., Pearce, A., Citro, C., Ameisen, E., Jones, A., Cunningham, H., Turner, N. L., McDougall, C., MacDiarmid, M., Freeman, C. D., Sumers, T. R., Rees, E., Batson, J., Jermyn, A., Carter, S., Olah, C., and Henighan, T · 2024
Later among the works it cites.
Diffusion lens: Interpreting text encoders in text-to-image pipelines
Toker, M., Orgad, H., Ventura, M., Arad, D., and Belinkov, Y · 2024
Later among the works it cites.
Scissorhands: Scrub data influence via connection sensitivity in networks
Wu, J. and Harandi, M · 2024
Later among the works it cites.
Erasediff: Erasing data influence in diffusion models
Wu, J., Le, T., Hayat, M., and Harandi, M · 2024
Later among the works it cites.
Scaling and evaluating sparse autoencoders
Gao, L., la Tour, T. D., Tillman, H., Goh, G., Troll, R., Radford, A., Sutskever, I., Leike, J., and Wu, J · 2025
Closest in time.
Interpreting and editing vision-language representations to mitigate hallucinations
Jiang, N., Kachinthaya, A., Petryk, S., and Gandelsman, Y · 2025
Closest in time.
Concept steerers: Leveraging k-sparse autoencoders for controllable generations
Kim, D. and Ghadiyaram, D · 2025
Closest in time.
Concept pinpoint eraser for text-to-image diffusion models via residual attention gate
Lee, B. H., Lim, S., Lee, S., Kang, D. U., and Chun, S. Y · 2025
Closest in time.
Sparse feature circuits: Discovering and editing interpretable causal graphs in language models
Marks, S., Rager, C., Michaud, E. J., Belinkov, Y., Bau, D., and Mueller, A · 2025
Closest in time.
SAFREE: Training-free and adaptive guard for safe text-to-image and video generation
Yoon, J., Yu, S., Patil, V., Yao, H., and Bansal, M · 2025
Closest in time.
To generate or not? safety-driven unlearned diffusion models are still easy to generate unsafe images… for now
Zhang, Y., Jia, J., Chen, X., Chen, A., Zhang, Y., Liu, J., Ding, K., and Liu, S · 2025
Closest in time.