Fetching the paper…
Reading the bibliography…
Do the rich representations of multi-modal diffusion transformers (DiTs) exhibit unique properties that enhance their interpretability? We introduce ConceptAttention, a novel method that leverages the expressive power of DiT attention layers to generate high-quality saliency maps that precisely locate textual concepts within images.
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer, September 2023
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 1910
Earlier work this paper cites.
Understanding and Improving Layer Normalization, November 2019
Xu, J., Sun, X., Zhang, Z., Zhao, G., and Lin, J · 1911
Earlier work this paper cites.
PyTorch: An Imperative Style, High-Performance Deep Learning Library, December 2019
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Köpf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., and Chintala, S · 1912
Earlier work this paper cites.
The Hungarian method for the assignment problem
Kuhn, H. W · 1931
Earlier work this paper cites.
Quantifying Attention Flow in Transformers, May 2020
Abnar, S. and Zuidema, W · 2005
Earlier work this paper cites.
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale, June 2021
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N · 2010
Earlier work this paper cites.
Transformer Interpretability Beyond Attention Visualization, April 2021
Chefer, H., Gur, S., and Wolf, L · 2012
Earlier work this paper cites.
ImageNet Auto-Annotation with Segmentation Propagation
Guillaumin, M., Küttel, D., and Ferrari, V · 2014
Earlier work this paper cites.
The Pascal Visual Object Classes Challenge: A Retrospective
Everingham, M., Eslami, S. M. A., Van Gool, L., Williams, C. K. I., Winn, J., and Zisserman, A · 2015
Earlier work this paper cites.
U-Net: Convolutional Networks for Biomedical Image Segmentation, May 2015
Ronneberger, O., Fischer, P., and Brox, T · 2015
Earlier work this paper cites.
Layer-wise Relevance Propagation for Neural Networks with Local Renormalization Layers, April 2016
Binder, A., Montavon, G., Bach, S., Müller, K.-R., and Samek, W · 2016
Earlier work this paper cites.
Grad-CAM: Visual Explanations from Deep Networks via Gradient-based Localization, December 2019
Selvaraju, R. R., Cogswell, M., Das, A., Vedantam, R., Parikh, D., and Batra, D · 2019
Earlier work this paper cites.
Grad-CAM: Visual Explanations from Deep Networks via Gradient-based Localization
Selvaraju, R. R., Cogswell, M., Das, A., Vedantam, R., Parikh, D., and Batra, D · 2020
Earlier work this paper cites.
Emerging Properties in Self-Supervised Vision Transformers, May 2021
Caron, M., Touvron, H., Misra, I., Jégou, H., Mairal, J., Bojanowski, P., and Joulin, A · 2021
Earlier work this paper cites.
Cho, J. H., Mall, U., Bala, K., and Hariharan, B · 2021
Earlier work this paper cites.
Learning Transferable Visual Models From Natural Language Supervision, February 2021
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I · 2021
Earlier work this paper cites.
SegDiff: Image Segmentation with Diffusion Probabilistic Models, September 2022
Amit, T., Shaharbany, T., Nachmani, E., and Wolf, L · 2022
Earlier work this paper cites.
Label-Efficient Semantic Segmentation with Diffusion Models, March 2022
Baranchuk, D., Rubachev, I., Voynov, A., Khrulkov, V., and Babenko, A · 2022
Cited alongside, same era.
Unsupervised Semantic Segmentation by Distilling Feature Correspondences, March 2022
Hamilton, M., Zhang, Z., Hariharan, B., Snavely, N., and Freeman, W. T · 2022
Cited alongside, same era.
Prompt-to-Prompt Image Editing with Cross Attention Control, August 2022
Hertz, A., Mokady, R., Tenenbaum, J., Aberman, K., Pritch, Y., and Cohen-Or, D · 2022
Cited alongside, same era.
Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow, September 2022
Liu, X., Gong, C., and Liu, Q · 2022
Cited alongside, same era.
High-Resolution Image Synthesis with Latent Diffusion Models, April 2022
FluxSpace: Disentangled Semantic Editing in Rectified Flow Transformers, December 2024
Dalva, Y., Venkatesh, K., and Yanardag, P · 2024
Later among the works it cites.
Vision Transformers Need Registers, April 2024
Darcet, T., Oquab, M., Mairal, J., and Bojanowski, P · 2024
Later among the works it cites.
Scaling Rectified Flow Transformers for High-Resolution Image Synthesis, March 2024
Esser, P., Kulal, S., Blattmann, A., Entezari, R., Müller, J., Saini, H., Levi, Y., Lorenz, D., Sauer, A., Boesel, F., Podell, D., Dockhorn, T., English, Z., Lacey, K., Goodwin, A., Marek, Y., and Rombach, R · 2024
Later among the works it cites.
Interpreting CLIP’s Image Representation via Text-Based Decomposition, March 2024
Gandelsman, Y., Efros, A. A., and Steinhardt, J · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B · 2022
Cited alongside, same era.
What the DAAM: Interpreting Stable Diffusion Using Cross Attention, December 2022
Tang, R., Liu, L., Pandey, A., Jiang, Z., Yang, G., Kumar, K., Stenetorp, P., Lin, J., and Ture, F · 2022
Cited alongside, same era.
iBOT: Image BERT Pre-Training with Online Tokenizer, January 2022
Zhou, J., Wei, C., Wang, H., Shen, W., Xie, C., Yuille, A., and Kong, T · 2022
Cited alongside, same era.
InstructPix2Pix: Learning to Follow Image Editing Instructions, January 2023
Brooks, T., Holynski, A., and Efros, A. A · 2023
Cited alongside, same era.
Extracting Training Data from Diffusion Models, January 2023
Carlini, N., Hayes, J., Nasr, M., Jagielski, M., Sehwag, V., Tramèr, F., Balle, B., Ippolito, D., and Wallace, E · 2023
Cited alongside, same era.
Attend-and-Excite: Attention-Based Semantic Guidance for Text-to-Image Diffusion Models
Chefer, H., Alaluf, Y., Vinker, Y., Wolf, L., and Cohen-Or, D · 2023
Cited alongside, same era.
Training-Free Layout Control with Cross-Attention Guidance, November 2023
Chen, M., Laina, I., and Vedaldi, A · 2023
Cited alongside, same era.
Diffusion Self-Guidance for Controllable Image Generation, June 2023
Epstein, D., Jabri, A., Poole, B., Efros, A. A., and Holynski, A · 2023
Cited alongside, same era.
Gupta, G., Yadav, K., Gal, Y., Batra, D., Kira, Z., Lu, C., and Rudner, T. G. J · 2024
Later among the works it cites.
Kadkhodaie, Z., Guth, F., Simoncelli, E. P., and Mallat, S · 2024
Later among the works it cites.
Diffusion Models for Open-Vocabulary Segmentation, September 2024
Karazija, L., Laina, I., Vedaldi, A., and Rupprecht, C · 2024
Later among the works it cites.
Liu, B., Wang, C., Cao, T., Jia, K., and Huang, J · 2024
Later among the works it cites.
Marcos-Manchón, P., Alcover-Couso, R., SanMiguel, J. C., and Martínez, J. M · 2024
Later among the works it cites.
CONFORM: Contrast is All You Need for High-Fidelity Text-to-Image Diffusion Models
Meral, T. H. S., Simsar, E., Tombari, F., and Yanardag, P · 2024
Later among the works it cites.
DINOv2: Learning Robust Visual Features without Supervision, February 2024
Oquab, M., Darcet, T., Moutakanni, T., Vo, H., Szafraniec, M., Khalidov, V., Fernandez, P., Haziza, D., Massa, F., El-Nouby, A., Assran, M., Ballas, N., Galuba, W., Howes, R., Huang, P.-Y., Li, S.-W., Misra, I., Rabbat, M., Sharma, V., Synnaeve, G., Xu, H., Jegou, H., Mairal, J., Labatut, P., Joulin, A., and Bojanowski, P · 2024
Later among the works it cites.
SAM 2: Segment Anything in Images and Videos, October 2024
Ravi, N., Gabeur, V., Hu, Y.-T., Hu, R., Ryali, C., Ma, T., Khedr, H., Rädle, R., Rolland, C., Gustafson, L., Mintun, E., Pan, J., Alwala, K. V., Carion, N., Wu, C.-Y., Girshick, R., Dollár, P., and Feichtenhofer, C · 2024
Later among the works it cites.
CLIP as RNN: Segment Countless Visual Concepts without Training Endeavor, May 2024
Sun, S., Li, R., Torr, P., Gu, X., and Li, S · 2024
Later among the works it cites.
Diffuse, Attend, and Segment: Unsupervised Zero-Shot Segmentation using Stable Diffusion, April 2024
Tian, J., Aggarwal, L., Colaco, A., Kira, Z., and Gonzalez-Franco, M · 2024
Later among the works it cites.
Diffusion Models Learn Low-Dimensional Distributions via Subspace Clustering, December 2024
Wang, P., Zhang, H., Zhang, Z., Chen, S., Ma, Y., and Qu, Q · 2024
Later among the works it cites.
CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer, March 2025
Yang, Z., Teng, J., Zheng, W., Ding, M., Huang, S., Xu, J., Yang, Y., Hong, W., Zhang, X., Feng, G., Yin, D., Zhang, Y., Wang, W., Cheng, Y., Xu, B., Gu, X., Dong, Y., and Tang, J · 2025
Closest in time.