Fetching the paper…
Reading the bibliography…
Diffusion models have revolted the field of text-to-image generation recently.
Multi-concept customization of text-to-image diffusion,
N. Kumari, B. Zhang, R. Zhang, E. Shechtman, J.-Y. Zhu, · 1941
Earlier work this paper cites.
The hungarian method for the assignment problem,
H. W. Kuhn, · 1955
Earlier work this paper cites.
Normalized cuts and image segmentation,
J. Shi, J. Malik, · 2000
Earlier work this paper cites.
The pascal visual object classes (VOC) challenge,
M. Everingham, L. V. Gool, C. K. I. Williams, J. M. Winn, A. Zisserman, · 2010
Earlier work this paper cites.
Efficient inference in fully connected crfs with gaussian edge potentials,
P. Krähenbühl, V. Koltun, · 2011
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality,
T. Mikolov, I. Sutskever, K. Chen, G. S. Corrado, J. Dean, · 2013
Earlier work this paper cites.
Anomaly detection and localization in crowded scenes,
W. Li, V. Mahadevan, N. Vasconcelos, · 2014
Earlier work this paper cites.
Microsoft COCO: common objects in context,
T. Lin, M. Maire, S. J. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, C. L. Zitnick, · 2014
Earlier work this paper cites.
Auto-encoding variational bayes,
D. P. Kingma, M. Welling, · 2014
Earlier work this paper cites.
Fully convolutional networks for semantic segmentation,
J. Long, E. Shelhamer, T. Darrell, · 2015
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation,
O. Ronneberger, P. Fischer, T. Brox, · 2015
Earlier work this paper cites.
Learning deep features for discriminative localization,
B. Zhou, A. Khosla, À. Lapedriza, A. Oliva, A. Torralba, · 2016
Earlier work this paper cites.
Attention is all you need,
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, I. Polosukhin, · 2017
Earlier work this paper cites.
Mask r-cnn,
K. He, G. Gkioxari, P. Dollár, R. B. Girshick, · 2017
Earlier work this paper cites.
Grad-cam: Visual explanations from deep networks via gradient-based localization,
R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh, D. Batra, · 2017
Earlier work this paper cites.
Affordancenet: An end-to-end deep learning approach for object affordance detection,
T. Do, A. Nguyen, I. D. Reid, · 2018
Earlier work this paper cites.
Film: Visual reasoning with a general conditioning layer,
E. Perez, F. Strub, H. de Vries, V. Dumoulin, A. C. Courville, · 2018
Earlier work this paper cites.
Learning pixel-level semantic affinity with image-level supervision for weakly supervised semantic segmentation,
J. Ahn, S. Kwak, · 2018
Earlier work this paper cites.
Youtube-vos: Sequence-to-sequence video object segmentation,
N. Xu, L. Yang, Y. Fan, J. Yang, D. Yue, Y. Liang, B. L. Price, S. Cohen, T. S. Huang, · 2018
Earlier work this paper cites.
Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs,
L. Chen, G. Papandreou, I. Kokkinos, K. Murphy, A. L. Yuille, · 2018
Earlier work this paper cites.
Aleatoric uncertainty estimation with test-time augmentation for medical image segmentation with convolutional neural networks,
G. Wang, W. Li, M. Aertsen, J. Deprest, S. Ourselin, T. Vercauteren, · 2019
Earlier work this paper cites.
Generative modeling by estimating gradients of the data distribution,
Y. Song, S. Ermon, · 2019
Earlier work this paper cites.
Zero-shot semantic segmentation,
M. Bucher, T. Vu, M. Cord, P. Pérez, · 2019
Cited alongside, same era.
Composing text and image for image retrieval - an empirical odyssey,
N. Vo, L. Jiang, C. Sun, K. Murphy, L. Li, L. Fei-Fei, J. Hays, · 2019
Cited alongside, same era.
Weakly supervised learning of instance segmentation with inter-pixel relations,
J. Ahn, S. Cho, S. Kwak, · 2019
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library,
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, A. Desmaison, A. Köpf, E. Z. Yang, Z. DeVito, M. Raison, A. Tejani, S. Chilamkurthy, B. Steiner, L. Fang, J. Bai, S. Chintala, · 2019
Cited alongside, same era.
Denoising diffusion probabilistic models,
J. Ho, A. Jain, P. Abbeel, · 2020
Cited alongside, same era.
Weakly-supervised semantic segmentation via sub-category exploration,
Diffusion models for implicit image segmentation ensembles,
J. Wolleb, R. Sandkühler, F. Bieder, P. Valmaggia, P. C. Cattin, · 2022
Later among the works it cites.
Fashionvlp: Vision language transformer for fashion retrieval with feedback,
S. Goenka, Z. Zheng, A. Jaiswal, R. Chada, Y. Wu, V. Hedau, P. Natarajan, · 2022
Later among the works it cites.
LAION-5b: An open large-scale dataset for training next generation image-text models,
C. Schuhmann, R. Beaumont, R. Vencu, C. W. Gordon, R. Wightman, M. Cherti, T. Coombes, A. Katta, C. Mullis, M. Wortsman, P. Schramowski, S. R. Kundurthy, K. Crowson, L. Schmidt, R. Kaczmarczyk, J. Jitsev, · 2022
Later among the works it cites.
Weakly supervised semantic segmentation using out-of-distribution data,
J. Lee, S. J. Oh, S. Yun, J. Choe, E. Kim, S. Yoon, · 2022
Later among the works it cites.
Multi-class token transformer for weakly supervised semantic segmentation,
L. Xu, W. Ouyang, M. Bennamoun, F. Boussaïd, D. Xu, · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Chang, Q. Wang, W. Hung, R. Piramuthu, Y. Tsai, M. Yang, · 2020
Cited alongside, same era.
Self-supervised equivariant attention mechanism for weakly supervised semantic segmentation,
Y. Wang, J. Zhang, M. Kan, S. Shan, X. Chen, · 2020
Cited alongside, same era.
Learning integral objects with intra-class discriminator for weakly-supervised semantic segmentation,
J. Fan, Z. Zhang, C. Song, T. Tan, · 2020
Cited alongside, same era.
Weakly supervised semantic segmentation with boundary exploration,
L. Chen, W. Wu, C. Fu, X. Han, Y. Zhang, · 2020
Cited alongside, same era.
Learning transferable visual models from natural language supervision,
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, I. Sutskever, · 2021
Cited alongside, same era.
Erase then grow: Generating correct class activation maps for weakly-supervised semantic segmentation,
Y. Chong, X. Chen, Y. Tao, S. Pan, · 2021
Cited alongside, same era.
Open-vocabulary object detection using captions,
A. Zareian, K. D. Rosa, D. H. Hu, S. Chang, · 2021
Cited alongside, same era.
Self-supervised image-specific prototype exploration for weakly supervised semantic segmentation,
Q. Chen, L. Yang, J. Lai, X. Xie, · 2022
Later among the works it cites.
Threshold matters in WSSS: manipulating the activation for the robust and accurate segmentation model against thresholds,
M. Lee, D. Kim, H. Shim, · 2022
Later among the works it cites.
Deep vit features as dense visual descriptors,
S. Amir, Y. Gandelsman, S. Bagon, T. Dekel, · 2022
Later among the works it cites.
CLIP is also an efficient segmenter: A text-driven approach for weakly supervised semantic segmentation,
Y. Lin, M. Chen, W. Wang, B. Wu, K. Li, B. Lin, H. Liu, X. He, · 2023
Closest in time.
Open-vocabulary panoptic segmentation with text-to-image diffusion models,
J. Xu, S. Liu, A. Vahdat, W. Byeon, X. Wang, S. D. Mello, · 2023
Closest in time.
Segment everything everywhere all at once,
X. Zou, J. Yang, H. Zhang, F. Li, L. Li, J. Wang, L. Wang, J. Gao, Y. J. Lee, · 2023
Closest in time.
Weakly supervised semantic segmentation via self-supervised destruction learning,
J. Li, Z. Jie, X. Wang, Y. Zhou, L. Ma, J. Jiang, · 2023
Closest in time.
Your diffusion model is secretly a zero-shot classifier,
A. C. Li, M. Prabhudesai, S. Duggal, E. Brown, D. Pathak, · 2023
Closest in time.
Open-vocabulary object segmentation with diffusion models,
Z. Li, Q. Zhou, X. Zhang, Y. Zhang, Y. Wang, W. Xie, · 2023
Closest in time.
Diffumask: Synthesizing images with pixel-level annotations for semantic segmentation using diffusion models,
W. Wu, Y. Zhao, M. Z. Shou, H. Zhou, C. Shen, · 2023
Closest in time.
Pic2word: Mapping pictures to words for zero-shot composed image retrieval,
K. Saito, K. Sohn, X. Zhang, C. Li, C. Lee, K. Saenko, T. Pfister, · 2023
Closest in time.
Localizing object-level shape variations with text-to-image diffusion models,
O. Patashnik, D. Garibi, I. Azuri, H. Averbuch-Elor, D. Cohen-Or, · 2023
Closest in time.
Prompt-to-prompt image editing with cross-attention control,
A. Hertz, R. Mokady, J. Tenenbaum, K. Aberman, Y. Pritch, D. Cohen-or, · 2023
Closest in time.
Ernie-vilg 2.0: Improving text-to-image diffusion model with knowledge-enhanced mixture-of-denoising-experts,
Z. Feng, Z. Zhang, X. Yu, Y. Fang, L. Li, X. Chen, Y. Lu, J. Liu, W. Yin, S. Feng, Y. Sun, L. Chen, H. Tian, H. Wu, H. Wang, · 2023
Closest in time.
2023
Closest in time.