Fetching the paper…
Reading the bibliography…
Image editing aims to edit the given synthetic or real image to meet the specific requirements from users.
C. Wah, S. Branson, P. Welinder, P. Perona, and S. Belongie, “The caltech-ucsd birds-200-2011 dataset,” 2011
2011
Earlier work this paper cites.
A. Khosla, N. Jayadevaprakash, B. Yao, and F.-F. Li, “Novel dataset for fine-grained image categorization: Stanford dogs,” in
2011
Earlier work this paper cites.
A. Borji, D. N. Sihite, and L. Itti, “Salient object detection: A benchmark,” in
2012
Earlier work this paper cites.
I. J. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. C. Courville, and Y. Bengio, “Generative adversarial nets,” in
2014
Earlier work this paper cites.
D. P. Kingma and M. Welling, “Auto-encoding variational bayes,” in
2014
Earlier work this paper cites.
T. Lin, M. Maire, S. J. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick, “Microsoft COCO: common objects in context,” in
2014
Earlier work this paper cites.
M. Andriluka, L. Pishchulin, P. Gehler, and B. Schiele, “2d human pose estimation: New benchmark and state of the art analysis,” in
2014
Earlier work this paper cites.
O. Rippel, M. A. Gelbart, and R. P. Adams, “Learning ordered representations with nested dropout,” in
2014
Earlier work this paper cites.
L. Dinh, D. Krueger, and Y. Bengio, “NICE: non-linear independent components estimation,” in
2015
Earlier work this paper cites.
O. Ronneberger, P. Fischer, and T. Brox, “U-Net: Convolutional networks for biomedical image segmentation,” in
2015
Earlier work this paper cites.
Z. Liu, P. Luo, X. Wang, and X. Tang, “Deep learning face attributes in the wild,” in
2015
Earlier work this paper cites.
F. Yu, Y. Zhang, S. Song, A. Seff, and J. Xiao, “LSUN: construction of a large-scale image dataset using deep learning with humans in the loop,”
2015
Earlier work this paper cites.
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. S. Bernstein, A. C. Berg, and L. Fei-Fei, “Imagenet large scale visual recognition challenge,”
2015
Earlier work this paper cites.
L. Dinh, D. Krueger, and Y. Bengio, “NICE: non-linear independent components estimation,” in
2015
Earlier work this paper cites.
S. Xie and Z. Tu, “Holistically-nested edge detection,” in
2015
Earlier work this paper cites.
S. Liu and W. Deng, “Very deep convolutional neural network based image classification using small training sample size,” in
2015
Earlier work this paper cites.
B. Thomee, D. A. Shamma, G. Friedland, B. Elizalde, K. Ni, D. Poland, D. Borth, and L. Li, “YFCC100M: the new data in multimedia research,”
2016
Earlier work this paper cites.
L. Yu, P. Poirson, S. Yang, A. C. Berg, and T. L. Berg, “Modeling context in referring expressions,” in
2016
Earlier work this paper cites.
J. Johnson, A. Alahi, and L. Fei-Fei, “Perceptual losses for real-time style transfer and super-resolution,” in
2016
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” in
2017
Earlier work this paper cites.
A. van den Oord, O. Vinyals, and K. Kavukcuoglu, “Neural discrete representation learning,” in
2017
Earlier work this paper cites.
J. Wu, H. Zheng, B. Zhao, Y. Li, B. Yan, R. Liang, W. Wang, S. Zhou, G. Lin, Y. Fu
2017
Earlier work this paper cites.
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L. Li, D. A. Shamma, M. S. Bernstein, and L. Fei-Fei, “Visual genome: Connecting language and vision using crowdsourced dense image annotations,”
2017
Earlier work this paper cites.
L. Wang, H. Lu, Y. Wang, M. Feng, D. Wang, B. Yin, and X. Ruan, “Learning to detect salient objects with image-level supervision,” in
2017
Earlier work this paper cites.
M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter, “Gans trained by a two time-scale update rule converge to a local nash equilibrium,” in
2017
Earlier work this paper cites.
L. Dinh, J. Sohl-Dickstein, and S. Bengio, “Density estimation using real NVP,” in
2017
Earlier work this paper cites.
P. Isola, J. Zhu, T. Zhou, and A. A. Efros, “Image-to-image translation with conditional adversarial networks,” in
2017
Earlier work this paper cites.
J. Zhu, T. Park, P. Isola, and A. A. Efros, “Unpaired image-to-image translation using cycle-consistent adversarial networks,” in
2017
Earlier work this paper cites.
N. Shazeer, A. Mirhoseini, K. Maziarz, A. Davis, Q. V. Le, G. E. Hinton, and J. Dean, “Outrageously large neural networks: The sparsely-gated mixture-of-experts layer,” in
2017
Earlier work this paper cites.
P. Sharma, N. Ding, S. Goodman, and R. Soricut, “Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning,” in
2018
Earlier work this paper cites.
H. Caesar, J. R. R. Uijlings, and V. Ferrari, “Coco-stuff: Thing and stuff classes in context,” in
2018
Earlier work this paper cites.
A. Kuznetsova, H. Rom, N. Alldrin, J. R. R. Uijlings, I. Krasin, J. Pont-Tuset, S. Kamali, S. Popov, M. Malloci, T. Duerig, and V. Ferrari, “The open images dataset V4: unified image classification, object detection, and visual relationship detection at scale,”
2018
Earlier work this paper cites.
N. Xu, L. Yang, Y. Fan, D. Yue, Y. Liang, J. Yang, and T. Huang, “Youtube-vos: A large-scale video object segmentation benchmark,”
2018
Earlier work this paper cites.
T. Karras, T. Aila, S. Laine, and J. Lehtinen, “Progressive growing of gans for improved quality, stability, and variation,” in
2018
Earlier work this paper cites.
W. R. Tan, C. S. Chan, H. E. Aguirre, and K. Tanaka, “Improved artgan for conditional synthesis of natural image and artwork,”
2018
Earlier work this paper cites.
R. Zhang, P. Isola, A. A. Efros, E. Shechtman, and O. Wang, “The unreasonable effectiveness of deep features as a perceptual metric,” in
2018
Earlier work this paper cites.
H. Ding, X. Jiang, B. Shuai, A. Q. Liu, and G. Wang, “Context contrasted feature and gated multi-scale aggregation for scene segmentation,” in
2018
Earlier work this paper cites.
Y. Song and S. Ermon, “Generative modeling by estimating gradients of the data distribution,” in
2019
Earlier work this paper cites.
J. Devlin, M. Chang, K. Lee, and K. Toutanova, “BERT: pre-training of deep bidirectional transformers for language understanding,” in
2019
Earlier work this paper cites.
J. Li, C. Wang, H. Zhu, Y. Mao, H.-S. Fang, and C. Lu, “Crowdpose: Efficient crowded scenes pose estimation and a new benchmark,” in
2019
Earlier work this paper cites.
L. Yang, Y. Fan, and N. Xu, “Video instance segmentation,” in
2019
Earlier work this paper cites.
N. Zheng, X. Song, Z. Chen, L. Hu, D. Cao, and L. Nie, “Virtually trying on new clothing with arbitrary poses,” in
2019
Earlier work this paper cites.
A. Gupta, P. Dollár, and R. B. Girshick, “LVIS: A dataset for large vocabulary instance segmentation,” in
2019
Earlier work this paper cites.
T. Karras, S. Laine, and T. Aila, “A style-based generator architecture for generative adversarial networks,” in
2019
Earlier work this paper cites.
M. Li, Z. L. Lin, R. Mech, E. Yumer, and D. Ramanan, “Photo-sketching: Inferring contour drawings from images,” in
2019
Earlier work this paper cites.
J. Lee, K. Cho, and D. Kiela, “Countering language drift via visual grounding,” in
2019
Earlier work this paper cites.
H. Ding, X. Jiang, B. Shuai, A. Q. Liu, and G. Wang, “Semantic correlation promoted shape-variant context for segmentation,” in
2019
Earlier work this paper cites.
J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” in
2020
Earlier work this paper cites.
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu, “Exploring the limits of transfer learning with a unified text-to-text transformer,”
2020
Earlier work this paper cites.
C. Wu, Z. Lin, S. Cohen, T. Bui, and S. Maji, “Phrasecut: Language-based image segmentation in the wild,” in
2020
Earlier work this paper cites.
W. Cong, J. Zhang, L. Niu, L. Liu, Z. Ling, W. Li, and L. Zhang, “Dovenet: Deep image harmonization via domain verification,” in
2020
Earlier work this paper cites.
M. Bijelic, T. Gruber, F. Mannan, F. Kraus, W. Ritter, K. Dietmayer, and F. Heide, “Seeing through fog without seeing fog: Deep multimodal sensor fusion in unseen adverse weather,” in
2020
Earlier work this paper cites.
F. Yu, H. Chen, X. Wang, W. Xian, Y. Chen, F. Liu, V. Madhavan, and T. Darrell, “BDD100K: A diverse driving dataset for heterogeneous multitask learning,” in
2020
Earlier work this paper cites.
P. Zhu, R. Abdal, Y. Qin, and P. Wonka, “Improved stylegan embedding: Where are the good latents?”
2020
Earlier work this paper cites.
Y. Lu, S. Singhal, F. Strub, A. C. Courville, and O. Pietquin, “Countering language drift with seeded iterated learning,” in
2020
Earlier work this paper cites.
B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Barron, R. Ramamoorthi, and R. Ng, “Nerf: Representing scenes as neural radiance fields for view synthesis,” in
2020
Earlier work this paper cites.
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger
2020
Earlier work this paper cites.
C. Schuhmann, R. Vencu, R. Beaumont, R. Kaczmarczyk, C. Mullis, A. Katta, T. Coombes, J. Jitsev, and A. Komatsuzaki, “LAION-400M: open dataset of clip-filtered 400 million image-text pairs,”
2021
Earlier work this paper cites.
K. Desai, G. Kaul, Z. Aysola, and J. Johnson, “Redcaps: Web-curated image-text data created by the people, for the people,” in
2021
Earlier work this paper cites.
K. Srinivasan, K. Raman, J. Chen, M. Bendersky, and M. Najork, “Wit: Wikipedia-based image text dataset for multimodal multilingual machine learning,” in
2021
Earlier work this paper cites.
J. Song, C. Meng, and S. Ermon, “Denoising diffusion implicit models,” in
2021
Earlier work this paper cites.
P. Dhariwal and A. Q. Nichol, “Diffusion models beat gans on image synthesis,” in
2021
Earlier work this paper cites.
Z. Wang, “Score-based generative modeling through backward stochastic differential equations: Inversion and generation,”
2021
Earlier work this paper cites.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever, “Learning transferable visual models from natural language supervision,” in
2021
Earlier work this paper cites.
P. Esser, R. Rombach, and B. Ommer, “Taming transformers for high-resolution image synthesis,” in
2021
Earlier work this paper cites.
W. Wang, M. Feiszli, H. Wang, and D. Tran, “Unidentified video objects: A benchmark for dense, open-world segmentation,” in
2021
Earlier work this paper cites.
S. Choi, S. Park, M. Lee, and J. Choo, “VITON-HD: high-resolution virtual try-on via misalignment-aware normalization,” in
2021
Earlier work this paper cites.
K. Kärkkäinen and J. Joo, “Fairface: Face attribute dataset for balanced race, gender, and age for bias measurement and mitigation,” in
2021
Earlier work this paper cites.
H. Ding, C. Liu, S. Wang, and X. Jiang, “Vision-language transformer and query generation for referring segmentation,” in
2021
Earlier work this paper cites.
Z. Cao, G. Hidalgo, T. Simon, S. Wei, and Y. Sheikh, “Openpose: Realtime multi-person 2d pose estimation using part affinity fields,”
2021
Earlier work this paper cites.
M. Caron, H. Touvron, I. Misra, H. Jégou, J. Mairal, P. Bojanowski, and A. Joulin, “Emerging properties in self-supervised vision transformers,” in
2021
Earlier work this paper cites.
Y. Kasten, D. Ofri, O. Wang, and T. Dekel, “Layered neural atlases for consistent video editing,”
2021
Earlier work this paper cites.
Z. Teed and J. Deng, “RAFT: recurrent all-pairs field transforms for optical flow (extended abstract),” in
2021
Earlier work this paper cites.
C. Schuhmann, R. Beaumont, R. Vencu, C. Gordon, R. Wightman, M. Cherti, T. Coombes, A. Katta, C. Mullis, M. Wortsman, P. Schramowski, S. Kundurthy, K. Crowson, L. Schmidt, R. Kaczmarczyk, and J. Jitsev, “LAION-5B: an open large-scale dataset for training next generation image-text models,” in
2022
Earlier work this paper cites.
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, “High-resolution image synthesis with latent diffusion models,” in
2022
Earlier work this paper cites.
C. Saharia, W. Chan, S. Saxena, L. Li, J. Whang, E. L. Denton, S. K. S. Ghasemipour, R. G. Lopes, B. K. Ayan, T. Salimans, J. Ho, D. J. Fleet, and M. Norouzi, “Photorealistic text-to-image diffusion models with deep language understanding,” in
2022
Earlier work this paper cites.
A. Ramesh, P. Dhariwal, A. Nichol, C. Chu, and M. Chen, “Hierarchical text-conditional image generation with clip latents,”
2022
Earlier work this paper cites.
J. Ho, T. Salimans, A. A. Gritsenko, W. Chan, M. Norouzi, and D. J. Fleet, “Video diffusion models,” in
2022
Earlier work this paper cites.
J. Ho, W. Chan, C. Saharia, J. Whang, R. Gao, A. A. Gritsenko, D. P. Kingma, B. Poole, M. Norouzi, D. J. Fleet, and T. Salimans, “Imagen video: High definition video generation with diffusion models,”
2022
Earlier work this paper cites.
Z. Dong, P. Wei, and L. Lin, “DreamArtist: Towards controllable one-shot text-to-image generation via contrastive prompt-tuning,”
2022
Earlier work this paper cites.
E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen, “Lora: Low-rank adaptation of large language models,” in
2022
Earlier work this paper cites.
J. Ackermann and M. Li, “High-resolution image editing via multi-stage blended diffusion,”
2022
Earlier work this paper cites.
M. Brack, P. Schramowski, F. Friedrich, D. Hintersdorf, and K. Kersting, “The stable artist: Steering semantics in diffusion latent space,”
2022
Earlier work this paper cites.
D. Zhou, W. Wang, H. Yan, W. Lv, Y. Zhu, and J. Feng, “Magicvideo: Efficient video generation with latent diffusion models,”
2022
Earlier work this paper cites.
O. Bar-Tal, D. Ofri-Amar, R. Fridman, Y. Kasten, and T. Dekel, “Text2live: Text-driven layered image and video editing,” in
2022
Earlier work this paper cites.
A. Kazerouni, E. K. Aghdam, M. Heidari, R. Azad, M. Fayyaz, I. Hacihaliloglu, and D. Merhof, “Diffusion models for medical image analysis: A comprehensive survey,”
2022
Earlier work this paper cites.
C. Meng, Y. He, Y. Song, J. Song, J. Wu, J. Zhu, and S. Ermon, “Sdedit: Guided image synthesis and editing with stochastic differential equations,” in
2022
Earlier work this paper cites.
Y. Zhao, H. Ding, H. Huang, and N.-M. Cheung, “A closer look at few-shot image generation,” in
2022
Earlier work this paper cites.
J. Ho and T. Salimans, “Classifier-free diffusion guidance,”
2022
Earlier work this paper cites.
T. Karras, M. Aittala, T. Aila, and S. Laine, “Elucidating the design space of diffusion-based generative models,” in
2022
Earlier work this paper cites.
A. Q. Nichol, P. Dhariwal, A. Ramesh, P. Shyam, P. Mishkin, B. McGrew, I. Sutskever, and M. Chen, “GLIDE: towards photorealistic image generation and editing with text-guided diffusion models,” in
2022
Earlier work this paper cites.
J. Miao, X. Wang, Y. Wu, W. Li, X. Zhang, Y. Wei, and Y. Yang, “Large-scale video panoptic segmentation in the wild: A benchmark,” in
2022
Earlier work this paper cites.
Y. Zheng, H. Yang, T. Zhang, J. Bao, D. Chen, Y. Huang, L. Yuan, D. Chen, M. Zeng, and F. Wen, “General facial representation learning in a visual-linguistic manner,” in
2022
Earlier work this paper cites.
O. Avrahami, D. Lischinski, and O. Fried, “Blended diffusion for text-driven editing of natural images,” in
2022
Earlier work this paper cites.
A. Lugmayr, M. Danelljan, A. Romero, F. Yu, R. Timofte, and L. V. Gool, “RePaint: Inpainting using denoising diffusion probabilistic models,” in
2022
Earlier work this paper cites.
R. Gal, O. Patashnik, H. Maron, A. H. Bermano, G. Chechik, and D. Cohen-Or, “Stylegan-nada: Clip-guided domain adaptation of image generators,”
2022
Earlier work this paper cites.
G. Kim, T. Kwon, and J. C. Ye, “Diffusionclip: Text-guided diffusion models for robust image manipulation,” in
2022
Earlier work this paper cites.
K. Preechakul, N. Chatthee, S. Wizadwongsa, and S. Suwajanakorn, “Diffusion autoencoders: Toward a meaningful and decodable representation,” in
2022
Earlier work this paper cites.
M. Zhao, F. Bao, C. Li, and J. Zhu, “EGSDE: unpaired image-to-image translation via energy-guided stochastic differential equations,” in
2022
Earlier work this paper cites.
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. L. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, J. Schulman, J. Hilton, F. Kelton, L. Miller, M. Simens, A. Askell, P. Welinder, P. F. Christiano, J. Leike, and R. Lowe, “Training language models to follow instructions with human feedback,” in
2022
Earlier work this paper cites.
R. Ranftl, K. Lasinger, D. Hafner, K. Schindler, and V. Koltun, “Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer,”
2022
Earlier work this paper cites.
H. Xu, J. Zhang, J. Cai, H. Rezatofighi, and D. Tao, “Gmflow: Learning optical flow via global matching,” in
2022
Earlier work this paper cites.
B. Cheng, I. Misra, A. G. Schwing, A. Kirillov, and R. Girdhar, “Masked-attention mask transformer for universal image segmentation,” in
2022
Earlier work this paper cites.
P. Truong, M. Danelljan, F. Yu, and L. V. Gool, “Probabilistic warp consistency for weakly-supervised semantic correspondences,” in
2022
Earlier work this paper cites.
H. Ding, C. Liu, S. He, X. Jiang, and C. C. Loy, “MeViS: A large-scale benchmark for video segmentation with motion expressions,” in
2023
Cited alongside, same era.
N. Tumanyan, M. Geyer, S. Bagon, and T. Dekel, “Plug-and-play diffusion features for text-driven image-to-image translation,” in
2023
Cited alongside, same era.
M. Cao, X. Wang, Z. Qi, Y. Shan, X. Qie, and Y. Zheng, “MasaCtrl: Tuning-free mutual self-attention control for consistent image synthesis and editing,” in
2023
Cited alongside, same era.
G. Parmar, K. K. Singh, R. Zhang, Y. Li, J. Lu, and J. Zhu, “Zero-shot image-to-image translation,” in
2023
Cited alongside, same era.
B. Kawar, S. Zada, O. Lang, O. Tov, H. Chang, T. Dekel, I. Mosseri, and M. Irani, “Imagic: Text-based real image editing with diffusion models,” in
2023
Cited alongside, same era.
M. Wang, H. Ding, J. H. Liew, J. Liu, Y. Zhao, and Y. Wei, “SegRefiner: Towards model-agnostic segmentation refinement with discrete diffusion process,” in
2023
Later among the works it cites.
C. Lin, J. Gao, L. Tang, T. Takikawa, X. Zeng, X. Huang, K. Kreis, S. Fidler, M. Liu, and T. Lin, “Magic3d: High-resolution text-to-3d content creation,” in
2023
Later among the works it cites.
C. Liu, H. Ding, and X. Jiang, “GRES: Generalized referring expression segmentation,” in
2023
Later among the works it cites.
Y. Shi, P. Wang, J. Ye, M. Long, K. Li, and X. Yang, “MVDream: Multi-view diffusion for 3d generation,”
2023
Later among the works it cites.
Z. Wang, C. Lu, Y. Wang, F. Bao, C. Li, H. Su, and J. Zhu, “Prolificdreamer: High-fidelity and diverse text-to-3d generation with variational score distillation,” in
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
B. Poole, A. Jain, J. T. Barron, and B. Mildenhall, “Dreamfusion: Text-to-3d using 2d diffusion,” in
2023
Cited alongside, same era.
A. Haque, M. Tancik, A. A. Efros, A. Holynski, and A. Kanazawa, “Instruct-nerf2nerf: Editing 3d scenes with instructions,” in
2023
Cited alongside, same era.
G. Liu, M. Xia, Y. Zhang, H. Chen, J. Xing, X. Wang, Y. Yang, and Y. Shan, “StyleCrafter: Enhancing stylized text-to-video generation with style adapter,”
2023
Cited alongside, same era.
B. Yang, S. Gu, B. Zhang, T. Zhang, X. Chen, X. Sun, D. Chen, and F. Wen, “Paint by Example: Exemplar-based image editing with diffusion models,” in
2023
Cited alongside, same era.
A. Hertz, R. Mokady, J. Tenenbaum, K. Aberman, Y. Pritch, and D. Cohen-Or, “Prompt-to-prompt image editing with cross-attention control,” in
2023
Cited alongside, same era.
Y. Zhang, N. Huang, F. Tang, H. Huang, C. Ma, W. Dong, and C. Xu, “Inversion-based style transfer with diffusion models,” in
2023
Cited alongside, same era.
L. Zhang, A. Rao, and M. Agrawala, “Adding conditional control to text-to-image diffusion models,” in
2023
Cited alongside, same era.
K. Sohn, L. Jiang, J. Barber, K. Lee, N. Ruiz, D. Krishnan, H. Chang, Y. Li, I. Essa, M. Rubinstein, Y. Hao, G. Entis, I. Blok, and D. C. Chin, “StyleDrop: Text-to-image synthesis of any style,” in
2023
Later among the works it cites.
T. Li, M. Ku, C. Wei, and W. Chen, “DreamEdit: Subject-driven image editing,”
2023
Later among the works it cites.
A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W. Lo, P. Dollár, and R. B. Girshick, “Segment anything,” in
2023
Later among the works it cites.
H. Liu, C. Li, Q. Wu, and Y. J. Lee, “Visual instruction tuning,” in
2023
Later among the works it cites.
H. Ding, C. Liu, S. He, X. Jiang, P. H. S. Torr, and S. Bai, “MOSE: A new dataset for video object segmentation in complex scenes,” in
2023
Later among the works it cites.
A. Athar, J. Luiten, P. Voigtlaender, T. Khurana, A. Dave, B. Leibe, and D. Ramanan, “BURST: A benchmark for unifying object recognition, segmentation and tracking in video,” in
2023
Later among the works it cites.
X. Yu, M. Xu, Y. Zhang, H. Liu, C. Ye, Y. Wu, Z. Yan, C. Zhu, Z. Xiong, T. Liang, G. Chen, S. Cui, and X. Han, “Mvimgnet: A large-scale dataset of multi-view images,” in
2023
Later among the works it cites.
A. Hertz, A. Voynov, S. Fruchter, and D. Cohen-Or, “Style aligned image generation via shared attention,”
2023
Later among the works it cites.
W. Xia, Y. Zhang, Y. Yang, J. Xue, B. Zhou, and M. Yang, “GAN inversion: A survey,”
2023
Later among the works it cites.
Y. Gu, X. Wang, J. Z. Wu, Y. Shi, Y. Chen, Z. Fan, W. Xiao, R. Zhao, S. Chang, W. Wu, Y. Ge, Y. Shan, and M. Z. Shou, “Mix-of-show: Decentralized low-rank adaptation for multi-concept customization of diffusion models,” in
2023
Later among the works it cites.
B. Wallace, M. Dang, R. Rafailov, L. Zhou, A. Lou, S. Purushwalkam, S. Ermon, C. Xiong, S. Joty, and N. Naik, “Diffusion model alignment using direct preference optimization,”
2023
Later among the works it cites.
J. Li, D. Li, S. Savarese, and S. C. H. Hoi, “BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models,” in
2023
Later among the works it cites.
2023
Later among the works it cites.
S. Chen and J. Huang, “FEC: three finetuning-free methods to enhance consistency for real image editing,”
2023
Later among the works it cites.
G. Kwon and J. C. Ye, “Diffusion-based image translation using disentangled style and content representation,” in
2023
Later among the works it cites.
G. Y. Park, J. Kim, B. Kim, S. W. Lee, and J. C. Ye, “Energy-based cross attention for bayesian context update in text-to-image diffusion models,” in
2023
Later among the works it cites.
P. Ling, L. Chen, P. Zhang, H. Chen, and Y. Jin, “FreeDrag: Point tracking is not what you need for interactive point-based image editing,”
2023
Later among the works it cites.
X. Pan, A. Tewari, T. Leimkühler, L. Liu, A. Meka, and C. Theobalt, “Drag your GAN: interactive point-based manipulation on the generative image manifold,” in
2023
Later among the works it cites.
H. Ding, C. Liu, S. Wang, and X. Jiang, “VLT: Vision-language transformer and query generation for referring segmentation,”
2023
Later among the works it cites.
M. Oquab, T. Darcet, T. Moutakanni, H. Vo, M. Szafraniec, V. Khalidov, P. Fernandez, D. Haziza, F. Massa, A. El-Nouby
2023
Later among the works it cites.
S. Luo, Y. Tan, L. Huang, J. Li, and H. Zhao, “Latent Consistency Models: Synthesizing high-resolution images with few-step inference,”
2023
Later among the works it cites.
D. Valevski, D. Lumen, Y. Matias, and Y. Leviathan, “Face0: Instantaneously conditioning a text-to-image model on a face,” in
2023
Later among the works it cites.
N. Ruiz, Y. Li, V. Jampani, W. Wei, T. Hou, Y. Pritch, N. Wadhwa, M. Rubinstein, and K. Aberman, “HyperDreamBooth: Hypernetworks for fast personalization of text-to-image models,”
2023
Later among the works it cites.
W. Chen, H. Hu, C. Saharia, and W. W. Cohen, “Re-imagen: Retrieval-augmented text-to-image generator,” in
2023
Later among the works it cites.
S. Liu, Z. Zeng, T. Ren, F. Li, H. Zhang, J. Yang, C. Li, J. Yang, H. Su, J. Zhu, and L. Zhang, “Grounding DINO: marrying DINO with grounded pre-training for open-set object detection,”
2023
Later among the works it cites.
Y. Li, H. Liu, Q. Wu, F. Mu, J. Yang, J. Gao, C. Li, and Y. J. Lee, “GLIGEN: open-set grounded text-to-image generation,” in
2023
Later among the works it cites.
L. Khachatryan, A. Movsisyan, V. Tadevosyan, R. Henschel, Z. Wang, S. Navasardyan, and H. Shi, “Text2video-zero: Text-to-image diffusion models are zero-shot video generators,” in
2023
Later among the works it cites.
Z. Hu and D. Xu, “VideoControlNet: A motion-guided video-to-video translation framework by using diffusion model with controlnet,”
2023
Later among the works it cites.
M. Zhao, R. Wang, F. Bao, C. Li, and J. Zhu, “ControlVideo: Adding conditional control for one shot text-to-video editing,”
2023
Later among the works it cites.
S. Zhang, J. Wang, Y. Zhang, K. Zhao, H. Yuan, Z. Qin, X. Wang, D. Zhao, and J. Zhou, “I2VGen-XL: High-quality image-to-video synthesis via cascaded diffusion models,”
2023
Later among the works it cites.
H. Xu, J. Zhang, J. Cai, H. Rezatofighi, F. Yu, D. Tao, and A. Geiger, “Unifying flow, stereo and depth estimation,”
2023
Later among the works it cites.
W. Wang, K. Xie, Z. Liu, H. Chen, Y. Cao, X. Wang, and C. Shen, “Zero-shot video editing using off-the-shelf image diffusion models,”
2023
Later among the works it cites.
S. Yang, Y. Zhou, Z. Liu, and C. C. Loy, “Rerender A video: Zero-shot text-guided video-to-video translation,” in
2023
Later among the works it cites.
D. Bolya, C. Fu, X. Dai, P. Zhang, C. Feichtenhofer, and J. Hoffman, “Token merging: Your vit but faster,” in
2023
Later among the works it cites.
G. Y. Park, J. Kim, B. Kim, S. W. Lee, and J. C. Ye, “Energy-based cross attention for bayesian context update in text-to-image diffusion models,”
2024
Closest in time.
X. Chen, L. Huang, Y. Liu, Y. Shen, D. Zhao, and H. Zhao, “AnyDoor: Zero-shot object-level image customization,”
2024
Closest in time.
C. Liu, X. Li, and H. Ding, “Referring image editing: Object-level image editing via referring expressions,” in
2024
Closest in time.
Y. Chen, Z. Chen, C. Zhang, F. Wang, X. Yang, Y. Wang, Z. Cai, L. Yang, H. Liu, and G. Lin, “Gaussianeditor: Swift and controllable 3d editing with gaussian splatting,”
2024
Closest in time.
Y. Cong, M. Xu, C. Simon, S. Chen, J. Ren, Y. Xie, J. Pérez-Rúa, B. Rosenhahn, T. Xiang, and S. He, “FLATTEN: optical flow-guided attention for consistent text-to-video editing,”
2024
Closest in time.
S. Xu, Y. Huang, J. Pan, Z. Ma, and J. Chai, “Inversion-free image editing with natural language,”
2024
Closest in time.
M. Brack, F. Friedrich, K. Kornmeier, L. Tsaban, P. Schramowski, K. Kersting, and A. Passos, “LEDITS++: limitless image editing using text-to-image models,”
2024
Closest in time.
Y. Alaluf, D. Garibi, O. Patashnik, H. Averbuch-Elor, and D. Cohen-Or, “Cross-image attention for zero-shot appearance transfer,”
2024
Closest in time.
Y. Jia, Y. Yuan, A. Cheng, C. Wang, J. Li, H. Jia, and S. Zhang, “DesignEdit: Multi-layered latent decomposition and fusion for unified & accurate image editing,”
2024
Closest in time.
Y. Shi, C. Xue, J. Pan, W. Zhang, V. Y. F. Tan, and S. Bai, “DragDiffusion: Harnessing diffusion models for interactive point-based image editing,”
2024
Closest in time.
X. He, Z. Cao, N. Kolkin, L. Yu, H. Rhodin, and R. Kalarot, “A data perspective on enhanced identity preservation for diffusion personalization,”
2024
Closest in time.
K. Lee, S. Kwak, K. Sohn, and J. Shin, “Direct consistency optimization for compositional text-to-image personalization,”
2024
Closest in time.
P. Qiao, L. Shang, C. Liu, B. Sun, X. Ji, and J. Chen, “FaceChain-SuDe: Building derived class to inherit category attributes for one-shot subject-driven generation,”
2024
Closest in time.
Y. Cai, Y. Wei, Z. Ji, J. Bai, H. Han, and W. Zuo, “Decoupled textual embeddings for customized image generation,” in
2024
Closest in time.
L. Han, S. Wen, Q. Chen, Z. Zhang, K. Song, M. Ren, R. Gao, A. Stathopoulos, X. He, Y. Chen, D. Liu, Q. Zhangli, J. Jiang, Z. Xia, A. Srivastava, and D. Metaxas, “ProxEdit: Improving tuning-free real image editing with proximal guidance,” in
2024
Closest in time.
X. Ju, A. Zeng, Y. Bian, S. Liu, and Q. Xu, “PnP Inversion: Boosting diffusion-based editing with 3 lines of code,”
2024
Closest in time.
I. Huberman-Spiegelglas, V. Kulikov, and T. Michaeli, “An edit friendly DDPM noise space: Inversion and manipulations,”
2024
Closest in time.
Q. Guo and T. Lin, “Focus on your instruction: Fine-grained and multi-instruction image editing by attention modulation,”
2024
Closest in time.
B. Liu, C. Wang, T. Cao, K. Jia, and J. Huang, “Towards understanding cross and self-attention in stable diffusion for text-guided image editing,”
2024
Closest in time.
J. Nam, H. Kim, D. Lee, S. Jin, S. Kim, and S. Chang, “DreamMatcher: Appearance matching self-attention for semantically-consistent text-to-image personalization,”
2024
Closest in time.
J. Chung, S. Hyun, and J. Heo, “Style injection in diffusion: A training-free approach for adapting large-scale diffusion models for style transfer,”
2024
Closest in time.
Z. Yang, D. Gui, W. Wang, H. Chen, B. Zhuang, and C. Shen, “Object-aware inversion and reassembly for image editing,”
2024
Closest in time.
S. Li, B. Zeng, Y. Feng, S. Gao, X. Liu, J. Liu, L. Lin, X. Tang, Y. Hu, J. Liu, and B. Zhang, “ZONE: zero-shot instruction-guided local editing,”
2024
Closest in time.
P. Li, Q. Nie, Y. Chen, X. Jiang, K. Wu, Y. Lin, Y. Liu, J. Peng, C. Wang, and F. Zheng, “Tuning-free image customization with image and text guidance,”
2024
Closest in time.
X. Song, J. Cui, H. Zhang, J. Chen, R. Hong, and Y. Jiang, “Doubly abductive counterfactual inference for text-based image editing,”
2024
Closest in time.
H. Cho, J. Lee, S. B. Kim, T. Oh, and Y. Jeong, “Noise Map Guidance: Inversion with spatial context for real image editing,”
2024
Closest in time.
C. Mou, X. Wang, J. Song, Y. Shan, and J. Zhang, “DragonDiffusion: Enabling drag-style manipulation on diffusion models,”
2024
Closest in time.
S. Yang, L. Zhang, L. Ma, Y. Liu, J. Fu, and Y. He, “Magicremover: Tuning-free text-guided image inpainting with diffusion models,”
2024
Closest in time.
S. Mo, F. Mu, K. H. Lin, Y. Liu, B. Guan, Y. Li, and B. Zhou, “FreeControl: Training-free spatial control of any text-to-image diffusion model with any condition,”
2024
Closest in time.
C. Mou, X. Wang, J. Song, Y. Shan, and J. Zhang, “DiffEditor: Boosting accuracy and flexibility on diffusion-based image editing,”
2024
Closest in time.
H. Lv, J. Xiao, L. Li, and Q. Huang, “Pick-and-Draw: Training-free semantic guidance for text-to-image personalization,”
2024
Closest in time.
H. Nam, G. Kwon, G. Y. Park, and J. C. Ye, “Contrastive denoising score for text-guided latent diffusion image editing,”
2024
Closest in time.
H. Chang, J. Chang, and J. C. Ye, “Ground-A-Score: Scaling up the score distillation for multi-attribute editing,”
2024
Closest in time.
T. Fu, W. Hu, X. Du, W. Y. Wang, Y. Yang, and Z. Gan, “Guiding instruction-based image editing via multimodal large language models,”
2024
Closest in time.
Y. Huang, L. Xie, X. Wang, Z. Yuan, X. Cun, Y. Ge, J. Zhou, C. Dong, R. Huang, R. Zhang, and Y. Shan, “SmartEdit: Exploring complex instruction-based image editing with multimodal large language models,”
2024
Closest in time.
X. Zhang, J. Guo, P. Yoo, Y. Matsuo, and Y. Iwasawa, “Paste, inpaint and harmonize via denoising: Subject-driven image editing with pre-trained diffusion model,”
2024
Closest in time.
C. Mou, X. Wang, L. Xie, Y. Wu, J. Zhang, Z. Qi, and Y. Shan, “T2I-Adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models,” in
2024
Closest in time.
Z. Jiang, C. Mao, Y. Pan, Z. Han, and J. Zhang, “SCEdit: Efficient and controllable image diffusion generation via skip connection editing,”
2024
Closest in time.
G. Parmar, T. Park, S. Narasimhan, and J. Zhu, “One-step image translation with text-to-image models,”
2024
Closest in time.
X. Jia, Y. Zhao, K. C. K. Chan, Y. Li, H. Zhang, B. Gong, T. Hou, H. Wang, and Y. Su, “Taming encoder for zero fine-tuning image customization with text-to-image diffusion models,”
2024
Closest in time.
J. Shi, W. Xiong, Z. Lin, and H. J. Jung, “InstantBooth: Personalized text-to-image generation without test-time finetuning,”
2024
Closest in time.
Y. Zhou, R. Zhang, T. Sun, and J. Xu, “Enhancing detail preservation for customized text-to-image generation: A regularization-free approach,”
2024
Closest in time.
Q. Wang, X. Bai, H. Wang, Z. Qin, and A. Chen, “InstantID: Zero-shot identity-preserving generation in seconds,”
2024
Closest in time.
H. Hu, K. C. K. Chan, Y. Su, W. Chen, Y. Li, K. Sohn, Y. Zhao, X. Ben, B. Gong, W. W. Cohen, M. Chang, and X. Jia, “Instruct-Imagen: Image generation with multi-modal instruction,”
2024
Closest in time.
S. Lee, Y. Zhang, S. Wu, and J. Wu, “Language-informed visual concept learning,”
2024
Closest in time.
E. Richardson, Y. Alaluf, A. Mahdavi-Amiri, and D. Cohen-Or, “pOps: Photo-inspired diffusion operators,”
2024
Closest in time.
M. Ku, C. Wei, W. Ren, H. Yang, and W. Chen, “AnyV2V: A plug-and-play framework for any video-to-video editing tasks,”
2024
Closest in time.
F. Liang, B. Wu, J. Wang, L. Yu, K. Li, Y. Zhao, I. Misra, J. Huang, P. Zhang, P. Vajda, and D. Marculescu, “FlowVid: Taming imperfect optical flows for consistent video-to-video synthesis,”
2024
Closest in time.
C. Shin, H. Kim, C. H. Lee, S. Lee, and S. Yoon, “Edit-A-Video: single video editing with object-aware consistency,”
2024
Closest in time.
S. Liu, Y. Zhang, W. Li, Z. Lin, and J. Jia, “Video-P2P: Video editing with cross-attention control,”
2024
Closest in time.
B. Wu, C. Chuang, X. Wang, Y. Jia, K. Krishnakumar, T. Xiao, F. Liang, L. Yu, and P. Vajda, “Fairy: Fast parallelized instruction-guided video-to-video synthesis,”
2024
Closest in time.
X. Li, C. Ma, X. Yang, and M. Yang, “VidToMe: Video token merging for zero-shot video editing,”
2024
Closest in time.
L. Yang, Z. Zhang, Y. Song, S. Hong, R. Xu, Y. Zhao, W. Zhang, B. Cui, and M. Yang, “Diffusion models: A comprehensive survey of methods and applications,”
2024
Closest in time.
P. Cao, F. Zhou, Q. Song, and L. Yang, “Controllable generation with text-to-image diffusion models: A survey,”
2024
Closest in time.
B. B. Moser, A. S. Shanbhag, F. Raue, S. Frolov, S. Palacio, and A. Dengel, “Diffusion models, image super-resolution and everything: A survey,”
2024
Closest in time.
Y. Huang, J. Huang, Y. Liu, M. Yan, J. Lv, J. Liu, W. Xiong, H. Zhang, S. Chen, and L. Cao, “Diffusion model-based image editing: A survey,”
2024
Closest in time.
J. Chen, J. Yu, C. Ge, L. Yao, E. Xie, Y. Wu, Z. Wang, J. T. Kwok, P. Luo, H. Lu, and Z. Li, “Pixart-
2024
Closest in time.
X. Lai, Z. Tian, Y. Chen, Y. Li, Y. Yuan, S. Liu, and J. Jia, “Lisa: Reasoning segmentation via large language model,” in
2024
Closest in time.
G. Ilharco, M. Wortsman, R. Wightman, C. Gordon, N. Carlini, R. Taori, A. Dave, V. Shankar, H. Namkoong, J. Miller, H. Hajishirzi, A. Farhadi, and L. Schmidt, “Openclip,” Jan. 2024
2024
Closest in time.
H. Liu, C. Xu, Y. Yang, L. Zeng, and S. He, “Drag your noise: Interactive point-based editing via diffusion semantic propagation,”
2024
Closest in time.
N. Huang, Y. Zhang, F. Tang, C. Ma, H. Huang, Y. Zhang, W. Dong, and C. Xu, “DiffStyler: Controllable dual diffusion for text-driven image stylization,”
2024
Closest in time.
J. Wu, X. Li, S. Xu, H. Yuan, H. Ding, Y. Yang, X. Li, J. Zhang, Y. Tong, X. Jiang
2024
Closest in time.
Z. Chen, S. Fang, W. Liu, Q. He, M. Huang, and Z. Mao, “Dreamidentity: Enhanced editability for efficient face-identity preserved image generation,” in
2024
Closest in time.
H. Jeong and J. C. Ye, “Ground-A-Video: Zero-shot grounded video editing using text-to-image diffusion models,”
2024
Closest in time.