Fetching the paper…
Reading the bibliography…
Diffusion models have recently motivated great success in many generation tasks like object removal.
A universal image quality index
Zhou Wang and Alan C Bovik · 2002
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P. Kingma and Max Welling · 2014
Earlier work this paper cites.
Context encoders: Feature learning by inpainting
Deepak Pathak, Philipp Krahenbuhl, Jeff Donahue, Trevor Darrell, and Alexei A Efros · 2016
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter · 2017
Earlier work this paper cites.
Learning to generate images with perceptual similarity metrics
Jake Snell, Karl Ridgeway, Renjie Liao, Brett D Roads, Michael C Mozer, and Richard S Zemel · 2017
Earlier work this paper cites.
Image inpainting for irregular holes using partial convolutions
Guilin Liu, Fitsum A Reda, Kevin J Shih, Ting-Chun Wang, Andrew Tao, and Bryan Catanzaro · 2018
Earlier work this paper cites.
Fast and robust segmentation of white blood cell images by self-supervised learning
Xin Zheng, Yong Wang, Guoyou Wang, and Jianguo Liu · 2018
Earlier work this paper cites.
Machine learning approach of automatic identification and counting of blood cells
Mohammad Mahmudul Alam and Mohammad Tariqul Islam · 2019
Earlier work this paper cites.
An improved method for semantic image inpainting with gans: progressive inpainting
Yizhen Chen and Haifeng Hu · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu · 2020
Earlier work this paper cites.
Occluded prohibited items detection: An x-ray security inspection benchmark and de-occlusion attention module
Yanlu Wei, Renshuai Tao, Zhangjie Wu, Yuqing Ma, Libo Zhang, and Xianglong Liu · 2020
Earlier work this paper cites.
Segdiff: Image segmentation with diffusion probabilistic models
Tomer Amit, Tal Shaharbany, Eliya Nachmani, and Lior Wolf · 2021
Earlier work this paper cites.
Split then refine: stacked attention-guided resunets for blind single image visible watermark removal
Xiaodong Cun and Chi-Man Pun · 2021
Earlier work this paper cites.
Image segmentation using deep learning: A survey
Shervin Minaee, Yuri Boykov, Fatih Porikli, Antonio Plaza, Nasser Kehtarnavaz, and Demetri Terzopoulos · 2021
Earlier work this paper cites.
Encoding in style: a stylegan encoder for image-to-image translation
Elad Richardson, Yuval Alaluf, Or Patashnik, Yotam Nitzan, Yaniv Azar, Stav Shapiro, and Daniel Cohen-Or · 2021
Earlier work this paper cites.
Denoising diffusion implicit models
Jiaming Song, Chenlin Meng, and Stefano Ermon · 2021
Earlier work this paper cites.
Root mean square error (rmse) or mean absolute error (mae): When to use them or not
Timothy O Hodson · 2022
Earlier work this paper cites.
Partial convolution for padding, inpainting, and image synthesis
Guilin Liu, Aysegul Dundar, Kevin J Shih, Ting-Chun Wang, Fitsum A Reda, Karan Sapra, Zhiding Yu, Xiaodong Yang, Andrew Tao, and Bryan Catanzaro · 2022
Earlier work this paper cites.
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer · 2022
Earlier work this paper cites.
Magicvideo: Efficient video generation with latent diffusion models
Daquan Zhou, Weimin Wang, Hanshu Yan, Weiwei Lv, Yizhe Zhu, and Jiashi Feng · 2022
Earlier work this paper cites.
Improving image generation with better captions
James Betker, Gabriel Goh, Li Jing, Tim Brooks, Jianfeng Wang, Linjie Li, Long Ouyang, Juntang Zhuang, Joyce Lee, Yufei Guo, et al · 2023
Earlier work this paper cites.
Instructpix2pix: Learning to follow image editing instructions
Tim Brooks, Aleksander Holynski, and Alexei A Efros · 2023
Earlier work this paper cites.
Explore in-context learning for 3d point cloud understanding
Zhongbin Fang, Xiangtai Li, Xia Li, Joachim M Buhmann, Chen Change Loy, and Mengyuan Liu · 2023
Earlier work this paper cites.
Factormatte: Redefining video matting for re-composition tasks
Zeqi Gu, Wenqi Xian, Noah Snavely, and Abe Davis · 2023
Earlier work this paper cites.
Scalable diffusion models with transformers
William Peebles and Saining Xie · 2023
Cited alongside, same era.
Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation
Jay Zhangjie Wu, Yixiao Ge, Xintao Wang, Stan Weixian Lei, Yuchao Gu, Yufei Shi, Wynne Hsu, Ying Shan, Xiaohu Qie, and Mike Zheng Shou · 2023
Cited alongside, same era.
Smartbrush: Text and shape guided object inpainting with diffusion model
Shaoan Xie, Zhifei Zhang, Zhe Lin, Tobias Hinz, and Kun Zhang · 2023
Cited alongside, same era.
Ip-adapter: Text compatible image prompt adapter for text-to-image diffusion models
Hu Ye, Jun Zhang, Sibo Liu, Xiao Han, and Wei Yang · 2023
Cited alongside, same era.
Inpaint anything: Segment anything meets image inpainting
Tao Yu, Runseng Feng, Ruoyu Feng, Jinming Liu, Xin Jin, Wenjun Zeng, and Zhibo Chen · 2023
Cited alongside, same era.
Towards language-driven video inpainting via multimodal large language models
Jianzong Wu, Xiangtai Li, Chenyang Si, Shangchen Zhou, Jingkang Yang, Jiangning Zhang, Yining Li, Kai Chen, Yunhai Tong, Ziwei Liu, et al · 2024
Later among the works it cites.
Motionbooth: Motion-aware customized text-to-video generation
Jianzong Wu, Xiangtai Li, Yanhong Zeng, Jiangning Zhang, Qianyu Zhou, Yining Li, Yunhai Tong, and Kai Chen · 2024
Later among the works it cites.
Ufogen: You forward once large scale text-to-image generation via diffusion gans
Yanwu Xu, Yang Zhao, Zhisheng Xiao, and Tingbo Hou · 2024
Later among the works it cites.
Generative image layer decomposition with visual effects
Jinrui Yang, Qing Liu, Yijun Li, Soo Ye Kim, Daniil Pakhomov, Mengwei Ren, Jianming Zhang, Zhe Lin, Cihang Xie, and Yuyin Zhou · 2024
Later among the works it cites.
Layerpano3d: Layered 3d panorama for hyper-immersive scene generation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Adding conditional control to text-to-image diffusion models
Lvmin Zhang, Anyi Rao, and Maneesh Agrawala · 2023
Cited alongside, same era.
What makes good examples for visual in-context learning?
Yuanhan Zhang, Kaiyang Zhou, and Ziwei Liu · 2023
Cited alongside, same era.
Sequential modeling enables scalable learning for large vision models
Yutong Bai, Xinyang Geng, Karttikeya Mangalam, Amir Bar, Alan L Yuille, Trevor Darrell, Jitendra Malik, and Alexei A Efros · 2024
Cited alongside, same era.
Lumiere: A space-time diffusion model for video generation
Omer Bar-Tal, Hila Chefer, Omer Tov, Charles Herrmann, Roni Paiss, Shiran Zada, Ariel Ephrat, Junhwa Hur, Guanghui Liu, Amit Raj, et al · 2024
Cited alongside, same era.
Inverse painting: Reconstructing the painting process
Bowei Chen, Yifan Wang, Brian Curless, Ira Kemelmacher-Shlizerman, and Steven M Seitz · 2024
Cited alongside, same era.
Zero-shot image editing with reference imitation
Xi Chen, Yutong Feng, Mengting Chen, Yiyang Wang, Shilong Zhang, Yu Liu, Yujun Shen, and Hengshuang Zhao · 2024
Cited alongside, same era.
Enhance image-to-image generation with llava-generated prompts
Zhicheng Ding, Panfeng Li, Qikai Yang, and Siyang Li · 2024
Cited alongside, same era.
Shuai Yang, Jing Tan, Mengchen Zhang, Tong Wu, Yixuan Li, Gordon Wetzstein, Ziwei Liu, and Dahua Lin · 2024
Later among the works it cites.
Anyedit: Mastering unified high-quality image editing for any idea
Qifan Yu, Wei Chow, Zhongqi Yue, Kaihang Pan, Yang Wu, Xiaoyang Wan, Juncheng Li, Siliang Tang, Hanwang Zhang, and Yueting Zhuang · 2024
Later among the works it cites.
Promptfix: You prompt and we fix the photo
Yongsheng Yu, Ziyun Zeng, Hang Hua, Jianlong Fu, and Jiebo Luo · 2024
Later among the works it cites.
Transparent image layer diffusion using latent transparency
Lvmin Zhang and Maneesh Agrawala · 2024
Later among the works it cites.
A task is worth one word: Learning with task prompts for high-quality versatile image inpainting
Junhao Zhuang, Yanhong Zeng, Wenran Liu, Chun Yuan, and Kai Chen · 2024
Later among the works it cites.
Layer-animate for transparent video generation
Jingqi Bai, Jingkai Zhou, Benzhi Wang, Weihua Chen, Yang Yang, Zhen Lei, and Fan Wang · 2025
Closest in time.
Edit transfer: Learning image editing via vision in-context relations
Lan Chen, Qi Mao, Yuchao Gu, and Mike Zheng Shou · 2025
Closest in time.
Transanimate: Taming layer diffusion to generate rgba video
Xuewei Chen, Zhimin Chen, and Yiren Song · 2025
Closest in time.
Smarteraser: Remove anything from images using masked-region guidance
Longtao Jiang, Zhendong Wang, Jianmin Bao, Wengang Zhou, Dongdong Chen, Lei Shi, Dong Chen, and Houqiang Li · 2025
Closest in time.
Exploiting diffusion prior for real-world image dehazing with unpaired training
Yunwei Lan, Zhigao Cui, Chang Liu, Jialun Peng, Nian Wang, Xin Luo, and Dong Liu · 2025
Closest in time.
Garmentdiffusion: 3d garment sewing pattern generation with multimodal diffusion transformers
Xinyu Li, Qi Yao, and Yuanda Wang · 2025
Closest in time.
Nighthaze: Nighttime image dehazing via self-prior learning
Beibei Lin, Yeying Jin, Yan Wending, Wei Ye, Yuan Yuan, and Robby T Tan · 2025
Closest in time.
Efficient portrait matte creation with layer diffusion and connectivity priors
Zhiyuan Lu, Hao Lu, and Hua Huang · 2025
Closest in time.
Model see model do: Speech-driven facial animation with style control
Yifang Pan, Karan Singh, and Luiz Gustavo Hafemann · 2025
Closest in time.
Insert anything: Image insertion via in-context editing in dit
Wensong Song, Hong Jiang, Zongxing Yang, Ruijie Quan, and Yi Yang · 2025
Closest in time.
Makeanything: Harnessing diffusion transformers for multi-domain procedural sequence generation
Yiren Song, Cheng Liu, and Mike Zheng Shou · 2025
Closest in time.
Anywhere: A multi-agent framework for user-guided, reliable, and diverse foreground-conditioned image generation
Xie Tianyidan, Rui Ma, Qian Wang, Xiaoqian Ye, Feixuan Liu, Ying Tai, Zhenyu Zhang, Lanjun Wang, and Zili Yi · 2025
Closest in time.
Explore in-context segmentation via latent diffusion models
Chaoyang Wang, Xiangtai Li, Henghui Ding, Lu Qi, Jiangning Zhang, Yunhai Tong, Chen Change Loy, and Shuicheng Yan · 2025
Closest in time.
Pixelhacker: Image inpainting with structural and semantic consistency
Ziyang Xu, Kangsheng Duan, Xiaolei Shen, Zhifeng Ding, Wenyu Liu, Xiaohu Ruan, Xiaoxin Chen, and Xinggang Wang · 2025
Closest in time.
Unified dense prediction of video diffusion
Lehan Yang, Lu Qi, Xiangtai Li, Sheng Li, Varun Jampani, and Ming-Hsuan Yang · 2025
Closest in time.
Zechuan Zhang, Ji Xie, Yu Lu, Zongxin Yang, and Yi Yang · 2025
Closest in time.