Fetching the paper…
Reading the bibliography…
Diffusion models gain increasing popularity for their generative capabilities.
Estimation of non-normalized statistical models by score matching
Aapo Hyvärinen and Peter Dayan. 2005 · 2005
Earlier work this paper cites.
Visualizing data using t-SNE
Laurens Van der Maaten and Geoffrey Hinton. 2008 · 2008
Earlier work this paper cites.
Oxford dictionary of English
Angus Stevenson. 2010 · 2010
Earlier work this paper cites.
A connection between score matching and denoising autoencoders
Pascal Vincent. 2011 · 2011
Earlier work this paper cites.
Generative Adversarial Nets. In NeurIPS
Ian J Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron C Courville, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
U-Net: Convolutional networks for biomedical image segmentation. In MICCAI
Olaf Ronneberger, Philipp Fischer, and Thomas Brox. 2015 · 2015
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics. In ICML
Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli. 2015 · 2015
Earlier work this paper cites.
Visual Relationship Detection with Language Priors. In ECCV
Cewu Lu, Ranjay Krishna, Michael Bernstein, and Fei-Fei Li. 2016 · 2016
Earlier work this paper cites.
Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, Michael Bernstein, and Fei-Fei Li. 2017 · 2017
Earlier work this paper cites.
Video Visual Relation Detection. In ACM MM
Xindi Shang, Tongwei Ren, Jingfan Guo, Hanwang Zhang, and Tat-Seng Chua. 2017 · 2017
Earlier work this paper cites.
Scene graph generation by iterative message passing. In CVPR
Danfei Xu, Yuke Zhu, Christopher B Choy, and Fei-Fei Li. 2017 · 2017
Earlier work this paper cites.
Visual relationship detection with internal and external linguistic knowledge distillation. In ICCV
Ruichi Yu, Ang Li, Vlad I Morariu, and Larry S Davis. 2017 · 2017
Earlier work this paper cites.
Towards context-aware interaction recognition for visual relationship detection. In ICCV
Bohan Zhuang, Lingqiao Liu, Chunhua Shen, and Ian Reid. 2017 · 2017
Earlier work this paper cites.
Representation learning with contrastive predictive coding
Aaron van den Oord, Yazhe Li, and Oriol Vinyals. 2018 · 2018
Earlier work this paper cites.
Decoupled weight decay regularization. In ICLR
Ilya Loshchilov and Frank Hutter. 2019 · 2019
Earlier work this paper cites.
A note on data biases in generative models. In NeurIPS Workshop
Patrick Esser, Robin Rombach, and Björn Ommer. 2020 · 2020
Earlier work this paper cites.
Momentum contrast for unsupervised visual representation learning. In CVPR
Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick. 2020 · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models. In NeurIPS
Jonathan Ho, Ajay Jain, and Pieter Abbeel. 2020 · 2020
Earlier work this paper cites.
Action genome: Actions as compositions of spatio-temporal scene graphs. In CVPR . 10236–10247
Jingwei Ji, Ranjay Krishna, Fei-Fei Li, and Juan Carlos Niebles. 2020 · 2020
Earlier work this paper cites.
End-to-end learning of visual representations from uncurated instructional videos. In CVPR . 9879–9889
Antoine Miech, Jean-Baptiste Alayrac, Lucas Smaira, Ivan Laptev, Josef Sivic, and Andrew Zisserman. 2020 · 2020
Earlier work this paper cites.
SegDiff: Image Segmentation with Diffusion Probabilistic Models
Tomer Amit, Eliya Nachmani, Tal Shaharbany, and Lior Wolf. 2021 · 2021
Earlier work this paper cites.
Structured denoising diffusion models in discrete state-spaces. In NeurIPS
Jacob Austin, Daniel D Johnson, Jonathan Ho, Daniel Tarlow, and Rianne van den Berg. 2021 · 2021
Earlier work this paper cites.
Diffusion Models Beat GANs on Image Synthesis. In NeurIPS
Prafulla Dhariwal and Alexander Nichol. 2021 · 2021
Earlier work this paper cites.
ImageBART: Bidirectional context with multinomial diffusion for autoregressive image synthesis. In NeurIPS
Patrick Esser, Robin Rombach, Andreas Blattmann, and Bjorn Ommer. 2021 · 2021
Earlier work this paper cites.
GLIDE: Towards photorealistic image generation and editing with text-guided diffusion models
Alex Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam, Pamela Mishkin, Bob McGrew, Ilya Sutskever, and Mark Chen. 2021 · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision. In ICML
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Cited alongside, same era.
This face does not exist… but it might be yours! identity leakage in generative models. In WACV
Patrick Tinsley, Adam Czajka, and Patrick Flynn. 2021 · 2021
Cited alongside, same era.
Label-efficient semantic segmentation with diffusion models. In ICLR
Dmitry Baranchuk, Ivan Rubachev, Andrey Voynov, Valentin Khrulkov, and Artem Babenko. 2022 · 2022
Cited alongside, same era.
An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion
Rinon Gal, Yuval Alaluf, Yuval Atzmon, Or Patashnik, Amit H. Bermano, Gal Chechik, and Daniel Cohen-Or. 2022 · 2022
Cited alongside, same era.
Diffusion Models as Plug-and-Play Priors. In NeurIPS
Alexandros Graikos, Nikolay Malkin, Nebojsa Jojic, and Dimitris Samaras. 2022 · 2022
Panoptic Scene Graph Generation. In ECCV . Springer, 178–196
Jingkang Yang, Yi Zhe Ang, Zujin Guo, Kaiyang Zhou, Wayne Zhang, and Ziwei Liu. 2022 · 2022
Later among the works it cites.
A neural space-time representation for text-to-image personalization
Yuval Alaluf, Elad Richardson, Gal Metzer, and Daniel Cohen-Or. 2023 · 2023
Closest in time.
Domain-agnostic tuning-encoder for fast personalization of text-to-image models
Moab Arar, Rinon Gal, Yuval Atzmon, Gal Chechik, Daniel Cohen-Or, Ariel Shamir, and Amit H Bermano. 2023 · 2023
Closest in time.
Align your Latents: High-Resolution Video Synthesis with Latent Diffusion Models. In CVPR
Andreas Blattmann, Robin Rombach, Huan Ling, Tim Dockhorn, Seung Wook Kim, Sanja Fidler, and Karsten Kreis. 2023 · 2023
Closest in time.
DisenBooth: Disentangled Parameter-Efficient Tuning for Subject-Driven Text-to-Image Generation
Hong Chen, Yipeng Zhang, Xin Wang, Xuguang Duan, Yuwei Zhou, and Wenwu Zhu. 2023b · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Vector quantized diffusion model for text-to-image synthesis. In CVPR
Shuyang Gu, Dong Chen, Jianmin Bao, Fang Wen, Bo Zhang, Dongdong Chen, Lu Yuan, and Baining Guo. 2022 · 2022
Cited alongside, same era.
Flexible Diffusion Modeling of Long Videos
William Harvey, Saeid Naderiparizi, Vaden Masrani, Christian Weilbach, and Frank Wood. 2022 · 2022
Cited alongside, same era.
Latent video diffusion models for high-fidelity video generation with arbitrary lengths
Yingqing He, Tianyu Yang, Yong Zhang, Ying Shan, and Qifeng Chen. 2022 · 2022
Cited alongside, same era.
Cascaded Diffusion Models for High Fidelity Image Generation
Jonathan Ho, Chitwan Saharia, William Chan, David J Fleet, Mohammad Norouzi, and Tim Salimans. 2022a · 2022
Cited alongside, same era.
Jonathan Ho, Tim Salimans, Alexey Gritsenko, William Chan, Mohammad Norouzi, and David J Fleet. 2022b · 2022
Cited alongside, same era.
LoRA: Low-Rank Adaptation of Large Language Models. In ICLR
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022 · 2022
Cited alongside, same era.
Text2human: Text-driven controllable human image generation
Yuming Jiang, Shuai Yang, Haonan Qju, Wayne Wu, Chen Change Loy, and Ziwei Liu. 2022 · 2022
Cited alongside, same era.
Closest in time.
Subject-driven Text-to-Image Generation via Apprenticeship Learning
Wenhu Chen, Hexiang Hu, Yandong Li, Nataniel Ruiz, Xuhui Jia, Ming-Wei Chang, and William W Cohen. 2023a · 2023
Closest in time.
Custom-Edit: Text-Guided Image Editing with Customized Diffusion Models
Jooyoung Choi, Yunjey Choi, Yunji Kim, Junho Kim, and Sungroh Yoon. 2023 · 2023
Closest in time.
Encoder-based domain tuning for fast personalization of text-to-image models
Rinon Gal, Moab Arar, Yuval Atzmon, Amit H Bermano, Gal Chechik, and Daniel Cohen-Or. 2023 · 2023
Closest in time.
TaleCrafter: Interactive Story Visualization with Multiple Characters
Yuan Gong, Youxin Pang, Xiaodong Cun, Menghan Xia, Haoxin Chen, Longyue Wang, Yong Zhang, Xintao Wang, Ying Shan, and Yujiu Yang. 2023 · 2023
Closest in time.
SVDiff: Compact Parameter Space for Diffusion Fine-Tuning
Ligong Han, Yinxiao Li, Han Zhang, Peyman Milanfar, Dimitris Metaxas, and Feng Yang. 2023 · 2023
Closest in time.
Collaborative Diffusion for Multi-Modal Face Generation and Editing. In CVPR
Ziqi Huang, Kelvin C.K. Chan, Yuming Jiang, and Ziwei Liu. 2023 · 2023
Closest in time.
Taming Encoder for Zero Fine-tuning Image Customization with Text-to-Image Diffusion
Xuhui Jia, Yang Zhao, Kelvin C.K. Chan, Yandong Li, Han Zhang, Boqing Gong, Tingbo Hou, Huisheng Wang, and Yu-Chuan Su. 2023 · 2023
Closest in time.
Dongxu Li, Junnan Li, and Steven CH Hoi. 2023a · 2023
Closest in time.
Generate Anything Anywhere in Any Scene
Yuheng Li, Haotian Liu, Yangming Wen, and Yong Jae Lee. 2023b · 2023
Closest in time.
Subject-diffusion: Open domain personalized text-to-image generation without test-time fine-tuning
Jian Ma, Junhao Liang, Chen Chen, and Haonan Lu. 2023 · 2023
Closest in time.
Localizing object-level shape variations with text-to-image diffusion models. In ICCV
Or Patashnik, Daniel Garibi, Idan Azuri, Hadar Averbuch-Elor, and Daniel Cohen-Or. 2023 · 2023
Closest in time.
p+: Extended textual conditioning in text-to-image generation
Andrey Voynov, Qinghao Chu, Daniel Cohen-Or, and Kfir Aberman. 2023 · 2023
Closest in time.
Diffusion Models Generate Images Like Painters: an Analytical Theory of Outline First, Details Later
Binxu Wang and John J. Vastola. 2023 · 2023
Closest in time.
ELITE: Encoding Visual Concepts into Textual Embeddings for Customized Text-to-Image Generation
Yuxiang Wei, Yabo Zhang, Zhilong Ji, Jinfeng Bai, Lei Zhang, and Wangmeng Zuo. 2023 · 2023
Closest in time.
Prompt-Free Diffusion: Taking "Text" out of Text-to-Image Diffusion Models
Xingqian Xu, Jiayi Guo, Zhangyang Wang, Gao Huang, Irfan Essa, and Humphrey Shi. 2023 · 2023
Closest in time.
Panoptic Video Scene Graph Generation. In CVPR
Jingkang Yang, Wenxuan Peng, Xiangtai Li, Zujin Guo, Liangyu Chen, Bo Li, Zheng Ma, Kaiyang Zhou, Wayne Zhang, Chen Change Loy, and Ziwei Liu. 2023 · 2023
Closest in time.
Ip-adapter: Text compatible image prompt adapter for text-to-image diffusion models
Hu Ye, Jun Zhang, Sibo Liu, Xiao Han, and Wei Yang. 2023 · 2023
Closest in time.
Yufan Zhou, Ruiyi Zhang, Tong Sun, and Jinhui Xu. 2023 · 2023
Closest in time.
Hyperdreambooth: Hypernetworks for fast personalization of text-to-image models. In CVPR
Nataniel Ruiz, Yuanzhen Li, Varun Jampani, Wei Wei, Tingbo Hou, Yael Pritch, Neal Wadhwa, Michael Rubinstein, and Kfir Aberman. 2024 · 2024
Closest in time.