Fetching the paper…
Reading the bibliography…
ControlNets are widely used for adding spatial control to text-to-image diffusion models with different conditions, such as depth maps, scribbles/sketches, and human poses.
The OpenCV Library
G. Bradski · 2000
Earlier work this paper cites.
Two-frame motion estimation based on polynomial expansion
G. Farnebäck · 2003
Earlier work this paper cites.
Two-frame motion estimation based on polynomial expansion
G. Farnebäck · 2003
Earlier work this paper cites.
Image quality assessment: from error visibility to structural similarity
Z. Wang, A. C. Bovik, H. R. Sheikh, and E. P. Simoncelli · 2004
Earlier work this paper cites.
Estimation lemma — Wikipedia, the free encyclopedia, 2010
Estimation lemma · 2010
Earlier work this paper cites.
Auto-encoding variational bayes
D. P. Kingma and M. Welling · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
J. Sohl-Dickstein, E. A. Weiss, N. Maheswaranathan, and S. Ganguli · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Realtime multi-person 2d pose estimation using part affinity fields
Z. Cao, T. Simon, S.-E. Wei, and Y. Sheikh · 2017
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter · 2017
Earlier work this paper cites.
Video generation from text
Y. Li, M. R. Min, D. Shen, D. Carlson, and L. Carin · 2017
Earlier work this paper cites.
The 2017 davis challenge on video object segmentation
J. Pont-Tuset, F. Perazzi, S. Caelles, P. Arbeláez, A. Sorkine-Hornung, and L. Van Gool · 2017
Earlier work this paper cites.
Optical flow estimation using a spatial pyramid network
A. Ranjan and M. J. Black · 2017
Earlier work this paper cites.
Outrageously Large Neural Networks: the Sparsely-Gated Mixture-of-Experts Layer
N. Shazeer, A. Mirhoseini, K. Maziarz, A. Davis, Q. Le, G. Hinton, and J. Dean · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
I. Loshchilov and F. Hutter · 2018
Earlier work this paper cites.
Unified perceptual parsing for scene understanding
T. Xiao, Y. Liu, B. Zhou, Y. Jiang, and J. Sun · 2018
Earlier work this paper cites.
Learning to forecast and refine residual motion for image-to-video generation
L. Zhao, X. Peng, Y. Tian, M. Kapadia, and D. Metaxas · 2018
Earlier work this paper cites.
Parameter-efficient transfer learning for nlp
N. Houlsby, A. Giurgiu, S. Jastrzebski, B. Morrone, Q. de Laroussilhe, A. Gesmundo, M. Attariyan, and S. Gelly · 2019
Earlier work this paper cites.
Storygan: A sequential conditional gan for story visualization
Y. Li, Z. Gan, Y. Shen, J. Liu, Y. Cheng, Y. Wu, L. Carin, D. Carlson, and J. Gao · 2019
Earlier work this paper cites.
Generative adversarial networks
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
J. Ho, A. Jain, and P. Abbeel · 2020
Earlier work this paper cites.
Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer
R. Ranftl, K. Lasinger, D. Hafner, K. Schindler, and V. Koltun · 2020
Earlier work this paper cites.
Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters
J. Rasley, S. Rajbhandari, O. Ruwase, and Y. He · 2020
Earlier work this paper cites.
Transformers: State-of-the-art natural language processing
T. Wolf, L. Debut, V. Sanh, J. Chaumond, C. Delangue, A. Moi, P. Cistac, T. Rault, R. Louf, M. Funtowicz, J. Davison, S. Shleifer, P. von Platen, C. Ma, Y. Jernite, J. Plu, C. Xu, T. L. Scao, S. Gugger, M. Drame, Q. Lhoest, and A. M. Rush · 2020
Earlier work this paper cites.
High-resolution image synthesis with latent diffusion models
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer · 2021
Earlier work this paper cites.
Segformer: Simple and efficient design for semantic segmentation with transformers
E. Xie, W. Wang, Z. Yu, A. Anandkumar, J. M. Alvarez, and P. Luo · 2021
Earlier work this paper cites.
Make-A-Scene: Scene-Based Text-to-Image Generation with Human Priors
O. Gafni, A. Polyak, O. Ashual, S. Sheynin, D. Parikh, and Y. Taigman · 2022
Earlier work this paper cites.
Latent video diffusion models for high-fidelity long video generation, 2022
Y. He, T. Yang, Y. Zhang, Y. Shan, and Q. Chen · 2022
Earlier work this paper cites.
Imagen video: High definition video generation with diffusion models
J. Ho, W. Chan, C. Saharia, J. Whang, R. Gao, A. Gritsenko, D. P. Kingma, B. Poole, M. Norouzi, D. J. Fleet, et al · 2022
Earlier work this paper cites.
xformers: A modular and hackable transformer modelling library
B. Lefaudeux, F. Massa, D. Liskovich, W. Xiong, V. Caggiano, S. Naren, M. Xu, J. Hu, M. Tintore, S. Zhang, P. Labatut, D. Haziza, L. Wehrstedt, J. Reizenstein, and G. Sizov · 2022
Earlier work this paper cites.
On aliased resizing and surprising subtleties in gan evaluation
G. Parmar, R. Zhang, and J.-Y. Zhu · 2022
Cited alongside, same era.
Hierarchical Text-Conditional Image Generation with CLIP Latents, 2022
A. Ramesh, P. Dhariwal, A. Nichol, C. Chu, and M. Chen · 2022
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer · 2022
Cited alongside, same era.
Low-rank adaptation for fast text-to-image diffusion fine-tuning, 2022
S. Ryu · 2022
Cited alongside, same era.
Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding
C. Saharia, W. Chan, S. Saxena, L. Li, J. Whang, E. Denton, S. K. S. Ghasemipour, B. K. Ayan, S. S. Mahdavi, R. G. Lopes, T. Salimans, J. Ho, D. J. Fleet, and M. Norouzi · 2022
Cited alongside, same era.
Laion-5b: An open large-scale dataset for training next generation image-text models
DreamBooth: Fine Tuning Text-to-Image Diffusion Models for Subject-Driven Generation
N. Ruiz, Y. Li, V. Jampani, Y. Pritch, M. Rubinstein, and K. Aberman · 2023
Later among the works it cites.
Laion pop: 600,000 high-resolution images with detailed descriptions
C. Schuhmann and P. Bevan · 2023
Later among the works it cites.
Make-a-video: Text-to-video generation without text-video data
U. Singer, A. Polyak, T. Hayes, X. Yin, J. An, S. Zhang, Q. Hu, H. Yang, O. Ashual, O. Gafni, et al · 2023
Later among the works it cites.
Plug-and-play diffusion features for text-driven image-to-image translation
N. Tumanyan, M. Geyer, S. Bagon, and T. Dekel · 2023
Later among the works it cites.
Gen-l-video: Multi-text to long video generation via temporal co-denoising
F.-Y. Wang, W. Chen, G. Song, H.-J. Ye, Y. Liu, and H. Li · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
C. Schuhmann, R. Beaumont, R. Vencu, C. Gordon, R. Wightman, M. Cherti, T. Coombes, A. Katta, C. Mullis, M. Wortsman, et al · 2022
Cited alongside, same era.
Stylegan-v: A continuous video generator with the price, image quality and perks of stylegan2
I. Skorokhodov, S. Tulyakov, and M. Elhoseiny · 2022
Cited alongside, same era.
Diffusers: State-of-the-art diffusion models
P. von Platen, S. Patil, A. Lozhkov, P. Cuenca, N. Lambert, K. Rasul, M. Davaadorj, and T. Wolf · 2022
Cited alongside, same era.
Vl-adapter: Parameter-efficient transfer learning for vision-and-language tasks
M. B. Yi-Lin Sung, Jaemin Cho · 2022
Cited alongside, same era.
Magicvideo: Efficient video generation with latent diffusion models
D. Zhou, W. Wang, H. Yan, W. Lv, Y. Zhu, and J. Feng · 2022
Cited alongside, same era.
SpaText: Spatio-Textual Representation for Controllable Image Generation
O. Avrahami, T. Hayes, O. Gafni, S. Gupta, Y. Taigman, D. Parikh, D. Lischinski, O. Fried, and X. Yin · 2023
Cited alongside, same era.
Stable video diffusion: Scaling latent video diffusion models to large datasets
A. Blattmann, T. Dockhorn, S. Kulal, D. Mendelevitch, M. Kilian, D. Lorenz, Y. Levi, Z. English, V. Voleti, A. Letts, et al · 2023
Cited alongside, same era.
Modelscope text-to-video technical report, 2023
J. Wang, H. Yuan, D. Chen, Y. Zhang, X. Wang, and S. Zhang · 2023
Later among the works it cites.
Cogvlm: Visual expert for pretrained language models
W. Wang, Q. Lv, W. Yu, W. Hong, J. Qi, Y. Wang, J. Ji, Z. Yang, L. Zhao, X. Song, et al · 2023
Later among the works it cites.
Dynamicrafter: Animating open-domain images with video diffusion priors
J. Xing, M. Xia, Y. Zhang, H. Chen, W. Yu, H. Liu, X. Wang, T.-T. Wong, and Y. Shan · 2023
Later among the works it cites.
Reco: Region-controlled text-to-image generation
Z. Yang, J. Wang, Z. Gan, L. Li, K. Lin, C. Wu, N. Duan, Z. Liu, C. Liu, M. Zeng, and L. Wang · 2023
Later among the works it cites.
NUWA-XL: Diffusion over diffusion for eXtremely long video generation
S. Yin, C. Wu, H. Yang, J. Wang, X. Wang, M. Ni, Z. Yang, L. Li, S. Liu, F. Yang, J. Fu, M. Gong, L. Wang, Z. Liu, H. Li, and N. Duan · 2023
Later among the works it cites.
Adding conditional control to text-to-image diffusion models
L. Zhang, A. Rao, and M. Agrawala · 2023
Later among the works it cites.
Adding conditional control to text-to-image diffusion models
L. Zhang, A. Rao, and M. Agrawala · 2023
Later among the works it cites.
Adding conditional control to text-to-image diffusion models
L. Zhang, A. Rao, and M. Agrawala · 2023
Later among the works it cites.
I2vgen-xl: High-quality image-to-video synthesis via cascaded diffusion models
S. Zhang, J. Wang, Y. Zhang, K. Zhao, H. Yuan, Z. Qin, X. Wang, D. Zhao, and J. Zhou · 2023
Later among the works it cites.
Videocrafter2: Overcoming data limitations for high-quality video diffusion models, 2024
H. Chen, Y. Zhang, X. Cun, M. Xia, X. Wang, C. Weng, and Y. Shan · 2024
Closest in time.
Pixart-$\alpha$: Fast training of diffusion transformer for photorealistic text-to-image synthesis
J. Chen, J. YU, C. GE, L. Yao, E. Xie, Z. Wang, J. Kwok, P. Luo, H. Lu, and Z. Li · 2024
Closest in time.
Panda-70m: Captioning 70m videos with multiple cross-modality teachers
T.-S. Chen, A. Siarohin, W. Menapace, E. Deyneka, H.-w. Chao, B. E. Jeon, Y. Fang, H.-Y. Lee, J. Ren, M.-H. Yang, and S. Tulyakov · 2024
Closest in time.
Scaling rectified flow transformers for high-resolution image synthesis, 2024
P. Esser, S. Kulal, A. Blattmann, R. Entezari, J. Müller, H. Saini, Y. Levi, D. Lorenz, A. Sauer, F. Boesel, D. Podell, T. Dockhorn, Z. English, K. Lacey, A. Goodwin, Y. Marek, and R. Rombach · 2024
Closest in time.
Dit-visualization
Q. Guo and D. Yue · 2024
Closest in time.
Animatediff: Animate your personalized text-to-image diffusion models without specific tuning
Y. Guo, C. Yang, A. Rao, Y. Wang, Y. Qiao, D. Lin, and B. Dai · 2024
Closest in time.
Videodrafter: Content-consistent multi-scene video generation with llm
F. Long, Z. Qiu, T. Yao, and T. Mei · 2024
Closest in time.
Sit: Exploring flow and diffusion-based generative models with scalable interpolant transformers
N. Ma, M. Goldstein, M. S. Albergo, N. M. Boffi, E. Vanden-Eijnden, and S. Xie · 2024
Closest in time.
Latte: Latent diffusion transformer for video generation
X. Ma, Y. Wang, G. Jia, X. Chen, Z. Liu, Y.-F. Li, C. Chen, and Y. Qiao · 2024
Closest in time.
Snap video: Scaled spatiotemporal transformers for text-to-video synthesis, 2024
W. Menapace, A. Siarohin, I. Skorokhodov, E. Deyneka, T.-S. Chen, A. Kag, Y. Fang, A. Stoliar, E. Ricci, J. Ren, and S. Tulyakov · 2024
Closest in time.
Video generation models as world simulators, 2024
OpenAI · 2024
Closest in time.
Sdxl: Improving latent diffusion models for high-resolution image synthesis
D. Podell, Z. English, K. Lacey, A. Blattmann, T. Dockhorn, J. Müller, J. Penna, and R. Rombach · 2024
Closest in time.
X-adapter: Adding universal compatibility of plugins for upgraded diffusion model
L. Ran, X. Cun, J.-W. Liu, R. Zhao, S. Zijie, X. Wang, J. Keppo, and M. Z. Shou · 2024
Closest in time.
Videocomposer: Compositional video synthesis with motion controllability
X. Wang, H. Yuan, S. Zhang, D. Chen, J. Wang, Y. Zhang, Y. Shen, D. Zhao, and J. Zhou · 2024
Closest in time.
Controlvideo: Training-free controllable text-to-video generation
Y. Zhang, Y. Wei, D. Jiang, X. Zhang, W. Zuo, and Q. Tian · 2024
Closest in time.
Uni-controlnet: All-in-one control to text-to-image diffusion models
S. Zhao, D. Chen, Y.-C. Chen, J. Bao, S. Hao, L. Yuan, and K.-Y. K. Wong · 2024
Closest in time.