Fetching the paper…
Reading the bibliography…
3D meshes are widely used in computer vision and graphics for their efficiency in animation and minimal memory use, playing a crucial role in movies, games, AR, and VR.
Towards accurate generative models of video: A new metric & challenges
T. Unterthiner, S. Van Steenkiste, K. Kurach, R. Marinier, M. Michalski, and S. Gelly · 2018
Earlier work this paper cites.
Denoising diffusion probabilistic models
J. Ho, A. Jain, and P. Abbeel · 2020
Earlier work this paper cites.
Denoising diffusion implicit models
J. Song, C. Meng, and S. Ermon · 2020
Earlier work this paper cites.
Transformers: State-of-the-art natural language processing
T. Wolf, L. Debut, V. Sanh, J. Chaumond, C. Delangue, A. Moi, P. Cistac, T. Rault, R. Louf, M. Funtowicz, J. Davison, S. Shleifer, P. von Platen, C. Ma, Y. Jernite, J. Plu, C. Xu, T. L. Scao, S. Gugger, M. Drame, Q. Lhoest, and A. M. Rush · 2020
Earlier work this paper cites.
Frozen in time: A joint video and image encoder for end-to-end retrieval
M. Bain, A. Nagrani, G. Varol, and A. Zisserman · 2021
Earlier work this paper cites.
Text2mesh: Text-driven neural stylization for meshes
O. Michel, R. Bar-On, R. Liu, S. Benaim, and R. Hanocka · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Earlier work this paper cites.
High-resolution image synthesis with latent diffusion models
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer · 2021
Earlier work this paper cites.
Laion-400m: Open dataset of clip-filtered 400 million image-text pairs
C. Schuhmann, R. Vencu, R. Beaumont, R. Kaczmarczyk, C. Mullis, A. Katta, T. Coombes, J. Jitsev, and A. Komatsuzaki · 2021
Earlier work this paper cites.
Text2live: Text-driven layered image and video editing
O. Bar-Tal, D. Ofri-Amar, R. Fridman, Y. Kasten, and T. Dekel · 2022
Earlier work this paper cites.
Tango: Text-driven photorealistic and robust 3d stylization via lighting decomposition
Y. Chen, R. Chen, J. Lei, Y. Zhang, and K. Jia · 2022
Earlier work this paper cites.
Latent video diffusion models for high-fidelity long video generation
Y. He, T. Yang, Y. Zhang, Y. Shan, and Q. Chen · 2022
Earlier work this paper cites.
Imagen video: High definition video generation with diffusion models
J. Ho, W. Chan, C. Saharia, J. Whang, R. Gao, A. Gritsenko, D. P. Kingma, B. Poole, M. Norouzi, D. J. Fleet, et al · 2022
Earlier work this paper cites.
Cogvideo: Large-scale pretraining for text-to-video generation via transformers
W. Hong, M. Ding, W. Zheng, X. Liu, and J. Tang · 2022
Earlier work this paper cites.
Clip-mesh: Generating textured meshes from text using pretrained image-text models
N. M. Khalid, T. Xie, E. Belilovsky, and P. Tiberiu · 2022
Earlier work this paper cites.
Latent-nerf for shape-guided generation of 3d shapes and textures
G. Metzer, E. Richardson, O. Patashnik, R. Giryes, and D. Cohen-Or · 2022
Earlier work this paper cites.
Dreamfusion: Text-to-3d using 2d diffusion
B. Poole, A. Jain, J. T. Barron, and B. Mildenhall · 2022
Earlier work this paper cites.
Progressive distillation for fast sampling of diffusion models
T. Salimans and J. Ho · 2022
Earlier work this paper cites.
Make-a-video: Text-to-video generation without text-video data
U. Singer, A. Polyak, T. Hayes, X. Yin, J. An, S. Zhang, Q. Hu, H. Yang, O. Ashual, O. Gafni, et al · 2022
Earlier work this paper cites.
Diffusers: State-of-the-art diffusion models
P. von Platen, S. Patil, A. Lozhkov, P. Cuenca, N. Lambert, K. Rasul, M. Davaadorj, and T. Wolf · 2022
Earlier work this paper cites.
Magicvideo: Efficient video generation with latent diffusion models
D. Zhou, W. Wang, H. Yan, W. Lv, Y. Zhu, and J. Feng · 2022
Earlier work this paper cites.
Align your latents: High-resolution video synthesis with latent diffusion models
A. Blattmann, R. Rombach, H. Ling, T. Dockhorn, S. W. Kim, S. Fidler, and K. Kreis · 2023
Earlier work this paper cites.
Texfusion: Synthesizing 3d textures with text-guided image diffusion models
T. Cao, K. Kreis, S. Fidler, N. Sharp, and K. Yin · 2023
Earlier work this paper cites.
Pix2video: Video editing using image diffusion
D. Ceylan, C.-H. Huang, and N. J. Mitra · 2023
Earlier work this paper cites.
Scenetex: High-quality texture synthesis for indoor scenes via diffusion priors
D. Z. Chen, H. Li, H.-Y. Lee, S. Tulyakov, and M. Nießner · 2023
Earlier work this paper cites.
Text2tex: Text-driven texture synthesis via diffusion models
D. Z. Chen, Y. Siddiqui, H.-Y. Lee, S. Tulyakov, and M. Nießner · 2023
Earlier work this paper cites.
Videocrafter1: Open diffusion models for high-quality video generation
H. Chen, M. Xia, Y. He, Y. Zhang, X. Cun, S. Yang, J. Xing, Y. Liu, Q. Chen, X. Wang, C. Weng, and Y. Shan · 2023
Earlier work this paper cites.
Structure and content-guided video synthesis with diffusion models
P. Esser, J. Chiu, P. Atighehchian, J. Granskog, and A. Germanidis · 2023
Cited alongside, same era.
Tokenflow: Consistent diffusion features for consistent video editing
M. Geyer, O. Bar-Tal, S. Bagon, and T. Dekel · 2023
Cited alongside, same era.
Sparsectrl: Adding sparse controls to text-to-video diffusion models
Y. Guo, C. Yang, A. Rao, M. Agrawala, D. Lin, and B. Dai · 2023
Cited alongside, same era.
Animatediff: Animate your personalized text-to-image diffusion models without specific tuning
Y. Guo, C. Yang, A. Rao, Z. Liang, Y. Wang, Y. Qiao, M. Agrawala, D. Lin, and B. Dai · 2023
Cited alongside, same era.
Text2video-zero: Text-to-image diffusion models are zero-shot video generators
I2vgen-xl: High-quality image-to-video synthesis via cascaded diffusion models
S. Zhang, J. Wang, Y. Zhang, K. Zhao, H. Yuan, Z. Qing, X. Wang, D. Zhao, and J. Zhou · 2023
Later among the works it cites.
Controlvideo: Training-free controllable text-to-video generation
Y. Zhang, Y. Wei, D. Jiang, X. Zhang, W. Zuo, and Q. Tian · 2023
Later among the works it cites.
Meta 3d texturegen: Fast and consistent texture generation for 3d objects
R. Bensadoun, Y. Kleiman, I. Azuri, O. Harosh, A. Vedaldi, N. Neverova, and O. Gafni · 2024
Closest in time.
Generative rendering: Controllable 4d-guided video generation with 2d diffusion models
S. Cai, D. Ceylan, M. Gadelha, C.-H. Huang, T. Wang, and G. Wetzstein · 2024
Closest in time.
Videocrafter2: Overcoming data limitations for high-quality video diffusion models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
L. Khachatryan, A. Movsisyan, V. Tadevosyan, R. Henschel, Z. Wang, S. Navasardyan, and H. Shi · 2023
Cited alongside, same era.
Vivid-1-to-3: Novel view synthesis with video diffusion models
J.-g. Kwak, E. Dong, Y. Jin, H. Ko, S. Mahajan, and K. M. Yi · 2023
Cited alongside, same era.
Videogen: A reference-guided latent diffusion approach for high definition text-to-video generation
X. Li, W. Chu, Y. Wu, W. Yuan, F. Liu, Q. Zhang, F. Li, H. Feng, E. Ding, and J. Wang · 2023
Cited alongside, same era.
Magic3d: High-resolution text-to-3d content creation
C.-H. Lin, J. Gao, L. Tang, T. Takikawa, X. Zeng, X. Huang, K. Kreis, S. Fidler, M.-Y. Liu, and T.-Y. Lin · 2023
Cited alongside, same era.
Syncdreamer: Generating multiview-consistent images from a single-view image
Y. Liu, C. Lin, Z. Zeng, X. Long, L. Liu, T. Komura, and W. Wang · 2023
Cited alongside, same era.
Text-guided texturing by synchronized multi-view diffusion
Y. Liu, M. Xie, H. Liu, and T.-T. Wong · 2023
Cited alongside, same era.
Wonder3d: Single image to 3d using cross-domain diffusion
X. Long, Y.-C. Guo, C. Lin, Y. Liu, Z. Dou, L. Liu, Y. Ma, S.-H. Zhang, M. Habermann, C. Theobalt, et al · 2023
Cited alongside, same era.
Fatezero: Fusing attentions for zero-shot text-based video editing
C. Qi, X. Cun, Y. Zhang, C. Lei, X. Wang, Y. Shan, and Q. Chen · 2023
Cited alongside, same era.
H. Chen, Y. Zhang, X. Cun, M. Xia, X. Wang, C. Weng, and Y. Shan · 2024
Closest in time.
Blender - a 3D modelling and rendering package
B. O. Community · 2024
Closest in time.
Latentman: Generating consistent animated characters using image diffusion models
A. Eldesokey and P. Wonka · 2024
Closest in time.
Texgen: Text-guided 3d texture generation with multi-view sampling and resampling
D. Huo, Z. Guo, X. Zuo, Z. Shi, J. Lu, P. Dai, S. Xu, L. Cheng, and Y.-H. Yang · 2024
Closest in time.
Vivid-zoo: Multi-view video generation with diffusion model
B. Li, C. Zheng, W. Zhu, J. Mai, B. Zhang, P. Wonka, and B. Ghanem · 2024
Closest in time.
H. Lin, J. Cho, A. Zala, and M. Bansal · 2024
Closest in time.
Freelong: Training-free long video generation with spectralblend temporal attention
Y. Lu, Y. Liang, L. Zhu, and Y. Yang · 2024
Closest in time.
Controlnext: Powerful and efficient control for image and video generation
B. Peng, J. Wang, Y. Zhang, W. Li, M.-C. Yang, and J. Jia · 2024
Closest in time.
Compositional 3d scene generation using locally conditioned diffusion
R. Po and G. Wetzstein · 2024
Closest in time.
L4gm: Large 4d gaussian reconstruction model
J. Ren, C. Xie, A. Mirzaei, K. Kreis, Z. Liu, A. Torralba, S. Fidler, S. W. Kim, H. Ling, et al · 2024
Closest in time.
S. Tang, J. Chen, D. Wang, C. Tang, F. Zhang, Y. Fan, V. Chandra, Y. Furukawa, and R. Ranjan · 2024
Closest in time.
SV3D: Novel multi-view synthesis and 3D generation from a single image using latent video diffusion
V. Voleti, C.-H. Yao, M. Boss, A. Letts, D. Pankratz, D. Tochilkin, C. Laforte, R. Rombach, and V. Jampani · 2024
Closest in time.
Cat4d: Create anything in 4d with multi-view video diffusion models
R. Wu, R. Gao, B. Poole, A. Trevithick, C. Zheng, J. T. Barron, and A. Holynski · 2024
Closest in time.
SV4D: Dynamic 3d content generation with multi-frame and multi-view consistency
Y. Xie, C.-H. Yao, V. Voleti, H. Jiang, and V. Jampani · 2024
Closest in time.
Magicanimate: Temporally consistent human image animation using diffusion model
Z. Xu, J. Zhang, J. H. Liew, H. Yan, J.-W. Liu, C. Zhang, J. Feng, and M. Z. Shou · 2024
Closest in time.
Cogvideox: Text-to-video diffusion models with an expert transformer
Z. Yang, J. Teng, W. Zheng, M. Ding, S. Huang, J. Xu, Y. Yang, W. Hong, X. Zhang, G. Feng, et al · 2024
Closest in time.
Paint3d: Paint anything 3d with lighting-less texture diffusion models
X. Zeng, X. Chen, Z. Qi, W. Liu, Z. Zhao, Z. Wang, B. Fu, Y. Liu, and G. Yu · 2024
Closest in time.
Clay: A controllable large-scale generative model for creating high-quality 3d assets
L. Zhang, Z. Wang, Q. Zhang, Q. Qiu, A. Pang, H. Jiang, W. Yang, L. Xu, and J. Yu · 2024
Closest in time.
Genxd: Generating any 3d and 4d scenes
Y. Zhao, C.-C. Lin, K. Lin, Z. Yan, L. Li, Z. Yang, J. Wang, G. H. Lee, and L. Wang · 2024
Closest in time.
Cogview3: Finer and faster text-to-image generation via relay diffusion
W. Zheng, J. Teng, Z. Yang, W. Wang, J. Chen, X. Gu, Y. Dong, M. Ding, and J. Tang · 2024
Closest in time.
Diffusionrenderer: Neural inverse and forward rendering with video diffusion models
R. Liang, Z. Gojcic, H. Ling, J. Munkberg, J. Hasselgren, Z.-H. Lin, J. Gao, A. Keller, N. Vijaykumar, S. Fidler, et al · 2025
Closest in time.
Freqprior: Improving video diffusion models with frequency filtering gaussian noise
Y. Yuan, Y. Guo, C. Wang, W. Zhang, H. Xu, and L. Zhang · 2025
Closest in time.