Fetching the paper…
Reading the bibliography…
Recent techniques for text-to-4D generation synthesize dynamic 3D scenes using supervision from pre-trained text-to-video models.
Kingma, D.P., Ba, J.: Adam: A method for stochastic optimization. Proc. ICLR (2015)
2015
Earlier work this paper cites.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, Ł., Polosukhin, I.: Attention is all you need. Proc. NeurIPS (2017)
2017
Earlier work this paper cites.
Chen, K., Choy, C.B., Savva, M., Chang, A.X., Funkhouser, T., Savarese, S.: Text2Shape: Generating shapes from natural language by learning joint embeddings. Proc. ACCV (2018)
2018
Earlier work this paper cites.
Mildenhall, B., Srinivasan, P.P., Tancik, M., Barron, J.T., Ramamoorthi, R., Ng, R.: NeRF: Representing scenes as neural radiance fields for view synthesis. Proc. ECCV (2020)
2020
Earlier work this paper cites.
Bain, M., Nagrani, A., Varol, G., Zisserman, A.: Frozen in time: A joint video and image encoder for end-to-end retrieval. Proc. ICCV (2021)
2021
Earlier work this paper cites.
DeVries, T., Bautista, M.A., Srivastava, N., Taylor, G.W., Susskind, J.M.: Unconstrained scene generation with locally conditioned radiance fields. Proc. ICCV (2021)
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
Park, D.H., Azadi, S., Liu, X., Darrell, T., Rohrbach, A.: Benchmark for compositional text-to-image synthesis. In: Proc. NeurIPS (2021)
2021
Earlier work this paper cites.
Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., Sutskever, I.: Learning transferable visual models from natural language supervision. Proc. ICML (2021)
2021
Earlier work this paper cites.
Chan, E.R., Lin, C.Z., Chan, M.A., Nagano, K., Pan, B., De Mello, S., Gallo, O., Guibas, L.J., Tremblay, J., Khamis, S., et al.: Efficient geometry-aware 3D generative adversarial networks. Proc. CVPR (2022)
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
Jain, A., Mildenhall, B., Barron, J.T., Abbeel, P., Poole, B.: Zero-shot text-guided object generation with dream fields. Proc. CVPR (2022)
2022
Earlier work this paper cites.
Müller, T., Evans, A., Schied, C., Keller, A.: Instant neural graphics primitives with a multiresolution hash encoding. ACM Trans. Graph. 41
2022
Earlier work this paper cites.
Or-El, R., Luo, X., Shan, M., Shechtman, E., Park, J.J., Kemelmacher-Shlizerman, I.: StyleSDF: High-resolution 3D-consistent image and geometry generation. Proc. CVPR (2022)
2022
Earlier work this paper cites.
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., Ommer, B.: High-resolution image synthesis with latent diffusion models. In: Proc. CVPR (2022)
2022
Earlier work this paper cites.
Sanghi, A., Chu, H., Lambourne, J.G., Wang, Y., Cheng, C.Y., Fumero, M., Malekshan, K.R.: CLIP-Forge: Towards zero-shot text-to-shape generation. Proc. CVPR (2022)
2022
Earlier work this paper cites.
Schwarz, K., Sauer, A., Niemeyer, M., Liao, Y., Geiger, A.: VoxGRAF: Fast 3D-aware image synthesis with sparse voxel grids. Proc. NeurIPS (2022)
2022
Earlier work this paper cites.
Skorokhodov, I., Tulyakov, S., Elhoseiny, M.: StyleGAN-V: A continuous video generator with the price, image quality and perks of StyleGAN2. Proc. CVPR (2022)
2022
Earlier work this paper cites.
Wang, C., Chai, M., He, M., Chen, D., Liao, J.: Clip-NeRF: Text-and-image driven manipulation of neural radiance fields. Proc. CVPR (2022)
2022
Earlier work this paper cites.
Xue, H., Hang, T., Zeng, Y., Sun, Y., Liu, B., Yang, H., Fu, J., Guo, B.: Advancing high-resolution video-language representation with large-scale video transcriptions. In: Proc. CVPR (2022)
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
Zeroscope text-to-video model. https://huggingface.co/cerspense/zeroscope_v2_576w , accessed: 2023-10-31
2023
Earlier work this paper cites.
Bahmani, S., Park, J.J., Paschalidou, D., Tang, H., Wetzstein, G., Guibas, L., Van Gool, L., Timofte, R.: 3D-aware video generation. TMLR (2023)
2023
Earlier work this paper cites.
Bahmani, S., Park, J.J., Paschalidou, D., Yan, X., Wetzstein, G., Guibas, L., Tagliasacchi, A.: CC3D: Layout-conditioned generation of compositional 3D scenes. Proc. ICCV (2023)
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
Blattmann, A., Rombach, R., Ling, H., Dockhorn, T., Kim, S.W., Fidler, S., Kreis, K.: Align your latents: High-resolution video synthesis with latent diffusion models. Proc. CVPR (2023)
2023
Earlier work this paper cites.
Chan, E.R., Nagano, K., Chan, M.A., Bergman, A.W., Park, J.J., Levy, A., Aittala, M., De Mello, S., Karras, T., Wetzstein, G.: Generative novel view synthesis with 3D-aware diffusion models. Proc. ICCV (2023)
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
Cohen-Bar, D., Richardson, E., Metzer, G., Giryes, R., Cohen-Or, D.: Set-the-scene: Global-local training for generating controllable NeRF scenes. Proc. ICCV Workshops (2023)
2023
Earlier work this paper cites.
Gao, W., Aigerman, N., Groueix, T., Kim, V., Hanocka, R.: TextDeformer: Geometry manipulation using text guidance. Proc. SIGGRAPH (2023)
2023
Earlier work this paper cites.
Gu, J., Trevithick, A., Lin, K.E., Susskind, J.M., Theobalt, C., Liu, L., Ramamoorthi, R.: NerfDiff: Single-image view synthesis with NeRF-guided distillation from 3D-aware diffusion. Proc. ICML (2023)
2023
Earlier work this paper cites.
Guo, Y.C., Liu, Y.T., Shao, R., Laforte, C., Voleti, V., Luo, G., Chen, C.H., Zou, Z.X., Wang, C., Cao, Y.P., Zhang, S.H.: threestudio: A unified framework for 3D content generation. https://github.com/threestudio-project/threestudio (2023)
2023
Earlier work this paper cites.
Jiang, Y., Zhang, L., Gao, J., Hu, W., Yao, Y.: Consistent4D: Consistent 360 ∘
2023
Earlier work this paper cites.
Kim, S.W., Brown, B., Yin, K., Kreis, K., Schwarz, K., Li, D., Rombach, R., Torralba, A., Fidler, S.: NeuralField-LDM: Scene generation with hierarchical latent diffusion models. Proc. CVPR (2023)
2023
Earlier work this paper cites.
Li, R., Tancik, M., Kanazawa, A.: NerfAcc: A general NeRF acceleration toolbox. Proc. ICCV (2023)
2023
Earlier work this paper cites.
2023
Cited alongside, same era.
Lin, C.H., Gao, J., Tang, L., Takikawa, T., Zeng, X., Huang, X., Kreis, K., Fidler, S., Liu, M.Y., Lin, T.Y.: Magic3D: High-resolution text-to-3D content creation. Proc. CVPR (2023)
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Liu, R., Wu, R., Van Hoorick, B., Tokmakov, P., Zakharov, S., Vondrick, C.: Zero-1-to-3: Zero-shot one image to 3D object. Proc. ICCV (2023)
2023
Cited alongside, same era.
2024
Closest in time.
Katzir, O., Patashnik, O., Cohen-Or, D., Lischinski, D.: Noise-free score distillation. Proc. ICLR (2024)
2024
Closest in time.
Lee, K., Sohn, K., Shin, J.: DreamFlow: High-quality text-to-3D generation by approximating probability flow. Proc. ICLR (2024)
2024
Closest in time.
Li, J., Tan, H., Zhang, K., Xu, Z., Luan, F., Xu, Y., Hong, Y., Sunkavalli, K., Shakhnarovich, G., Bi, S.: Instant3D: Fast text-to-3D with sparse-view generation and large reconstruction model. Proc. ICLR (2024)
2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Masood, M., Nawaz, M., Malik, K.M., Javed, A., Irtaza, A., Malik, H.: Deepfakes generation and detection: State-of-the-art, open challenges, countermeasures, and way forward. Applied Intelligence 53
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Poole, B., Jain, A., Barron, J.T., Mildenhall, B.: DreamFusion: Text-to-3D using 2D diffusion. In: Proc. ICLR (2023)
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Singer, U., Polyak, A., Hayes, T., Yin, X., An, J., Zhang, S., Hu, Q., Yang, H., Ashual, O., Gafni, O., et al.: Make-a-video: Text-to-video generation without text-video data. Proc. ICLR (2023)
2023
Cited alongside, same era.
Singer, U., Sheynin, S., Polyak, A., Ashual, O., Makarov, I., Kokkinos, F., Goyal, N., Vedaldi, A., Parikh, D., Johnson, J., et al.: Text-to-4D dynamic scene generation. Proc. ICML (2023)
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Tewari, A., Yin, T., Cazenavette, G., Rezchikov, S., Tenenbaum, J., Durand, F., Freeman, B., Sitzmann, V.: Diffusion with forward models: Solving stochastic inverse problems without direct supervision. Proc. NeurIPS (2023)
2023
Cited alongside, same era.
2024
Closest in time.
Ling, H., Kim, S.W., Torralba, A., Fidler, S., Kreis, K.: Align your Gaussians: Text-to-4D with dynamic 3D Gaussians and composed diffusion models. Proc. CVPR (2024)
2024
Closest in time.
2024
Closest in time.
Liu, X., Zhan, X., Tang, J., Shan, Y., Zeng, G., Lin, D., Liu, X., Liu, Z.: HumanGaussian: Text-driven 3D human generation with Gaussian splatting. Proc. CVPR (2024)
2024
Closest in time.
Liu, Y., Lin, C., Zeng, Z., Long, X., Liu, L., Komura, T., Wang, W.: SyncDreamer: Generating multiview-consistent images from a single-view image. Proc. ICLR (2024)
2024
Closest in time.
Long, X., Guo, Y.C., Lin, C., Liu, Y., Dou, Z., Liu, L., Ma, Y., Zhang, S.H., Habermann, M., Theobalt, C., et al.: Wonder3D: Single image to 3D using cross-domain diffusion. Proc. CVPR (2024)
2024
Closest in time.
2024
Closest in time.
Menapace, W., Siarohin, A., Skorokhodov, I., Deyneka, E., Chen, T.S., Kag, A., Fang, Y., Stoliar, A., Ricci, E., Ren, J., et al.: Snap video: Scaled spatiotemporal transformers for text-to-video synthesis. Proc. CVPR (2024)
2024
Closest in time.
2024
Closest in time.
Po, R., Wetzstein, G.: Compositional 3D scene generation using locally conditioned diffusion. Proc. 3DV (2024)
2024
Closest in time.
2024
Closest in time.
Qian, G., Mai, J., Hamdi, A., Ren, J., Siarohin, A., Li, B., Lee, H.Y., Skorokhodov, I., Wonka, P., Tulyakov, S., et al.: Magic123: One image to high-quality 3D object generation using both 2D and 3D diffusion priors. Proc. ICLR (2024)
2024
Closest in time.
2024
Closest in time.
Shi, Y., Wang, P., Ye, J., Mai, L., Li, K., Yang, X.: MVDream: Multi-view diffusion for 3D generation. Proc. ICLR (2024)
2024
Closest in time.
Sun, J., Zhang, B., Shao, R., Wang, L., Liu, W., Xie, Z., Liu, Y.: DreamCraft3D: Hierarchical 3D generation with bootstrapped diffusion prior. Proc. ICLR (2024)
2024
Closest in time.
Szymanowicz, S., Rupprecht, C., Vedaldi, A.: Splatter image: Ultra-fast single-view 3D reconstruction. Proc. CVPR (2024)
2024
Closest in time.
Tang, J., Chen, Z., Chen, X., Wang, T., Zeng, G., Liu, Z.: LGM: Large multi-view gaussian model for high-resolution 3d content creation. Proc. ECCV (2024)
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Wan, Z., Paschalidou, D., Huang, I., Liu, H., Shen, B., Xiang, X., Liao, J., Guibas, L.: CAD: Photorealistic 3D generation via adversarial distillation. Proc. CVPR (2024)
2024
Closest in time.
2024
Closest in time.
Wu, T., Yang, G., Li, Z., Zhang, K., Liu, Z., Guibas, L., Lin, D., Wetzstein, G.: GPT-4V(ision) is a human-aligned evaluator for text-to-3D generation. Proc. CVPR (2024)
2024
Closest in time.
Xie, K., Lorraine, J., Cao, T., Gao, J., Lucas, J., Torralba, A., Fidler, S., Zeng, X.: LATTE3D: Large-scale amortized text-to-enhanced3D synthesis. Proc. ECCV (2024)
2024
Closest in time.
Xu, Y., Shi, Z., Yifan, W., Chen, H., Yang, C., Peng, S., Shen, Y., Wetzstein, G.: GRM: Large Gaussian reconstruction model for efficient 3D reconstruction and generation. Proc. ECCV (2024)
2024
Closest in time.
Xu, Y., Tan, H., Luan, F., Bi, S., Wang, P., Li, J., Shi, Z., Sunkavalli, K., Wetzstein, G., Xu, Z., et al.: DMV3D: Denoising multi-view diffusion using 3D large reconstruction model. Proc. ICLR (2024)
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Yunus, R., Lenssen, J.E., Niemeyer, M., Liao, Y., Rupprecht, C., Theobalt, C., Pons-Moll, G., Huang, J.B., Golyanik, V., Ilg, E.: Recent trends in 3D reconstruction of general non-rigid scenes. Computer Graphics Forum (2024)
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Zhang, Q., Wang, C., Siarohin, A., Zhuang, P., Xu, Y., Yang, C., Lin, D., Zhou, B., Tulyakov, S., Lee, H.Y.: SceneWiz3D: Towards text-guided 3D scene composition. Proc. CVPR (2024)
2024
Closest in time.
Zheng, Y., Li, X., Nagano, K., Liu, S., Hilliges, O., De Mello, S.: A unified approach for text-and image-guided 4D scene generation. Proc. CVPR (2024)
2024
Closest in time.