Fetching the paper…
Reading the bibliography…
We present a two-stage text-to-3D generation system, namely 3DTopia, which generates high-quality general 3D assets within 5 minutes using hybrid diffusion priors.
Lorensen WE, Cline HE (1998) Marching cubes: A high resolution 3d surface construction algorithm. In: Seminal graphics: pioneering efforts that shaped the field, pp 347–353
1998
Earlier work this paper cites.
2015
Earlier work this paper cites.
Dai A, Chang AX, Savva M, Halber M, Funkhouser TA, Nießner M (2017) Scannet: Richly-annotated 3d reconstructions of indoor scenes. In: CVPR, pp 2432–2443, DOI 10.1109/CVPR.2017.261
2017
Earlier work this paper cites.
Nash C, Williams CK (2017) The shape variational autoencoder: A deep generative model of part-segmented 3d objects. In: Computer Graphics Forum, Wiley Online Library, vol 36, pp 1–12
2017
Earlier work this paper cites.
Chen K, Choy CB, Savva M, Chang AX, Funkhouser TA, Savarese S (2018) Text2shape: Generating shapes from natural language by learning joint embeddings. In: ACCV, vol 11363, pp 100–116, DOI 10.1007/978-3-030-20893-6\_7
2018
Earlier work this paper cites.
Chen Z, Zhang H (2019) Learning implicit fields for generative shape modeling. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp 5939–5948
2019
Earlier work this paper cites.
Ho J, Jain A, Abbeel P (2020) Denoising diffusion probabilistic models. Advances in neural information processing systems 33:6840–6851
2020
Earlier work this paper cites.
Laine S, Hellsten J, Karras T, Seol Y, Lehtinen J, Aila T (2020) Modular primitives for high-performance differentiable rendering. ToG 39(6)
2020
Earlier work this paper cites.
Mildenhall B, Srinivasan PP, Tancik M, Barron JT, Ramamoorthi R, Ng R (2020) Nerf: Representing scenes as neural radiance fields for view synthesis. In: ECCV
2020
Earlier work this paper cites.
Song J, Meng C, Ermon S (2020) Denoising diffusion implicit models. arXiv preprint arXiv:201002502
2020
Earlier work this paper cites.
Chen DZ, Gholami A, Nießner M, Chang AX (2021) Scan2cap: Context-aware dense captioning in RGB-D scans. In: CVPR, pp 3193–3203, DOI 10.1109/CVPR46437.2021.00321
2021
Earlier work this paper cites.
Radford A, Kim JW, Hallacy C, Ramesh A, Goh G, Agarwal S, Sastry G, Askell A, Mishkin P, Clark J, Krueger G, Sutskever I (2021) Learning transferable visual models from natural language supervision. In: ICML, vol 139, pp 8748–8763, URL http://proceedings.mlr.press/v139/radford21a.html
2021
Earlier work this paper cites.
Wang P, Liu L, Liu Y, Theobalt C, Komura T, Wang W (2021) Neus: Learning neural implicit surfaces by volume rendering for multi-view reconstruction. NeurIPS
2021
Earlier work this paper cites.
Chan ER, Lin CZ, Chan MA, Nagano K, Pan B, De Mello S, Gallo O, Guibas LJ, Tremblay J, Khamis S, et al. (2022) Efficient geometry-aware 3d generative adversarial networks. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp 16123–16133
2022
Earlier work this paper cites.
Chen A, Xu Z, Geiger A, Yu J, Su H (2022) Tensorf: Tensorial radiance fields. In: ECCV, Springer, pp 333–350
2022
Earlier work this paper cites.
Gao J, Shen T, Wang Z, Chen W, Yin K, Li D, Litany O, Gojcic Z, Fidler S (2022) Get3d: A generative model of high quality 3d textured shapes learned from images. Advances In Neural Information Processing Systems 35:31841–31854
2022
Earlier work this paper cites.
Ho J, Salimans T (2022) Classifier-free diffusion guidance. arXiv preprint arXiv:220712598
2022
Earlier work this paper cites.
Jain A, Mildenhall B, Barron JT, Abbeel P, Poole B (2022) Zero-shot text-guided object generation with dream fields. In: CVPR, pp 867–876
2022
Earlier work this paper cites.
Metzer G, Richardson E, Patashnik O, Giryes R, Cohen-Or D (2022) Latent-nerf for shape-guided generation of 3d shapes and textures. arXiv preprint arXiv:221107600
2022
Earlier work this paper cites.
Michel O, Bar-On R, Liu R, Benaim S, Hanocka R (2022) Text2mesh: Text-driven neural stylization for meshes. In: CVPR, pp 13492–13502
2022
Earlier work this paper cites.
Mohammad Khalid N, Xie T, Belilovsky E, Popa T (2022) Clip-mesh: Generating textured meshes from text using pretrained image-text models. In: SIGGRAPH Asia, pp 1–8
2022
Cited alongside, same era.
Müller T, Evans A, Schied C, Keller A (2022) Instant neural graphics primitives with a multiresolution hash encoding. ACM TOG
2022
Cited alongside, same era.
Nam G, Khlifi M, Rodriguez A, Tono A, Zhou L, Guerrero P (2022) 3d-ldm: Neural implicit 3d shape generation with latent diffusion models. arXiv preprint arXiv:221200842
2022
Cited alongside, same era.
Nichol A, Jun H, Dhariwal P, Mishkin P, Chen M (2022) Point-e: A system for generating 3d point clouds from complex prompts. arXiv preprint arXiv:221208751
2022
Cited alongside, same era.
Poole B, Jain A, Barron JT, Mildenhall B (2022) Dreamfusion: Text-to-3d using 2d diffusion. arXiv preprint arXiv:220914988
Lin CH, Gao J, Tang L, Takikawa T, Zeng X, Huang X, Kreis K, Fidler S, Liu MY, Lin TY (2023) Magic3d: High-resolution text-to-3d content creation. In: CVPR, pp 300–309
2023
Later among the works it cites.
Luo T, Rockwell C, Lee H, Johnson J (2023) Scalable 3d captioning with pretrained models. arXiv preprint arXiv:230607279 DOI 10.48550/ARXIV.2306.07279
2023
Later among the works it cites.
Müller N, Siddiqui Y, Porzi L, Bulo SR, Kontschieder P, Nießner M (2023) Diffrf: Rendering-guided 3d radiance field diffusion. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp 4328–4338
2023
Later among the works it cites.
Qiu L, Chen G, Gu X, zuo Q, Xu M, Wu Y, Yuan W, Dong Z, Bo L, Han X (2023) Richdreamer: A generalizable normal-depth diffusion model for detail richness in text-to-3d. arXiv preprint arXiv:231116918
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2022
Cited alongside, same era.
Rombach R, Blattmann A, Lorenz D, Esser P, Ommer B (2022) High-resolution image synthesis with latent diffusion models. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp 10684–10695
2022
Cited alongside, same era.
Saharia C, Chan W, Saxena S, Li L, Whang J, Denton EL, Ghasemipour K, Gontijo Lopes R, Karagol Ayan B, Salimans T, et al. (2022) Photorealistic text-to-image diffusion models with deep language understanding. NeurIPS 35:36479–36494
2022
Cited alongside, same era.
Sara Fridovich-Keil and Alex Yu, Tancik M, Chen Q, Recht B, Kanazawa A (2022) Plenoxels: Radiance fields without neural networks. In: CVPR
2022
Cited alongside, same era.
Tang J, Zhou H, Chen X, Hu T, Ding E, Wang J, Zeng G (2022) Delicate textured mesh recovery from nerf via adaptive surface refinement. arXiv preprint arXiv:230302091
2022
Cited alongside, same era.
Xu Q, Xu Z, Philip J, Bi S, Shu Z, Sunkavalli K, Neumann U (2022) Point-nerf: Point-based neural radiance fields. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp 5438–5448
2022
Cited alongside, same era.
Zeng X, Vahdat A, Williams F, Gojcic Z, Litany O, Fidler S, Kreis K (2022) Lion: Latent point diffusion models for 3d shape generation. arXiv preprint arXiv:221006978
2022
Cited alongside, same era.
Betker J, Goh G, Jing L, TimBrooks, Wang J, Li L, LongOuyang, JuntangZhuang, JoyceLee, YufeiGuo, WesamManassra, PrafullaDhariwal, CaseyChu, YunxinJiao, Ramesh A (2023) Improving image generation with better captions. URL https://api.semanticscholar.org/CorpusID:264403242
2023
Cited alongside, same era.
Raj A, Kaza S, Poole B, Niemeyer M, Ruiz N, Mildenhall B, Zada S, Aberman K, Rubinstein M, Barron J, et al. (2023) Dreambooth3d: Subject-driven text-to-3d generation. arXiv preprint arXiv:230313508
2023
Later among the works it cites.
Sella E, Fiebelman G, Hedman P, Averbuch-Elor H (2023) Vox-e: Text-guided voxel editing of 3d objects. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp 430–440
2023
Later among the works it cites.
Shi Y, Wang P, Ye J, Long M, Li K, Yang X (2023) Mvdream: Multi-view diffusion for 3d generation. arXiv preprint arXiv:230816512
2023
Later among the works it cites.
Shue JR, Chan ER, Po R, Ankner Z, Wu J, Wetzstein G (2023) 3d neural field generation using triplane diffusion. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp 20875–20886
2023
Later among the works it cites.
Singer U, Sheynin S, Polyak A, Ashual O, Makarov I, Kokkinos F, Goyal N, Vedaldi A, Parikh D, Johnson J, et al. (2023) Text-to-4d dynamic scene generation. arXiv preprint arXiv:230111280
2023
Later among the works it cites.
Tang J, Ren J, Zhou H, Liu Z, Zeng G (2023) Dreamgaussian: Generative gaussian splatting for efficient 3d content creation. arXiv preprint arXiv:230916653
2023
Later among the works it cites.
Tsalicoglou C, Manhardt F, Tonioni A, Niemeyer M, Tombari F (2023) Textmesh: Generation of realistic 3d meshes from text prompts. arXiv preprint arXiv:230412439
2023
Later among the works it cites.
Wei X, Xiang F, Bi S, Chen A, Sunkavalli K, Xu Z, Su H (2023) Neumanifold: Neural watertight manifold reconstruction with efficient and high-quality rendering support. arXiv preprint arXiv:230517134
2023
Later among the works it cites.
Xie H, Chen Z, Hong F, Liu Z (2023) Citydreamer: Compositional generative model of unbounded 3d cities. arXiv preprint arXiv:230900610
2023
Later among the works it cites.
Xue L, Gao M, Xing C, Martín-Martín R, Wu J, Xiong C, Xu R, Niebles JC, Savarese S (2023a) ULIP: learning a unified representation of language, images, and point clouds for 3d understanding. In: CVPR, pp 1179–1189, DOI 10.1109/CVPR52729.2023.00120
2023
Later among the works it cites.
Yi T, Fang J, Wu G, Xie L, Zhang X, Liu W, Tian Q, Wang X (2023) Gaussiandreamer: Fast generation from text to 3d gaussian splatting with point cloud priors. arXiv preprint arXiv:231008529
2023
Later among the works it cites.
Yu C, Zhou Q, Li J, Zhang Z, Wang Z, Wang F (2023) Points-to-3d: Bridging the gap between sparse points and shape-controllable text-to-3d generation. arXiv preprint arXiv:230713908
2023
Later among the works it cites.
Zhu J, Zhuang P (2023) Hifa: High-fidelity text-to-3d with advanced diffusion guidance. arXiv preprint arXiv:230518766
2023
Later among the works it cites.
Zhuang J, Wang C, Liu L, Lin L, Li G (2023) Dreameditor: Text-driven 3d scene editing with neural fields. arXiv preprint arXiv:230613455
2023
Later among the works it cites.
Hu T, Hong F, Liu Z (2024) Structldm: Structured latent diffusion for 3d human generation. 2404.01241
2024
Closest in time.
Wu T, Yang G, Li Z, Zhang K, Liu Z, Guibas L, Lin D, Wetzstein G (2024) Gpt-4v (ision) is a human-aligned evaluator for text-to-3d generation. arXiv preprint arXiv:240104092
2024
Closest in time.