Fetching the paper…
Reading the bibliography…
We present a latent diffusion model over 3D scenes, that can be trained using only 2D image data.
A volumetric method for building complex models from range images
B. Curless and M. Levoy · 1996
Earlier work this paper cites.
The lumigraph
S. J. Gortler, R. Grzeszczuk, R. Szeliski, and M. F. Cohen · 1996
Earlier work this paper cites.
Photorealistic scene reconstruction by voxel coloring
S. M. Seitz and C. R. Dyer · 1999
Earlier work this paper cites.
Multiple view geometry in computer vision
R. Hartley and A. Zisserman · 2003
Earlier work this paper cites.
A comparison and evaluation of multi-view stereo reconstruction algorithms
S. M. Seitz, B. Curless, J. Diebel, D. Scharstein, and R. Szeliski · 2006
Earlier work this paper cites.
Photo tourism: exploring photo collections in 3d
N. Snavely, S. M. Seitz, and R. Szeliski · 2006
Earlier work this paper cites.
Generative adversarial nets
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio · 2014
Earlier work this paper cites.
Auto-encoding variational bayes
D. P. Kingma and M. Welling · 2014
Earlier work this paper cites.
Stochastic backpropagation and approximate inference in deep generative models
D. J. Rezende, S. Mohamed, and D. Wierstra · 2014
Earlier work this paper cites.
Shapenet: An information-rich 3d model repository
A. X. Chang, T. Funkhouser, L. Guibas, P. Hanrahan, Q. Huang, Z. Li, S. Savarese, M. Savva, S. Song, H. Su, et al · 2015
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
O. Ronneberger, P. Fischer, and T. Brox · 2015
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
J. Sohl-Dickstein, E. A. Weiss, N. Maheswaranathan, and S. Ganguli · 2015
Earlier work this paper cites.
Structure-from-motion revisited
J. L. Schönberger and J.-M. Frahm · 2016
Earlier work this paper cites.
Wavenet: A generative model for raw audio
A. van den Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. Senior, and K. Kavukcuoglu · 2016
Earlier work this paper cites.
Pixel recurrent neural networks
A. Van Den Oord, N. Kalchbrenner, and K. Kavukcuoglu · 2016
Earlier work this paper cites.
Joint 2d-3d-semantic data for indoor scene understanding
I. Armeni, S. Sax, A. R. Zamir, and S. Savarese · 2017
Earlier work this paper cites.
Scannet: Richly-annotated 3d reconstructions of indoor scenes
A. Dai, A. X. Chang, M. Savva, M. Halber, T. Funkhouser, and M. Nießner · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. u. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Neural scene representation and rendering
S. A. Eslami, D. J. Rezende, F. Besse, F. Viola, A. S. Morcos, M. Garnelo, A. Ruderman, A. A. Rusu, I. Danihelka, K. Gregor, et al · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
A. Radford and K. Narasimhan · 2018
Earlier work this paper cites.
The unreasonable effectiveness of deep features as a perceptual metric
R. Zhang, P. Isola, A. A. Efros, E. Shechtman, and O. Wang · 2018
Earlier work this paper cites.
Stereo magnification: Learning view synthesis using multiplane images
T. Zhou, R. Tucker, J. Flynn, G. Fyffe, and N. Snavely · 2018
Earlier work this paper cites.
Learning single-image 3D reconstruction by generative modelling of shape, pose and shading
P. Henderson and V. Ferrari · 2019
Earlier work this paper cites.
A style-based generator architecture for generative adversarial networks
T. Karras, S. Laine, and T. Aila · 2019
Earlier work this paper cites.
Decoupled weight decay regularization
I. Loshchilov and F. Hutter · 2019
Earlier work this paper cites.
Hologan: Unsupervised learning of 3d representations from natural images
T. Nguyen-Phuoc, C. Li, L. Theis, C. Richardt, and Y.-L. Yang · 2019
Earlier work this paper cites.
Deepsdf: Learning continuous signed distance functions for shape representation
J. J. Park, P. Florence, J. Straub, R. Newcombe, and S. Lovegrove · 2019
Earlier work this paper cites.
Deepvoxels: Learning persistent 3d feature embeddings
V. Sitzmann, J. Thies, F. Heide, M. Nießner, G. Wetzstein, and M. Zollhöfer · 2019
Earlier work this paper cites.
The Pile: An 800gb dataset of diverse text for language modeling
L. Gao, S. Biderman, S. Black, L. Golding, T. Hoppe, C. Foster, J. Phang, H. He, A. Thite, N. Nabeshima, S. Presser, and C. Leahy · 2020
Earlier work this paper cites.
Leveraging 2D data to learn textured 3D mesh generation
P. Henderson, V. Tsiminaki, and C. Lampert · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
J. Ho, A. Jain, and P. Abbeel · 2020
Earlier work this paper cites.
Nerf: Representing scenes as neural radiance fields for view synthesis
B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Barron, R. Ramamoorthi, and R. Ng · 2020
Earlier work this paper cites.
Blockgan: Learning 3d object-aware scene representations from unlabelled images
T. Nguyen-Phuoc, C. Richardt, L. Mai, Y.-L. Yang, and N. Mitra · 2020
Earlier work this paper cites.
Convolutional occupancy networks
S. Peng, M. Niemeyer, L. Mescheder, M. Pollefeys, and A. Geiger · 2020
Earlier work this paper cites.
GRAF: generative radiance fields for 3d-aware image synthesis
K. Schwarz, Y. Liao, M. Niemeyer, and A. Geiger · 2020
Earlier work this paper cites.
Synsin: End-to-end view synthesis from a single image
O. Wiles, G. Gkioxari, R. Szeliski, and J. Johnson · 2020
Earlier work this paper cites.
Mip-nerf: A multiscale representation for anti-aliasing neural radiance fields
J. T. Barron, B. Mildenhall, M. Tancik, P. Hedman, R. Martin-Brualla, and P. P. Srinivasan · 2021
Earlier work this paper cites.
Mvsnerf: Fast generalizable radiance field reconstruction from multi-view stereo
A. Chen, Z. Xu, F. Zhao, X. Zhang, F. Xiang, J. Yu, and H. Su · 2021
Earlier work this paper cites.
Unconstrained scene generation with locally conditioned radiance fields
T. Devries, M. Á. Bautista, N. Srivastava, G. W. Taylor, and J. M. Susskind · 2021
Earlier work this paper cites.
Taming transformers for high-resolution image synthesis
P. Esser, R. Rombach, and B. Ommer · 2021
Earlier work this paper cites.
Unsupervised video prediction from a single frame by estimating 3d dynamic scene structure
P. Henderson, C. H. Lampert, and B. Bickel · 2021
Earlier work this paper cites.
Unsupervised learning of 3d object categories from videos in the wild
P. Henzler, J. Reizenstein, P. Labatut, R. Shapovalov, T. Ritschel, A. Vedaldi, and D. Novotny · 2021
Earlier work this paper cites.
Nerf-vae: A geometry aware 3d scene generative model
A. R. Kosiorek, H. Strathmann, D. Zoran, P. Moreno, R. Schneider, S. Mokrá, and D. J. Rezende · 2021
Earlier work this paper cites.
Diffusion probabilistic models for 3d point cloud generation
S. Luo and W. Hu · 2021
Earlier work this paper cites.
Common objects in 3d: Large-scale learning and evaluation of real-life 3d category reconstruction
J. Reizenstein, R. Shapovalov, P. Henzler, L. Sbordone, P. Labatut, and D. Novotny · 2021
Earlier work this paper cites.
Pixelsynth: Generating a 3d-consistent experience from a single image
C. Rockwell, D. F. Fouhey, and J. Johnson · 2021
Earlier work this paper cites.
Geometry-free view synthesis: Transformers and no 3d priors
R. Rombach, P. Esser, and B. Ommer · 2021
Earlier work this paper cites.
Denoising diffusion implicit models
J. Song, C. Meng, and S. Ermon · 2021
Cited alongside, same era.
Ibrnet: Learning multi-view image-based rendering
Q. Wang, Z. Wang, K. Genova, P. Srinivasan, H. Zhou, J. T. Barron, R. Martin-Brualla, N. Snavely, and T. Funkhouser · 2021
Cited alongside, same era.
pixelnerf: Neural radiance fields from one or few images
A. Yu, V. Ye, M. Tancik, and A. Kanazawa · 2021
Cited alongside, same era.
3d shape generation and completion through point-voxel diffusion
L. Zhou, Y. Du, and J. Wu · 2021
Cited alongside, same era.
Unsupervised causal generative understanding of images
T. Anciukevicius, P. Fox-Roberts, E. Rosten, and P. Henderson · 2022
Cited alongside, same era.
Gaudi: A neural architect for immersive 3d scene generation
M. A. Bautista, P. Guo, S. Abnar, W. Talbott, A. Toshev, Z. Chen, L. Dinh, S. Zhai, H. Goh, D. Ulbricht, A. Dehghan, and J. Susskind · 2022
Zero-1-to-3: Zero-shot one image to 3d object
R. Liu, R. Wu, B. V. Hoorick, P. Tokmakov, S. Zakharov, and C. Vondrick · 2023
Later among the works it cites.
Realfusion: 360° reconstruction of any object from a single image
L. Melas-Kyriazi, C. Rupprecht, I. Laina, and A. Vedaldi · 2023
Later among the works it cites.
Diffrf: Rendering-guided 3d radiance field diffusion
N. Müller, Y. Siddiqui, L. Porzi, S. R. Bulo, P. Kontschieder, and M. Nießner · 2023
Later among the works it cites.
Ganerf: Leveraging discriminators to optimize neural radiance fields
B. Roessle, N. Müller, L. Porzi, S. R. Bulò, P. Kontschieder, and M. Nießner · 2023
Later among the works it cites.
Viewset diffusion: (0-)image-conditioned 3D generative models from 2D data
S. Szymanowicz, C. Rupprecht, and A. Vedaldi · 2023
Later among the works it cites.
MVDiffusion: Enabling holistic multi-view image generation with correspondence-aware diffusion
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Efficient geometry-aware 3D generative adversarial networks
E. R. Chan, C. Z. Lin, M. A. Chan, K. Nagano, B. Pan, S. D. Mello, O. Gallo, L. Guibas, J. Tremblay, S. Khamis, T. Karras, and G. Wetzstein · 2022
Cited alongside, same era.
Tensorf: Tensorial radiance fields
A. Chen, Z. Xu, A. Geiger, J. Yu, and H. Su · 2022
Cited alongside, same era.
Gram: Generative radiance manifolds for 3d-aware image generation
Y. Deng, J. Yang, J. Xiang, and X. Tong · 2022
Cited alongside, same era.
Classifier-free diffusion guidance
J. Ho · 2022
Cited alongside, same era.
Imagen video: High definition video generation with diffusion models
J. Ho, W. Chan, C. Saharia, J. Whang, R. Gao, A. Gritsenko, D. P. Kingma, B. Poole, M. Norouzi, D. J. Fleet, et al · 2022
Cited alongside, same era.
Neural wavelet-domain diffusion for 3d shape generation
K.-H. Hui, R. Li, J. Hu, and C.-W. Fu · 2022
Cited alongside, same era.
S. Tang, F. Zhang, J. Chen, P. Wang, and Y. Furukawa · 2023
Later among the works it cites.
Diffusion with forward models: Solving stochastic inverse problems without direct supervision
A. Tewari, T. Yin, G. Cazenavette, S. Rezchikov, J. B. Tenenbaum, F. Durand, W. T. Freeman, and V. Sitzmann · 2023
Later among the works it cites.
Consistent view synthesis with pose-guided diffusion models
H.-Y. Tseng, Q. Li, C. Kim, S. Alsisan, J.-B. Huang, and J. Kopf · 2023
Later among the works it cites.
Score jacobian chaining: Lifting pretrained 2d diffusion models for 3d generation
H. Wang, X. Du, J. Li, R. A. Yeh, and G. Shakhnarovich · 2023
Later among the works it cites.
Novel view synthesis with diffusion models
D. Watson, W. Chan, R. M. Brualla, J. Ho, A. Tagliasacchi, and M. Norouzi · 2023
Later among the works it cites.
Multiview compressive coding for 3d reconstruction
C.-Y. Wu, J. Johnson, J. Malik, C. Feichtenhofer, and G. Gkioxari · 2023
Later among the works it cites.
Reconfusion: 3d reconstruction with diffusion priors
R. Wu, B. Mildenhall, P. Henzler, K. Park, R. Gao, D. Watson, P. P. Srinivasan, D. Verbin, J. T. Barron, B. Poole, and A. Holynski · 2023
Later among the works it cites.
Omniobject3d: Large-vocabulary 3d object dataset for realistic perception, reconstruction and generation
T. Wu, J. Zhang, X. Fu, Y. Wang, L. P. Jiawei Ren, W. Wu, L. Yang, J. Wang, C. Qian, D. Lin, and Z. Liu · 2023
Later among the works it cites.
DiffusioNeRF: Regularizing Neural Radiance Fields with Denoising Diffusion Models
J. Wynn and D. Turmukhambetov · 2023
Later among the works it cites.
Dreamsparse: Escaping from plato’s cave with 2d frozen diffusion model given sparse views
P. Yoo, J. Guo, Y. Matsuo, and S. S. Gu · 2023
Later among the works it cites.
Long-term photometric consistent novel view synthesis with diffusion models
J. J. Yu, F. Forghani, K. G. Derpanis, and M. A. Brubaker · 2023
Later among the works it cites.
Mvimgnet: A large-scale dataset of multi-view images
X. Yu, M. Xu, Y. Zhang, H. Liu, C. Ye, Y. Wu, Z. Yan, T. Liang, G. Chen, S. Cui, and X. Han · 2023
Later among the works it cites.
Sparsefusion: Distilling view-conditioned diffusion for 3d reconstruction
Z. Zhou and S. Tulsiani · 2023
Later among the works it cites.
Hifa: High-fidelity text-to-3d with advanced diffusion guidance
J. Zhu and P. Zhuang · 2023
Later among the works it cites.
Sparse3d: Distilling multiview-consistent diffusion for object reconstruction from sparse views
Z.-X. Zou, W. Cheng, Y.-P. Cao, S.-S. Huang, Y. Shan, and S.-H. Zhang · 2023
Later among the works it cites.
Triplane meets gaussian splatting: Fast and generalizable single-view 3d reconstruction with transformers
Z.-X. Zou, Z. Yu, Y.-C. Guo, Y. Li, D. Liang, Y.-P. Cao, and S.-H. Zhang · 2023
Later among the works it cites.
Denoising diffusion via image-based rendering
T. Anciukevičius, F. Manhardt, F. Tombari, and P. Henderson · 2024
Closest in time.
Lightplane: Highly-scalable components for neural 3d fields
A. Cao, J. Johnson, A. Vedaldi, and D. Novotny · 2024
Closest in time.
Mvsplat: Efficient 3d gaussian splatting from sparse multi-view images
Y. Chen, H. Xu, C. Zheng, B. Zhuang, M. Pollefeys, A. Geiger, T.-J. Cham, and J. Cai · 2024
Closest in time.
Objaverse-xl: A universe of 10m+ 3d objects
M. Deitke, R. Liu, M. Wallingford, H. Ngo, O. Michel, A. Kusupati, A. Fan, C. Laforte, V. Voleti, S. Y. Gadre, et al · 2024
Closest in time.
Cat3d: Create anything in 3d with multi-view diffusion models
R. Gao, A. Holynski, P. Henzler, A. Brussee, R. Martin-Brualla, P. Srinivasan, J. T. Barron, and B. Poole · 2024
Closest in time.
Viewdiff: 3d-consistent image generation with text-to-image models
L. Höllein, A. Božič, N. Müller, D. Novotny, H.-Y. Tseng, C. Richardt, M. Zollhöfer, and M. Nießner · 2024
Closest in time.
Mvd-fusion: Single-view 3d via depth-consistent multi-view generation
H. Hu, Z. Zhou, V. Jampani, and S. Tulsiani · 2024
Closest in time.
Eschernet: A generative model for scalable view synthesis
X. Kong, S. Liu, X. Lyu, M. Taher, X. Qi, and A. J. Davison · 2024
Closest in time.
One-2-3-45: Any single image to 3d mesh in 45 seconds without per-shape optimization
M. Liu, C. Xu, H. Jin, L. Chen, M. Varma T, Z. Xu, and H. Su · 2024
Closest in time.
Syncdreamer: Generating multiview-consistent images from a single-view image
Y. Liu, C. Lin, Z. Zeng, X. Long, L. Liu, T. Komura, and W. Wang · 2024
Closest in time.
Wildfusion: Learning 3d-aware latent diffusion models in view space
K. Schwarz, S. Wook Kim, J. Gao, S. Fidler, A. Geiger, and K. Kreis · 2024
Closest in time.
Gamba: Marry gaussian splatting with mamba for single view 3d reconstruction
Q. Shen, X. Yi, Z. Wu, P. Zhou, H. Zhang, S. Yan, and X. Wang · 2024
Closest in time.
RealmDreamer: Text-driven 3d scene generation with inpainting and depth diffusion
J. Shriram, A. Trevithick, L. Liu, and R. Ramamoorthi · 2024
Closest in time.
Splatter image: Ultra-fast single-view 3d reconstruction
S. Szymanowicz, C. Rupprecht, and A. Vedaldi · 2024
Closest in time.
Lgm: Large multi-view gaussian model for high-resolution 3d content creation
J. Tang, Z. Chen, X. Chen, T. Wang, G. Zeng, and Z. Liu · 2024
Closest in time.
Dreamgaussian: Generative gaussian splatting for efficient 3d content creation
J. Tang, J. Ren, H. Zhou, Z. Liu, and G. Zeng · 2024
Closest in time.
S. Tang, J. Chen, D. Wang, C. Tang, F. Zhang, Y. Fan, V. Chandra, Y. Furukawa, and R. Ranjan · 2024
Closest in time.
Prolificdreamer: High-fidelity and diverse text-to-3d generation with variational score distillation
Z. Wang, C. Lu, Y. Wang, F. Bao, C. Li, H. Su, and J. Zhu · 2024
Closest in time.
latentsplat: Autoencoding variational gaussians for fast generalizable 3d reconstruction
C. Wewer, K. Raj, E. Ilg, B. Schiele, and J. E. Lenssen · 2024
Closest in time.
DMV3d: Denoising multi-view diffusion using 3d large reconstruction model
Y. Xu, H. Tan, F. Luan, S. Bi, P. Wang, J. Li, Z. Shi, K. Sunkavalli, G. Wetzstein, Z. Xu, and K. Zhang · 2024
Closest in time.
Gaussiandreamer: Fast generation from text to 3d gaussians by bridging 2d and 3d diffusion models
T. Yi, J. Fang, J. Wang, G. Wu, L. Xie, X. Zhang, W. Liu, Q. Tian, and X. Wang · 2024
Closest in time.
Grm: Large gaussian reconstruction model for efficient 3d reconstruction and generation
X. Yinghao, S. Zifan, Y. Wang, C. Hansheng, Y. Ceyuan, P. Sida, S. Yujun, and W. Gordon · 2024
Closest in time.
Gaussiancube: Structuring gaussian splatting using optimal transport for 3d generative modeling
B. Zhang, Y. Cheng, J. Yang, C. Wang, F. Zhao, Y. Tang, D. Chen, and B. Guo · 2024
Closest in time.
Gs-lrm: Large reconstruction model for 3d gaussian splatting
K. Zhang, S. Bi, H. Tan, Y. Xiangli, N. Zhao, K. Sunkavalli, and Z. Xu · 2024
Closest in time.
Gps-gaussian: Generalizable pixel-wise 3d gaussian splatting for real-time human novel view synthesis
S. Zheng, B. Zhou, R. Shao, B. Liu, S. Zhang, L. Nie, and Y. Liu · 2024
Closest in time.