Fetching the paper…
Reading the bibliography…
How can one efficiently generate high-quality, wide-scope 3D scenes from arbitrary single images? Existing methods suffer several drawbacks, such as requiring multi-view data, time-consuming per-scene optimization, distorted geometry in occluded areas, and low visual quality in backgrounds.
Z. Wang, A. C. Bovik, H. R. Sheikh, and E. P. Simoncelli, “Image quality assessment: From error visibility to structural similarity,”
2004
Earlier work this paper cites.
P. Griffiths and J. Harris,
2014
Earlier work this paper cites.
K. Simonyan, “Very deep convolutional networks for large-scale image recognition,”
2014
Earlier work this paper cites.
A. Dosovitskiy and T. Brox, “Generating images with perceptual similarity metrics based on deep networks,”
2016
Earlier work this paper cites.
J. L. Schönberger and J.-M. Frahm, “Structure-from-motion revisited,” in
2016
Earlier work this paper cites.
J. L. Schönberger, E. Zheng, M. Pollefeys, and J.-M. Frahm, “Pixelwise view selection for unstructured multi-view stereo,” in
2016
Earlier work this paper cites.
P. K. Diederik and J. Ba, “Adam: A method for stochastic optimization,”
2017
Earlier work this paper cites.
M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter, “GANS trained by a two time-scale update rule converge to a local Nash equilibrium,”
2017
Earlier work this paper cites.
P. Isola, J.-Y. Zhu, T. Zhou, and A. A. Efros, “Image-to-image translation with conditional adversarial networks,” in
2017
Earlier work this paper cites.
A. Knapitsch, J. Park, Q.-Y. Zhou, and V. Koltun, “Tanks and temples: Benchmarking large-scale scene reconstruction,”
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
R. Zhang, P. Isola, A. A. Efros, E. Shechtman, and O. Wang, “The unreasonable effectiveness of deep features as a perceptual metric,” in
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
T. Unterthiner, S. van Steenkiste, K. Kurach, R. Marinier, M. Michalski, and S. Gelly, “FVD: A new metric for video generation,” in
2019
Earlier work this paper cites.
J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,”
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
M. Caron, H. Touvron, I. Misra, H. Jégou, J. Mairal, P. Bojanowski, and A. Joulin, “Emerging properties in self-supervised vision transformers,” in
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
A. Jain, M. Tancik, and P. Abbeel, “Putting NeRF on a diet: Semantically consistent few-shot view synthesis,” in
2021
Earlier work this paper cites.
A. Liu, R. Tucker, V. Jampani, A. Makadia, N. Snavely, and A. Kanazawa, “Infinite nature: Perpetual view generation of natural scenes from a single image,” in
2021
Earlier work this paper cites.
B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Barron, R. Ramamoorthi, and R. Ng, “NeRF: Representing scenes as neural radiance fields for view synthesis,”
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
J. T. Barron, B. Mildenhall, D. Verbin, P. P. Srinivasan, and P. Hedman, “Mip-NeRF 360: Unbounded anti-aliased neural radiance fields supplemental materials,” in
2022
Earlier work this paper cites.
K. He, X. Chen, S. Xie, Y. Li, P. Dollár, and R. Girshick, “Masked autoencoders are scalable vision learners,” in
2022
Earlier work this paper cites.
H. Liang, N. Quader, Z. Chi, L. Chen, P. Dai, J. Lu, and Y. Wang, “Self-supervised spatiotemporal representation learning by exploiting video continuity,” in
2022
Earlier work this paper cites.
B. Poole, A. Jain, J. T. Barron, and B. Mildenhall, “DreamFusion: Text-to-3D using 2D diffusion,” in
2022
Earlier work this paper cites.
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, “High-resolution image synthesis with latent diffusion models,” in
2022
Earlier work this paper cites.
S. Bahmani, J. J. Park, D. Paschalidou, X. Yan, G. Wetzstein, L. Guibas, and A. Tagliasacchi, “CC3D: Layout-conditioned generation of compositional 3D scenes,” in
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
A. Blattmann, R. Rombach, H. Ling, T. Dockhorn, S. W. Kim, S. Fidler, and K. Kreis, “Align your latents: High-resolution video synthesis with latent diffusion models,” in
2023
Earlier work this paper cites.
E. R. Chan, K. Nagano, M. A. Chan, A. W. Bergman, J. J. Park, A. Levy, M. Aittala, S. De Mello, T. Karras, and G. Wetzstein, “Generative novel view synthesis with 3D-aware diffusion models,” in
2023
Earlier work this paper cites.
D. Charatan, S. L. Li, A. Tagliasacchi, and V. Sitzmann, “pixelSplat: 3D Gaussian splats from image pairs for scalable generalizable 3D reconstruction,” in
2023
Earlier work this paper cites.
H. Chen, M. Xia, Y. He, Y. Zhang, X. Cun, S. Yang, J. Xing, Y. Liu, Q. Chen, X. Wang
2023
Earlier work this paper cites.
J. Chen, J. Yu, C. Ge, L. Yao, E. Xie, Y. Wu, Z. Wang, J. Kwok, P. Luo, H. Lu
2023
Earlier work this paper cites.
R. Chen, Y. Chen, N. Jiao, and K. Jia, “Fantasia3D: Disentangling geometry and appearance for high-quality text-to-3D content creation,” in
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
T. Dao, “FlashAttention-2: Faster attention with better parallelism and work partitioning,”
2023
Earlier work this paper cites.
J. Gu, A. Trevithick, K.-E. Lin, J. M. Susskind, C. Theobalt, L. Liu, and R. Ramamoorthi, “NerfDiff: Single-image view synthesis with NeRF-guided distillation from 3D-aware diffusion,” in
2023
Earlier work this paper cites.
Guangcong, Z. Chen, C. C. Loy, and Z. Liu, “SparseNeRF: Distilling depth ranking for few-shot novel view synthesis,” in
2023
Earlier work this paper cites.
2023
Cited alongside, same era.
B. Kerbl, G. Kopanas, T. Leimkühler, and G. Drettakis, “3D Gaussian splatting for real-time radiance field rendering,”
2023
Cited alongside, same era.
S. W. Kim, B. Brown, K. Yin, K. Kreis, K. Schwarz, D. Li, R. Rombach, A. Torralba, and S. Fidler, “NeuralField-LDM: Scene generation with hierarchical latent diffusion models,” in
2023
Cited alongside, same era.
2023
Cited alongside, same era.
T. Lu, T. Shu, J. Xiao, L. Ye, J. Wang, C. Peng, C. Wei, D. Khashabi, R. Chellappa, A. Yuille
2024
Closest in time.
2024
Closest in time.
W. Menapace, A. Siarohin, I. Skorokhodov, E. Deyneka, T.-S. Chen, A. Kag, Y. Fang, A. Stoliar, E. Ricci, J. Ren
2024
Closest in time.
G. Qian, J. Mai, A. Hamdi, J. Ren, A. Siarohin, B. Li, H. Lee, I. Skorokhodov, P. Wonka, S. Tulyakov, and B. Ghanem, “Magic123: One image to high-quality 3D object generation using both 2D and 3D diffusion priors,” in
2024
Closest in time.
J. Ren, K. Xie, A. Mirzaei, H. Liang, X. Zeng, K. Kreis, Z. Liu, A. Torralba, S. Fidler, S. W. Kim
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2023
Cited alongside, same era.
W. Peebles and S. Xie, “Scalable diffusion models with transformers,” in
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Z. Wang, C. Lu, Y. Wang, F. Bao, C. Li, H. Su, and J. Zhu, “ProlificDreamer: High-fidelity and diverse text-to-3D generation with variational score distillation,”
2023
Cited alongside, same era.
2023
Cited alongside, same era.
L. Zhang, A. Rao, and M. Agrawala, “Adding conditional control to text-to-image diffusion models,” in
2023
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
S. Szymanowicz, C. Rupprecht, and A. Vedaldi, “Splatter image: Ultra-fast single-view 3D reconstruction,” in
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
V. Voleti, C.-H. Yao, M. Boss, A. Letts, D. Pankratz, D. Tochilkin, C. Laforte, R. Rombach, and V. Jampani, “SV3D: Novel multi-view synthesis and 3D generation from a single image using latent video diffusion,” in
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Z. Wang, Z. Yuan, X. Wang, Y. Li, T. Chen, M. Xia, P. Luo, and Y. Shan, “MotionCtrl: A unified and flexible motion controller for video generation,” in
2024
Closest in time.
F. Xiao, X. Liu, X. Wang, S. Peng, M. Xia, X. Shi, Z. Yuan, P. Wan, D. Zhang, and D. Lin, “3DTrajMaster: Mastering 3D trajectory for multi-entity motion in video generation,” in
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Y. Xu, Z. Shi, W. Yifan, H. Chen, C. Yang, S. Peng, Y. Shen, and G. Wetzstein, “GRM: Large Gaussian reconstruction model for efficient 3D reconstruction and generation,” in
2024
Closest in time.
Z. Yang, Z. Pan, C. Gu, and L. Zhang, “Diffusion
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
H.-X. Yu, H. Duan, J. Hur, K. Sargent, M. Rubinstein, W. T. Freeman, F. Cole, D. Sun, N. Snavely, J. Wu
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
R. Gao, A. Holynski, P. Henzler, A. Brussee, R. Martin Brualla, P. Srinivasan, J. Barron, and B. Poole, “CAT3D: Create anything in 3D with multi-view diffusion models,”
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
T. Li, G. Zheng, R. Jiang, T. Wu, Y. Lu, Y. Lin, X. Li
2025
Closest in time.
2025
Closest in time.