Fetching the paper…
Reading the bibliography…
Generating high-quality 3D content from text, single images, or sparse view images remains a challenging task with broad applications.
Shapenet: An information-rich 3d model repository
Chang, A. X., Funkhouser, T., Guibas, L., Hanrahan, P., Huang, Q., Li, Z., Savarese, S., Savva, M., Song, S., Su, H., et al · 2015
Earlier work this paper cites.
Arbitrary style transfer in real-time with adaptive instance normalization
Huang, X. and Belongie, S · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2017
Earlier work this paper cites.
The unreasonable effectiveness of deep features as a perceptual metric
Zhang, R., Isola, P., Efros, A. A., Shechtman, E., and Wang, O · 2018
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A · 2020
Earlier work this paper cites.
Emerging properties in self-supervised vision transformers
Caron, M., Touvron, H., Misra, I., Jégou, H., Mairal, J., Bojanowski, P., and Joulin, A · 2021
Earlier work this paper cites.
Gancraft: Unsupervised 3d neural rendering of minecraft worlds
Hao, Z., Mallya, A., Belongie, S., and Liu, M.-Y · 2021
Earlier work this paper cites.
Nerf: Representing scenes as neural radiance fields for view synthesis
Mildenhall, B., Srinivasan, P. P., Tancik, M., Barron, J. T., Ramamoorthi, R., and Ng, R · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I · 2021
Earlier work this paper cites.
pixelnerf: Neural radiance fields from one or few images
Yu, A., Ye, V., Tancik, M., and Kanazawa, A · 2021
Earlier work this paper cites.
Efficient geometry-aware 3d generative adversarial networks
Chan, E. R., Lin, C. Z., Chan, M. A., Nagano, K., Pan, B., De Mello, S., Gallo, O., Guibas, L. J., Tremblay, J., Khamis, S., et al · 2022
Earlier work this paper cites.
Google scanned objects: A high-quality dataset of 3d scanned household items
Downs, L., Francis, A., Koenig, N., Kinman, B., Hickman, R., Reymann, K., McHugh, T. B., and Vanhoucke, V · 2022
Earlier work this paper cites.
Rt-nerf: Real-time on-device neural radiance fields towards immersive ar/vr rendering
Li, C., Li, S., Zhao, Y., Zhu, W., and Lin, Y · 2022
Earlier work this paper cites.
DreamFusion: Text-to-3d using 2d diffusion
Poole, B., Jain, A., Barron, J. T., and Mildenhall, B · 2022
Earlier work this paper cites.
Fantasia3d: Disentangling geometry and appearance for high-quality text-to-3d content creation
Chen, R., Chen, Y., Jiao, N., and Jia, K · 2023
Earlier work this paper cites.
Emu: Enhancing image generation models using photogenic needles in a haystack
Dai, X., Hou, J., Ma, C.-Y., Tsai, S., Wang, J., Wang, R., Zhang, P., Vandenhende, S., Wang, X., Dubey, A., et al · 2023
Earlier work this paper cites.
Objaverse: A universe of annotated 3d objects
Deitke, M., Schwenk, D., Salvador, J., Weihs, L., Michel, O., VanderBilt, E., Schmidt, L., Ehsani, K., Kembhavi, A., and Farhadi, A · 2023
Earlier work this paper cites.
Hyperdiffusion: Generating implicit neural fields with weight-space diffusion
Erkoç, Z., Ma, F., Shan, Q., Nießner, M., and Dai, A · 2023
Earlier work this paper cites.
Emu video: Factorizing text-to-video generation by explicit image conditioning
Girdhar, R., Singh, M., Brown, A., Duval, Q., Azadi, S., Rambhatla, S. S., Shah, A., Yin, X., Parikh, D., and Misra, I · 2023
Cited alongside, same era.
Openlrm: Open-source large reconstruction models
He, Z. and Wang, T · 2023
Cited alongside, same era.
3d gaussian splatting for real-time radiance field rendering
Kerbl, B., Kopanas, G., Leimkühler, T., and Drettakis, G · 2023
Cited alongside, same era.
Vivid-1-to-3: Novel view synthesis with video diffusion models
Kwak, J.-g., Dong, E., Jin, Y., Ko, H., Mahajan, S., and Yi, K. M · 2023
Cited alongside, same era.
Att3d: Amortized text-to-3d object synthesis
Lorraine, J., Xie, K., Zeng, X., Lin, C.-H., Takikawa, T., Sharp, N., Lin, T.-Y., Liu, M.-Y., Fidler, S., and Lucas, J · 2023
Hunyuan3d 1.0: A unified framework for text-to-3d and image-to-3d generation
Hunyuan3D, T · 2024
Closest in time.
Real3d: Scaling up large reconstruction models with real-world images
Jiang, H., Huang, Q., and Pavlakos, G · 2024
Closest in time.
Ln3diff: Scalable latent neural fields diffusion for speedy 3d generation
Lan, Y., Hong, F., Yang, S., Zhou, S., Meng, X., Dai, B., Pan, X., and Loy, C. C · 2024
Closest in time.
Wonder3d: Single image to 3d using cross-domain diffusion
Long, X., Guo, Y.-C., Lin, C., Liu, Y., Dou, Z., Liu, L., Ma, Y., Zhang, S.-H., Habermann, M., Theobalt, C., et al · 2024
Closest in time.
Im-3d: Iterative multiview diffusion and reconstruction for high-quality 3d generation
Melas-Kyriazi, L., Laina, I., Rupprecht, C., Neverova, N., Vedaldi, A., Gafni, O., and Kokkinos, F · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Dinov2: Learning robust visual features without supervision
Oquab, M., Darcet, T., Moutakanni, T., Vo, H., Szafraniec, M., Khalidov, V., Fernandez, P., Haziza, D., Massa, F., El-Nouby, A., et al · 2023
Cited alongside, same era.
Xcube (x3): Large-scale 3d generative modeling using sparse voxel hierarchies
Ren, X., Huang, J., Zeng, X., Museth, K., Fidler, S., and Williams, F · 2023
Cited alongside, same era.
3d-gpt: Procedural 3d modeling with large language models
Sun, C., Han, J., Deng, W., Wang, X., Qin, Z., and Gould, S · 2023
Cited alongside, same era.
Mvdiffusion: Enabling holistic multi-view image generation with correspondence-aware diffusion
Tang, S., Zhang, F., Chen, J., Wang, P., and Yasutaka, F · 2023
Cited alongside, same era.
Imagedream: Image-prompt multi-view diffusion for 3d generation
Wang, P. and Shi, Y · 2023
Cited alongside, same era.
Mvimgnet: A large-scale dataset of multi-view images
Yu, X., Xu, M., Zhang, Y., Liu, H., Ye, C., Wu, Y., Yan, Z., Zhu, C., Xiong, Z., Liang, T., et al · 2023
Cited alongside, same era.
Zou, Z.-X., Yu, Z., Guo, Y.-C., Li, Y., Liang, D., Cao, Y.-P., and Zhang, S.-H · 2023
Cited alongside, same era.
Closest in time.
Robocasa: Large-scale simulation of everyday tasks for generalist robots
Nasiriany, S., Maddukuri, A., Zhang, L., Parikh, A., Lo, A., Joshi, A., Mandlekar, A., and Zhu, Y · 2024
Closest in time.
Richdreamer: A generalizable normal-depth diffusion model for detail richness in text-to-3d
Qiu, L., Chen, G., Gu, X., Zuo, Q., Xu, M., Wu, Y., Yuan, W., Dong, Z., Bo, L., and Han, X · 2024
Closest in time.
Meta 3d assetgen: Text-to-mesh generation with high-quality geometry, texture, and pbr materials
Siddiqui, Y., Monnier, T., Kokkinos, F., Kariya, M., Kleiman, Y., Garreau, E., Gafni, O., Neverova, N., Vedaldi, A., Shapovalov, R., et al · 2024
Closest in time.
Triposr: Fast 3d object reconstruction from a single image
Tochilkin, D., Pankratz, D., Liu, Z., Huang, Z., , Letts, A., Li, Y., Liang, D., Laforte, C., Jampani, V., and Cao, Y.-P · 2024
Closest in time.
Sv3d: Novel multi-view synthesis and 3d generation from a single image using latent video diffusion
Voleti, V., Yao, C.-H., Boss, M., Letts, A., Pankratz, D., Tochilkin, D., Laforte, C., Rombach, R., and Jampani, V · 2024
Closest in time.
Meshlrm: Large reconstruction model for high-quality mesh
Wei, X., Zhang, K., Bi, S., Tan, H., Luan, F., Deschaintre, V., Sunkavalli, K., Su, H., and Xu, Z · 2024
Closest in time.
Ouroboros3d: Image-to-3d generation via 3d-aware recursive diffusion
Wen, H., Huang, Z., Wang, Y., Chen, X., Qiao, Y., and Sheng, L · 2024
Closest in time.
Gs2mesh: Surface reconstruction from gaussian splatting via novel stereo views
Wolf, Y., Bracha, A., and Kimmel, R · 2024
Closest in time.
Harmonyview: Harmonizing consistency and diversity in one-image-to-3d
Woo, S., Park, B., Go, H., Kim, J.-Y., and Kim, C · 2024
Closest in time.
Consistent-1-to-3: Consistent image to 3d view synthesis via geometry-aware diffusion models
Ye, J., Wang, P., Li, K., Shi, Y., and Wang, H · 2024
Closest in time.
Ctrl123: Consistent novel view synthesis via closed-loop transcription
Zhao, H., Dai, X., Wang, J., Tong, S., Zhang, J., Wang, W., Zhang, L., and Ma, Y · 2024
Closest in time.
Videomv: Consistent multi-view generation based on large video generative model
Zuo, Q., Gu, X., Qiu, L., Dong, Y., Zhao, Z., Yuan, W., Peng, R., Zhu, S., Dong, Z., Bo, L., et al · 2024
Closest in time.