Fetching the paper…
Reading the bibliography…
Recent advancements in diffusion techniques have propelled image and video generation to unprecedented levels of quality, significantly accelerating the deployment and application of generative AI.
Marching cubes: A high resolution 3d surface construction algorithm
Lorensen, W. E. and Cline, H. E · 1987
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Kinectfusion: Real-time dense surface mapping and tracking
Newcombe, R. A., Izadi, S., Hilliges, O., Molyneaux, D., Kim, D., Davison, A. J., Kohli, P., Shotton, J., Hodges, S., and Fitzgibbon, A. W · 2011
Earlier work this paper cites.
Visualizing and understanding convolutional networks
Zeiler, M. D. and Fergus, R · 2014
Earlier work this paper cites.
Shapenet: An information-rich 3d model repository
Chang, A. X., Funkhouser, T., Guibas, L., Hanrahan, P., Huang, Q., Li, Z., Savarese, S., Savva, M., Song, S., Su, H., et al · 2015
Earlier work this paper cites.
Learning a predictable and generative vector representation for objects
Girdhar, R., Fouhey, D. F., Rodriguez, M., and Gupta, A · 2016
Earlier work this paper cites.
A point set generation network for 3d object reconstruction from a single image
Fan, H., Su, H., and Guibas, L. J · 2017
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
Heusel, M., Ramsauer, H., Unterthiner, T., Nessler, B., and Hochreiter, S · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Earlier work this paper cites.
Marrnet: 3d shape reconstruction via 2.5d sketches
Wu, J., Wang, Y., Xue, T., Sun, X., Freeman, B., and Tenenbaum, J · 2017
Earlier work this paper cites.
Robust watertight manifold surface generation method for shapenet models
Huang, J., Su, H., and Guibas, L. J · 2018
Earlier work this paper cites.
Pixel2mesh: Generating 3d mesh models from single RGB images
Wang, N., Zhang, Y., Li, Z., Fu, Y., Liu, W., and Jiang, Y · 2018
Earlier work this paper cites.
Occupancy networks: Learning 3d reconstruction in function space
Mescheder, L. M., Oechsle, M., Niemeyer, M., Nowozin, S., and Geiger, A · 2019
Earlier work this paper cites.
DISN: deep implicit surface network for high-quality single-view 3d reconstruction
Xu, Q., Wang, W., Ceylan, D., Mech, R., and Neumann, U · 2019
Earlier work this paper cites.
Root mean square layer normalization
Zhang, B. and Sennrich, R · 2019
Earlier work this paper cites.
Denoising diffusion probabilistic models
Ho, J., Jain, A., and Abbeel, P · 2020
Earlier work this paper cites.
Manifoldplus: A robust and scalable watertight manifold surface generation method for triangle soups
Huang, J., Zhou, Y., and Guibas, L. J · 2020
Earlier work this paper cites.
Nerf: Representing scenes as neural radiance fields for view synthesis
Mildenhall, B., Srinivasan, P. P., Tancik, M., Barron, J. T., Ramamoorthi, R., and Ng, R · 2020
Earlier work this paper cites.
PQ-NET: A generative part seq2seq network for 3d shapes
Wu, R., Zhuang, Y., Xu, K., Zhang, H., and Chen, B · 2020
Earlier work this paper cites.
Frozen in time: A joint video and image encoder for end-to-end retrieval
Bain, M., Nagrani, A., Varol, G., and Zisserman, A · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Earlier work this paper cites.
Zero-shot text-to-image generation
Ramesh, A., Pavlov, M., Goh, G., Gray, S., Voss, C., Radford, A., Chen, M., and Sutskever, I · 2021
Earlier work this paper cites.
Scaling vision with sparse mixture of experts
Riquelme, C., Puigcerver, J., Mustafa, B., Neumann, M., Jenatton, R., Susano Pinto, A., Keysers, D., and Houlsby, N · 2021
Earlier work this paper cites.
pixelnerf: Neural radiance fields from one or few images
Yu, A., Ye, V., Tancik, M., and Kanazawa, A · 2021
Earlier work this paper cites.
3d shape generation and completion through point-voxel diffusion
Zhou, L., Du, Y., and Wu, J · 2021
Earlier work this paper cites.
Neural wavelet-domain diffusion for 3d shape generation
Hui, K., Li, R., Hu, J., and Fu, C · 2022
Earlier work this paper cites.
Point-e: A system for generating 3d point clouds from complex prompts
Nichol, A., Jun, H., Dhariwal, P., Mishkin, P., and Chen, M · 2022
Cited alongside, same era.
Dreamfusion: Text-to-3d using 2d diffusion
Poole, B., Jain, A., Barron, J. T., and Mildenhall, B · 2022
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B · 2022
Cited alongside, same era.
Laion-5b: An open large-scale dataset for training next generation image-text models
Schuhmann, C., Beaumont, R., Vencu, R., Gordon, C., Wightman, R., Cherti, M., Coombes, T., Katta, A., Mullis, C., Wortsman, M., et al · 2022
Cited alongside, same era.
Diffusers: State-of-the-art diffusion models
von Platen, P., Patil, S., Lozhkov, A., Cuenca, P., Lambert, N., Rasul, K., Davaadorj, M., Nair, D., Paul, S., Berman, W., Xu, Y., Liu, S., and Wolf, T · 2022
Cited alongside, same era.
3d neural field generation using triplane diffusion
Shue, J. R., Chan, E. R., Po, R., Ankner, Z., Wu, J., and Wetzstein, G · 2023
Later among the works it cites.
Make-a-video: Text-to-video generation without text-video data
Singer, U., Polyak, A., Hayes, T., Yin, X., An, J., Zhang, S., Hu, Q., Yang, H., Ashual, O., Gafni, O., Parikh, D., Gupta, S., and Taigman, Y · 2023
Later among the works it cites.
Make-it-3d: High-fidelity 3d creation from A single image with diffusion prior
Tang, J., Wang, T., Zhang, B., Zhang, T., Yi, R., Ma, L., and Chen, D · 2023
Later among the works it cites.
Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation
Wu, J. Z., Ge, Y., Wang, X., Lei, S. W., Gu, Y., Shi, Y., Hsu, W., Shan, Y., Qie, X., and Shou, M. Z · 2023
Later among the works it cites.
Dmv3d: Denoising multi-view diffusion using 3d large reconstruction model
Xu, Y., Tan, H., Luan, F., Bi, S., Wang, P., Li, J., Shi, Z., Sunkavalli, K., Wetzstein, G., Xu, Z., et al · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dual octree graph networks for learning adaptive volumetric shape representations
Wang, P., Liu, Y., and Tong, X · 2022
Cited alongside, same era.
Multi-view mesh reconstruction with neural deferred shading
Worchel, M., Diaz, R., Hu, W., Schreer, O., Feldmann, I., and Eisert, P · 2022
Cited alongside, same era.
LION: latent point diffusion models for 3d shape generation
Zeng, X., Vahdat, A., Williams, F., Gojcic, Z., Litany, O., Fidler, S., and Kreis, K · 2022
Cited alongside, same era.
All are worth words: A vit backbone for diffusion models
Bao, F., Nie, S., Xue, K., Cao, Y., Li, C., Su, H., and Zhu, J · 2023
Cited alongside, same era.
Generative novel view synthesis with 3d-aware diffusion models
Chan, E. R., Nagano, K., Chan, M. A., Bergman, A. W., Park, J. J., Levy, A., Aittala, M., Mello, S. D., Karras, T., and Wetzstein, G · 2023
Cited alongside, same era.
Fantasia3d: Disentangling geometry and appearance for high-quality text-to-3d content creation
Chen, R., Chen, Y., Jiao, N., and Jia, K · 2023
Cited alongside, same era.
Sdfusion: Multimodal 3d shape completion, reconstruction, and generation
Cheng, Y., Lee, H., Tulyakov, S., Schwing, A. G., and Gui, L · 2023
Cited alongside, same era.
3dshape2vecset: A 3d shape representation for neural fields and generative diffusion models
Zhang, B., Tang, J., Niessner, M., and Wonka, P · 2023
Later among the works it cites.
Locally attentional SDF diffusion for controllable 3d shape generation
Zheng, X., Pan, H., Wang, P., Tong, X., Liu, Y., and Shum, H · 2023
Later among the works it cites.
Meta 3d texturegen: fast and consistent texture generation for 3d objects
Bensadoun, R., Kleiman, Y., Azuri, I., Harosh, O., Vedaldi, A., Neverova, N., and Gafni, O · 2024
Later among the works it cites.
Video generation models as world simulators
Brooks, T., Peebles, B., Holmes, C., DePue, W., Guo, Y., Jing, L., Schnurr, D., Taylor, J., Luhman, T., Luhman, E., Ng, C., Wang, R., and Ramesh, A · 2024
Later among the works it cites.
Objaverse-xl: A universe of 10m+ 3d objects
Deitke, M., Liu, R., Wallingford, M., Ngo, H., Michel, O., Kusupati, A., Fan, A., Laforte, C., Voleti, V., Gadre, S. Y., et al · 2024
Later among the works it cites.
Scaling rectified flow transformers for high-resolution image synthesis
Esser, P., Kulal, S., Blattmann, A., Entezari, R., Müller, J., Saini, H., Levi, Y., Lorenz, D., Sauer, A., Boesel, F., et al · 2024
Later among the works it cites.
Scaling diffusion transformers to 16 billion parameters
Fei, Z., Fan, M., Yu, C., Li, D., and Huang, J · 2024
Later among the works it cites.
3dtopia: Large text-to-3d generation model with hybrid diffusion priors
Hong, F., Tang, J., Cao, Z., Shi, M., Wu, T., Chen, Z., Wang, T., Pan, L., Lin, D., and Liu, Z · 2024
Later among the works it cites.
Ln3diff: Scalable latent neural fields diffusion for speedy 3d generation
Lan, Y., Hong, F., Yang, S., Zhou, S., Meng, X., Dai, B., Pan, X., and Loy, C. C · 2024
Later among the works it cites.
Luciddreamer: Towards high-fidelity text-to-3d generation via interval score matching
Liang, Y., Yang, X., Lin, J., Li, H., Xu, X., and Chen, Y · 2024
Later among the works it cites.
Wonder3d: Single image to 3d using cross-domain diffusion
Long, X., Guo, Y.-C., Lin, C., Liu, Y., Dou, Z., Liu, L., Ma, Y., Zhang, S.-H., Habermann, M., Theobalt, C., et al · 2024
Later among the works it cites.
Mvdream: Multi-view diffusion for 3d generation
Shi, Y., Wang, P., Ye, J., Mai, L., Li, K., and Yang, X · 2024
Later among the works it cites.
Dreamgaussian: Generative gaussian splatting for efficient 3d content creation
Tang, J., Ren, J., Zhou, H., Liu, Z., and Zeng, G · 2024
Later among the works it cites.
Movie gen: A cast of media foundation models
team at Meta, T. M. G · 2024
Later among the works it cites.
Triposr: Fast 3d object reconstruction from a single image
Tochilkin, D., Pankratz, D., Liu, Z., Huang, Z., Letts, A., Li, Y., Liang, D., Laforte, C., Jampani, V., and Cao, Y.-P · 2024
Later among the works it cites.
Crm: Single image to 3d textured mesh with convolutional reconstruction model
Wang, Z., Wang, Y., Chen, Y., Xiang, C., Chen, S., Yu, D., Li, C., Su, H., and Zhu, J · 2024
Later among the works it cites.
Meshlrm: Large reconstruction model for high-quality mesh
Wei, X., Zhang, K., Bi, S., Tan, H., Luan, F., Deschaintre, V., Sunkavalli, K., Su, H., and Xu, Z · 2024
Later among the works it cites.
Xu, J., Cheng, W., Gao, Y., Wang, X., Gao, S., and Shan, Y · 2024
Later among the works it cites.
Gaussiandreamer: Fast generation from text to 3d gaussians by bridging 2d and 3d diffusion models
Yi, T., Fang, J., Wang, J., Wu, G., Xie, L., Zhang, X., Liu, W., Tian, Q., and Wang, X · 2024
Later among the works it cites.
Michelangelo: Conditional 3d shape generation based on shape-image-text aligned latent representation
Zhao, Z., Liu, W., Chen, X., Zeng, X., Wang, R., Cheng, P., Fu, B., Chen, T., Yu, G., and Gao, S · 2024
Later among the works it cites.