Fetching the paper…
Reading the bibliography…
Autoregressive transformers have revolutionized generative models in language processing and shown substantial promise in image and video generation.
Real-time, continuous level of detail rendering of height fields
Lindstrom, P.; Koller, D.; Ribarsky, W.; Hodges, L. F.; Faust, N.; and Turner, G. A. 1996 · 1996
Earlier work this paper cites.
Language Models are Few-Shot Learners
Brown, T. B.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; Agarwal, S.; Herbert-Voss, A.; Krueger, G.; Henighan, T.; Child, R.; Ramesh, A.; Ziegler, D. M.; Wu, J.; Winter, C.; Hesse, C.; Chen, M.; Sigler, E.; teusz Litwin, M.; Gray, S.; Chess, B.; Clark, J.; Berner, C.; McCandlish, S.; Radford, A.; Sutskever, I.; and Amodei, D. 2020 · 2005
Earlier work this paper cites.
Neural discrete representation learning
Van Den Oord, A.; Vinyals, O.; et al. 2017 · 2017
Earlier work this paper cites.
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, L.; and Polosukhin, I. 2017 · 2017
Earlier work this paper cites.
Graph Neural Networks: A Review of Methods and Applications
Zhou, J.; Cui, G.; Zhang, Z.; Yang, C.; Liu, Z.; and Sun, M. 2018 · 2018
Earlier work this paper cites.
SurfaceNet+: An End-to-end 3D Neural Network for Very Sparse Multi-view Stereopsis
Ji, M.; Zhang, J.; Dai, Q.; and Fang, L. 2020 · 2020
Earlier work this paper cites.
Taming transformers for high-resolution image synthesis
Esser, P.; Rombach, R.; and Ommer, B. 2021 · 2021
Earlier work this paper cites.
Zero-shot text-to-image generation
Ramesh, A.; Pavlov, M.; Goh, G.; Gray, S.; Voss, C.; Radford, A.; Chen, M.; and Sutskever, I. 2021 · 2021
Earlier work this paper cites.
SurRF: Unsupervised multi-view stereopsis by learning surface radiance field
Zhang, J.; Ji, M.; Wang, G.; Xue, Z.; Wang, S.; and Fang, L. 2021 · 2021
Earlier work this paper cites.
Autoregressive image generation using residual quantization
Lee, D.; Kim, C.; Kim, S.; Cho, M.; and Han, W.-S. 2022 · 2022
Earlier work this paper cites.
Dreamfusion: Text-to-3d using 2d diffusion
Poole, B.; Jain, A.; Barron, J. T.; and Mildenhall, B. 2022 · 2022
Earlier work this paper cites.
High-resolution image synthesis with latent diffusion models
Rombach, R.; Blattmann, A.; Lorenz, D.; Esser, P.; and Ommer, B. 2022 · 2022
Cited alongside, same era.
Scaling autoregressive models for content-rich text-to-image generation
Yu, J.; Xu, Y.; Koh, J. Y.; Luong, T.; Baid, G.; Wang, Z.; Vasudevan, V.; Ku, A.; Yang, Y.; Ayan, B. K.; et al. 2022 · 2022
Cited alongside, same era.
SDFusion: Multimodal 3D Shape Completion, Reconstruction, and Generation
Cheng, Y.-C.; Lee, H.-Y.; Tuyakov, S.; Schwing, A.; and Gui, L. 2023 · 2023
Cited alongside, same era.
Objaverse: A universe of annotated 3d objects
Deitke, M.; Schwenk, D.; Salvador, J.; Weihs, L.; Michel, O.; VanderBilt, E.; Schmidt, L.; Ehsani, K.; Kembhavi, A.; and Farhadi, A. 2023 · 2023
Cited alongside, same era.
Lrm: Large reconstruction model for single image to 3d
Hong, Y.; Zhang, K.; Gu, J.; Bi, S.; Zhou, Y.; Liu, D.; Liu, F.; Sunkavalli, K.; Bui, T.; and Tan, H. 2023 · 2023
Language Model Beats Diffusion – Tokenizer is Key to Visual Generation
Yu, L.; Lezama, J.; Gundavarapu, N. B.; Versari, L.; Sohn, K.; Minnen, D. C.; Cheng, Y.; Gupta, A.; Gu, X.; Hauptmann, A. G.; Gong, B.; Yang, M.-H.; Essa, I.; Ross, D. A.; and Jiang, L. 2023 · 2023
Later among the works it cites.
LAM3D: Large Image-Point-Cloud Alignment Model for 3D Reconstruction from Single Image
Cui, R.; Song, X.; Sun, W.; Wang, S.; Liu, W.; Chen, S.; Shang, T.; Li, Y.; Barnes, N.; Li, H.; et al. 2024 · 2024
Closest in time.
Objaverse-xl: A universe of 10m+ 3d objects
Deitke, M.; Liu, R.; Wallingford, M.; Ngo, H.; Michel, O.; Kusupati, A.; Fan, A.; Laforte, C.; Voleti, V.; Gadre, S. Y.; et al. 2024 · 2024
Closest in time.
CraftsMan: High-fidelity Mesh Generation with 3D Native Generation and Interactive Geometry Refiner
Li, W.; Liu, J.; Chen, R.; Liang, Y.; Chen, X.; Tan, P.; and Long, X. 2024 · 2024
Closest in time.
Meshgpt: Generating triangle meshes with decoder-only transformers
Siddiqui, Y.; Alliegro, A.; Artemov, A.; Tommasi, T.; Sirigatti, D.; Rosov, V.; Dai, A.; and Nießner, M. 2024 · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Shap-e: Generating conditional 3d implicit functions
Jun, H.; and Nichol, A. 2023 · 2023
Cited alongside, same era.
3D Gaussian Splatting for Real-Time Radiance Field Rendering
Kerbl, B.; Kopanas, G.; Leimkühler, T.; and Drettakis, G. 2023 · 2023
Cited alongside, same era.
Instant3d: Fast text-to-3d with sparse-view generation and large reconstruction model
Li, J.; Tan, H.; Zhang, K.; Xu, Z.; Luan, F.; Xu, Y.; Hong, Y.; Sunkavalli, K.; Shakhnarovich, G.; and Bi, S. 2023 · 2023
Cited alongside, same era.
DINOv2: Learning Robust Visual Features without Supervision
Oquab, M.; Darcet, T.; Moutakanni, T.; Vo, H. V.; Szafraniec, M.; Khalidov, V.; Fernandez, P.; Haziza, D.; Massa, F.; El-Nouby, A.; Howes, R.; Huang, P.-Y.; Xu, H.; Sharma, V.; Li, S.-W.; Galuba, W.; Rabbat, M.; Assran, M.; Ballas, N.; Synnaeve, G.; Misra, I.; Jegou, H.; Mairal, J.; Labatut, P.; Joulin, A.; and Bojanowski, P. 2023 · 2023
Cited alongside, same era.
Flexible Isosurface Extraction for Gradient-Based Mesh Optimization
Shen, T.; Munkberg, J.; Hasselgren, J.; Yin, K.; Wang, Z.; Chen, W.; Gojcic, Z.; Fidler, S.; Sharp, N.; and Gao, J. 2023 · 2023
Cited alongside, same era.
PET-NeuS: Positional Encoding Triplanes for Neural Surfaces
Wang, Y.; Skorokhodov, I.; and Wonka, P. 2023 · 2023
Cited alongside, same era.
Visual Instruction Tuning
Liu, H.; Li, C.; Wu, Q.; and Lee, Y. J. 2023a
Cited in the paper.
Closest in time.
Lgm: Large multi-view gaussian model for high-resolution 3d content creation
Tang, J.; Chen, Z.; Chen, X.; Wang, T.; Zeng, G.; and Liu, Z. 2024 · 2024
Closest in time.
Triposr: Fast 3d object reconstruction from a single image
Tochilkin, D.; Pankratz, D.; Liu, Z.; Huang, Z.; Letts, A.; Li, Y.; Liang, D.; Laforte, C.; Jampani, V.; and Cao, Y.-P. 2024 · 2024
Closest in time.
Crm: Single image to 3d textured mesh with convolutional reconstruction model
Wang, Z.; Wang, Y.; Chen, Y.; Xiang, C.; Chen, S.; Yu, D.; Li, C.; Su, H.; and Zhu, J. 2024 · 2024
Closest in time.
Meshlrm: Large reconstruction model for high-quality mesh
Wei, X.; Zhang, K.; Bi, S.; Tan, H.; Luan, F.; Deschaintre, V.; Sunkavalli, K.; Su, H.; and Xu, Z. 2024 · 2024
Closest in time.
CLAY: A Controllable Large-scale Generative Model for Creating High-quality 3D Assets
Zhang, L.; Wang, Z.; Zhang, Q.; Qiu, Q.; Pang, A.; Jiang, H.; Yang, W.; Xu, L.; and Yu, J. 2024 · 2024
Closest in time.
Michelangelo: Conditional 3d shape generation based on shape-image-text aligned latent representation
Zhao, Z.; Liu, W.; Chen, X.; Zeng, X.; Wang, R.; Cheng, P.; Fu, B.; Chen, T.; Yu, G.; and Gao, S. 2024 · 2024
Closest in time.