Fetching the paper…
Reading the bibliography…
We present L4GM, the first 4D Large Reconstruction Model that produces animated objects from a single-view video input -- in a single feed-forward pass that takes only a second.
Activitynet: A large-scale video benchmark for human activity understanding
B. G. Fabian Caba Heilbron, Victor Escorcia and J. C. Niebles · 2015
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
O. Ronneberger, P. Fischer, and T. Brox · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Blender - a 3D modelling and rendering package
B. O. Community · 2018
Earlier work this paper cites.
Towards accurate generative models of video: A new metric & challenges
T. Unterthiner, S. Van Steenkiste, K. Kurach, R. Marinier, M. Michalski, and S. Gelly · 2018
Earlier work this paper cites.
The unreasonable effectiveness of deep features as a perceptual metric
R. Zhang, P. Isola, A. A. Efros, E. Shechtman, and O. Wang · 2018
Earlier work this paper cites.
Neural volumes: Learning dynamic renderable volumes from images
S. Lombardi, T. Simon, J. Saragih, G. Schwartz, A. Lehrmann, and Y. Sheikh · 2019
Earlier work this paper cites.
Dynamic view synthesis from dynamic monocular video
C. Gao, A. Saraf, J. Kopf, and J.-B. Huang · 2021
Earlier work this paper cites.
D-nerf: Neural radiance fields for dynamic scenes
A. Pumarola, E. Corona, G. Pons-Moll, and F. Moreno-Noguer · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Earlier work this paper cites.
Einops: Clear and reliable tensor manipulations with einstein-like notation
A. Rogozhnikov · 2021
Earlier work this paper cites.
Light field networks: Neural scene representations with single-evaluation rendering
V. Sitzmann, S. Rezchikov, B. Freeman, J. Tenenbaum, and F. Durand · 2021
Earlier work this paper cites.
Ibrnet: Learning multi-view image-based rendering
Q. Wang, Z. Wang, K. Genova, P. P. Srinivasan, H. Zhou, J. T. Barron, R. Martin-Brualla, N. Snavely, and T. Funkhouser · 2021
Earlier work this paper cites.
Space-time neural irradiance fields for free-viewpoint video
W. Xian, J.-B. Huang, J. Kopf, and C. Kim · 2021
Earlier work this paper cites.
Lasr: Learning articulated shape reconstruction from a monocular video
G. Yang, D. Sun, V. Jampani, D. Vlasic, F. Cole, H. Chang, D. Ramanan, W. T. Freeman, and C. Liu · 2021
Earlier work this paper cites.
pixelnerf: Neural radiance fields from one or few images
A. Yu, V. Ye, M. Tancik, and A. Kanazawa · 2021
Earlier work this paper cites.
Ego4d: Around the world in 3,000 hours of egocentric video
K. Grauman, A. Westbury, E. Byrne, Z. Chavis, A. Furnari, R. Girdhar, J. Hamburger, H. Jiang, M. Liu, X. Liu, et al · 2022
Earlier work this paper cites.
High-Resolution Image Synthesis with Latent Diffusion Models
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer · 2022
Earlier work this paper cites.
Is attention all that nerf needs?
P. Wang, X. Chen, T. Chen, S. Venugopalan, Z. Wang, et al · 2022
Earlier work this paper cites.
4d-fy: Text-to-4d generation using hybrid score distillation sampling
S. Bahmani, I. Skorokhodov, V. Rong, G. Wetzstein, L. Guibas, P. Wonka, S. Tulyakov, J. J. Park, A. Tagliasacchi, and D. B. Lindell · 2023
Earlier work this paper cites.
Flowibr: Leveraging pre-training for efficient neural image-based rendering of dynamic scenes
M. Büsching, J. Bengtson, D. Nilsson, and M. Björkman · 2023
Cited alongside, same era.
Hexplane: A fast representation for dynamic scenes
A. Cao and J. Johnson · 2023
Cited alongside, same era.
pixelsplat: 3d gaussian splats from image pairs for scalable generalizable 3d reconstruction
D. Charatan, S. Li, A. Tagliasacchi, and V. Sitzmann · 2023
Cited alongside, same era.
Y. Cheng, L. Li, Y. Xu, X. Li, Z. Yang, W. Wang, and Y. Yang · 2023
Cited alongside, same era.
Objaverse: A universe of annotated 3d objects
M. Deitke, D. Schwenk, J. Salvador, L. Weihs, O. Michel, E. VanderBilt, L. Schmidt, K. Ehsani, A. Kembhavi, and A. Farhadi · 2023
Cited alongside, same era.
Mononerf: Learning a generalizable dynamic radiance field from monocular videos
F. Tian, S. Du, and Y. Duan · 2023
Later among the works it cites.
Imagedream: Image-prompt multi-view diffusion for 3d generation
P. Wang and Y. Shi · 2023
Later among the works it cites.
ProlificDreamer: High-Fidelity and Diverse Text-to-3D Generation with Variational Score Distillation
Z. Wang, C. Lu, Y. Wang, F. Bao, C. Li, H. Su, and J. Zhu · 2023
Later among the works it cites.
4dgen: Grounded 4d content generation with spatial-temporal consistency
Y. Yin, D. Xu, Z. Wang, Y. Zhao, and Y. Wei · 2023
Later among the works it cites.
Text-to-3D with Classifier Score Distillation
X. Yu, Y.-C. Guo, Y. Li, D. Liang, S.-H. Zhang, and X. Qi · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
K-planes: Explicit radiance fields in space, time, and appearance
S. Fridovich-Keil, G. Meanti, F. R. Warburg, B. Recht, and A. Kanazawa · 2023
Cited alongside, same era.
Emu video: Factorizing text-to-video generation by explicit image conditioning
R. Girdhar, M. Singh, A. Brown, Q. Duval, S. Azadi, S. S. Rambhatla, A. Shah, X. Yin, D. Parikh, and I. Misra · 2023
Cited alongside, same era.
Openlrm: Open-source large reconstruction models
Z. He and T. Wang · 2023
Cited alongside, same era.
Lrm: Large reconstruction model for single image to 3d
Y. Hong, K. Zhang, J. Gu, S. Bi, Y. Zhou, D. Liu, F. Liu, K. Sunkavalli, T. Bui, and H. Tan · 2023
Cited alongside, same era.
Y. Jiang, L. Zhang, J. Gao, W. Hu, and Y. Yao · 2023
Cited alongside, same era.
3d gaussian splatting for real-time radiance field rendering
B. Kerbl, G. Kopanas, T. Leimkühler, and G. Drettakis · 2023
Cited alongside, same era.
Magic3D: High-Resolution Text-to-3D Content Creation
C.-H. Lin, J. Gao, L. Tang, T. Takikawa, X. Zeng, X. Huang, K. Kreis, S. Fidler, M.-Y. Liu, and T.-Y. Lin · 2023
Cited alongside, same era.
Later among the works it cites.
Animate124: Animating one image to 4d dynamic scene
Y. Zhao, Z. Yan, E. Xie, L. Hong, Z. Li, and G. H. Lee · 2023
Later among the works it cites.
A unified approach for text-and image-guided 4d scene generation
Y. Zheng, X. Li, K. Nagano, S. Liu, K. Kreis, O. Hilliges, and S. D. Mello · 2023
Later among the works it cites.
Mvsplat: Efficient 3d gaussian splatting from sparse multi-view images
Y. Chen, H. Xu, C. Zheng, B. Zhuang, M. Pollefeys, A. Geiger, T.-J. Cham, and J. Cai · 2024
Closest in time.
Veo: our most capable generative video model
G. Deepmind · 2024
Closest in time.
Gaussianflow: Splatting gaussian dynamics for 4d content creation
Q. Gao, Q. Xu, Z. Cao, B. Mildenhall, W. Ma, L. Chen, D. Tang, and U. Neumann · 2024
Closest in time.
Generalizable neural human renderer
M. Masuda, J. Park, S. Iwase, R. Khirodkar, and K. Kitani · 2024
Closest in time.
Im-3d: Iterative multiview diffusion and reconstruction for high-quality 3d generation
L. Melas-Kyriazi, I. Laina, C. Rupprecht, N. Neverova, A. Vedaldi, O. Gafni, and F. Kokkinos · 2024
Closest in time.
Fast dynamic 3d object generation from a single-view video, 2024
Z. Pan, Z. Yang, X. Zhu, and L. Zhang · 2024
Closest in time.
Creating video from text
B. Peebles, T. Brooks, C. Brooks, C. Ng, D. Schnurr, E. Luhman, J. Taylor, L. Jing, N. Summers, R. Wang, and et al · 2024
Closest in time.
Gamba: Marry gaussian splatting with mamba for single view 3d reconstruction
Q. Shen, X. Yi, Z. Wu, P. Zhou, H. Zhang, S. Yan, and X. Wang · 2024
Closest in time.
Lgm: Large multi-view gaussian model for high-resolution 3d content creation
J. Tang, Z. Chen, X. Chen, T. Wang, G. Zeng, and Z. Liu · 2024
Closest in time.
Grm: Large gaussian reconstruction model for efficient 3d reconstruction and generation
Y. Xu, Z. Shi, W. Yifan, H. Chen, C. Yang, S. Peng, Y. Shen, and G. Wetzstein · 2024
Closest in time.
Diffusion 2 : Dynamic 3d content generation via score composition of orthogonal diffusion models
Z. Yang, Z. Pan, C. Gu, and L. Zhang · 2024
Closest in time.
Stag4d: Spatial-temporal anchored generative 4d gaussians
Y. Zeng, Y. Jiang, S. Zhu, Y. Lu, Y. Lin, H. Zhu, W. Hu, X. Cao, and Y. Yao · 2024
Closest in time.
Gs-lrm: Large reconstruction model for 3d gaussian splatting
K. Zhang, S. Bi, H. Tan, Y. Xiangli, N. Zhao, K. Sunkavalli, and Z. Xu · 2024
Closest in time.
Pseudo-generalized dynamic view synthesis from a video, 2024
X. Zhao, A. Colburn, F. Ma, M. A. Bautista, J. M. Susskind, and A. G. Schwing · 2024
Closest in time.