Fetching the paper…
Reading the bibliography…
We propose 4DGT, a 4D Gaussian-based Transformer model for dynamic scene reconstruction, trained entirely on real-world monocular posed videos.
Image quality assessment: from error visibility to structural similarity
Z. Wang, A. C. Bovik, H. R. Sheikh, and E. P. Simoncelli · 2004
Earlier work this paper cites.
Cop3d: Context-aware overlay tree for content-based control systems
M. Caporuscio and A. Navarra · 2009
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
A benchmark for the evaluation of rgb-d slam systems
J. Sturm, N. Engelhard, F. Endres, W. Burgard, and D. Cremers · 2012
Earlier work this paper cites.
Dynamicfusion: Reconstruction and tracking of non-rigid scenes in real-time
R. A. Newcombe, D. Fox, and S. M. Seitz · 2015
Earlier work this paper cites.
J. L. Ba, J. R. Kiros, and G. E. Hinton · 2016
Earlier work this paper cites.
Gaussian error linear units (gelus)
D. Hendrycks and K. Gimpel · 2016
Earlier work this paper cites.
Perceptual losses for real-time style transfer and super-resolution
J. Johnson, A. Alahi, and L. Fei-Fei · 2016
Earlier work this paper cites.
Sgdr: Stochastic gradient descent with warm restarts
I. Loshchilov and F. Hutter · 2016
Earlier work this paper cites.
Structure-from-motion revisited
J. L. Schonberger and J.-M. Frahm · 2016
Earlier work this paper cites.
Decoupled weight decay regularization
I. Loshchilov and F. Hutter · 2017
Earlier work this paper cites.
Killingfusion: Non-rigid 3d reconstruction without correspondences
M. Slavcheva, M. Baust, D. Cremers, and S. Ilic · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Scaling egocentric vision: The epic-kitchens dataset
D. Damen, H. Doughty, G. M. Farinella, S. Fidler, A. Furnari, E. Kazakos, D. Moltisanti, J. Munro, T. Perrett, W. Price, et al · 2018
Earlier work this paper cites.
The unreasonable effectiveness of deep features as a perceptual metric
R. Zhang, P. Isola, A. A. Efros, E. Shechtman, and O. Wang · 2018
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, A. Desmaison, A. Kopf, E. Yang, Z. DeVito, M. Raison, A. Tejani, S. Chilamkurthy, B. Steiner, L. Fang, J. Bai, and S. Chintala · 2019
Earlier work this paper cites.
Deepdeform: Learning non-rigid rgb-d reconstruction with semi-supervised data
A. Bozic, M. Zollhofer, C. Theobalt, and M. Niessner · 2020
Earlier work this paper cites.
Immersive light field video with a layered mesh representation
M. Broxton, J. Flynn, R. Overbeck, D. Erickson, P. Hedman, M. Duvall, J. Dourgarian, J. Busch, M. Whalen, and P. Debevec · 2020
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al · 2020
Earlier work this paper cites.
Nerf: Representing scenes as neural radiance fields for view synthesis
B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Barron, R. Ramamoorthi, and R. Ng · 2020
Cited alongside, same era.
Accelerating 3d deep learning with pytorch3d
N. Ravi, J. Reizenstein, D. Novotny, T. Gordon, W.-Y. Lo, J. Johnson, and G. Gkioxari · 2020
Cited alongside, same era.
Monocular dynamic view synthesis: A reality check
H. Gao, R. Li, S. Tulsiani, B. Russell, and A. Kanazawa · 2022
Cited alongside, same era.
Neural 3d video synthesis from multi-view video
T. Li, M. Slavcheva, M. Zollhoefer, S. Green, C. Lassner, C. Kim, T. Schmidt, S. Lovegrove, M. Goesele, R. Newcombe, et al · 2022
Cited alongside, same era.
Project aria: A new tool for egocentric multi-modal ai research
J. Engel, K. Somasundaram, M. Goesele, A. Sun, A. Gamino, A. Turner, A. Talattof, A. Yuan, B. Souti, B. Meredith, et al · 2023
Cited alongside, same era.
2d gaussian splatting for geometrically accurate radiance fields
B. Huang, Z. Yu, A. Chen, A. Geiger, and S. Gao · 2024
Later among the works it cites.
MoSca: Dynamic gaussian fusion from casual videos via 4D motion scaffolds
J. Lei, Y. Weng, A. Harley, L. Guibas, and K. Daniilidis · 2024
Later among the works it cites.
Megasam: Accurate, fast, and robust structure and motion from casual dynamic videos
Z. Li, R. Tucker, F. Cole, Q. Wang, L. Jin, V. Ye, A. Kanazawa, A. Holynski, and N. Snavely · 2024
Later among the works it cites.
Feed-forward bullet-time reconstruction of dynamic scenes from monocular videos
H. Liang, J. Ren, A. Mirzaei, A. Torralba, Z. Liu, I. Gilitschenski, S. Fidler, C. Oztireli, H. Ling, Z. Gojcic, et al · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
K-planes: Explicit radiance fields in space, time, and appearance
S. Fridovich-Keil, G. Meanti, F. R. Warburg, B. Recht, and A. Kanazawa · 2023
Cited alongside, same era.
Lrm: Large reconstruction model for single image to 3d
Y. Hong, K. Zhang, J. Gu, S. Bi, Y. Zhou, D. Liu, F. Liu, K. Sunkavalli, T. Bui, and H. Tan · 2023
Cited alongside, same era.
Consistent4d: Consistent 360
Y. Jiang, L. Zhang, J. Gao, W. Hu, and Y. Yao · 2023
Cited alongside, same era.
3d gaussian splatting for real-time radiance field rendering
B. Kerbl, G. Kopanas, T. Leimkühler, and G. Drettakis · 2023
Cited alongside, same era.
Dinov2: Learning robust visual features without supervision
M. Oquab, T. Darcet, T. Moutakanni, H. Vo, M. Szafraniec, V. Khalidov, P. Fernandez, D. Haziza, F. Massa, A. El-Nouby, et al · 2023
Cited alongside, same era.
Aria digital twin: A new benchmark dataset for egocentric 3d machine perception
X. Pan, N. Charron, Y. Yang, S. Peters, T. Whelan, C. Kong, O. Parkhi, R. Newcombe, and Y. C. Ren · 2023
Cited alongside, same era.
Common pets in 3d: Dynamic new-view synthesis of real-life deformable categories
S. Sinha, R. Shapovalov, J. Reizenstein, I. Rocco, N. Neverova, A. Vedaldi, and D. Novotny · 2023
Cited alongside, same era.
Z. Lv, N. Charron, P. Moulon, A. Gamino, C. Peng, C. Sweeney, E. Miller, H. Tang, J. Meissner, J. Dong, et al · 2024
Later among the works it cites.
Nymeria: A massive collection of multimodal egocentric daily motion in the wild
L. Ma, Y. Ye, F. Hong, V. Guzov, Y. Jiang, R. Postyeni, L. Pesqueira, A. Gamino, V. Baiyya, H. J. Kim, et al · 2024
Later among the works it cites.
Efficient4d: Fast dynamic 3d object generation from a single-view video
Z. Pan, Z. Yang, X. Zhu, and L. Zhang · 2024
Later among the works it cites.
Unidepth: Universal monocular metric depth estimation
L. Piccinelli, Y.-H. Yang, C. Sakaridis, M. Segu, S. Li, L. Van Gool, and F. Yu · 2024
Later among the works it cites.
L4gm: Large 4d gaussian reconstruction model
J. Ren, C. Xie, A. Mirzaei, K. Kreis, Z. Liu, A. Torralba, S. Fidler, S. W. Kim, H. Ling, et al · 2024
Later among the works it cites.
Flashattention-3: Fast and accurate attention with asynchrony and low-precision
J. Shah, G. Bikshandi, Y. Zhang, V. Thakkar, P. Ramani, and T. Dao · 2024
Later among the works it cites.
Shape of motion: 4d reconstruction from a single video
Q. Wang, V. Ye, H. Gao, J. Austin, Z. Li, and A. Kanazawa · 2024
Later among the works it cites.
Cat4d: Create anything in 4d with multi-view video diffusion models
R. Wu, R. Gao, B. Poole, A. Trevithick, C. Zheng, J. T. Barron, and A. Holynski · 2024
Later among the works it cites.
Stablenormal: Reducing diffusion variance for stable and sharp normal
C. Ye, L. Qiu, X. Gu, Q. Zuo, Y. Wu, Z. Dong, L. Bo, Y. Xiu, and X. Han · 2024
Later among the works it cites.
Long-lrm: Long-sequence large reconstruction model for wide-coverage gaussian splats
C. Ziwen, H. Tan, K. Zhang, S. Bi, F. Luan, Y. Hong, L. Fuxin, and Z. Xu · 2024
Later among the works it cites.
Z. Li, D. Wang, K. Chen, Z. Lv, T. Nguyen-Phuoc, M. Lee, J.-B. Huang, L. Xiao, C. Zhang, Y. Zhu, et al · 2025
Closest in time.
MASt3R-SLAM: Real-time dense SLAM with 3D reconstruction priors
R. Murai, E. Dexheimer, and A. J. Davison · 2025
Closest in time.
Unidepthv2: Universal monocular metric depth estimation made simpler
L. Piccinelli, C. Sakaridis, Y.-H. Yang, M. Segu, S. Li, W. Abbeloos, and L. Van Gool · 2025
Closest in time.
Storm: Spatio-temporal reconstruction model for large-scale outdoor scenes
J. Yang, J. Huang, Y. Chen, Y. Wang, B. Li, Y. You, M. Igl, A. Sharma, P. Karkus, D. Xu, B. Ivanovic, Y. Wang, and M. Pavone · 2025
Closest in time.
gsplat: An open-source library for gaussian splatting
V. Ye, R. Li, J. Kerr, M. Turkulainen, B. Yi, Z. Pan, O. Seiskari, J. Ye, J. Hu, M. Tancik, et al · 2025
Closest in time.