Fetching the paper…
Reading the bibliography…
Recent advances in camera-controllable video generation have been constrained by the reliance on static-scene datasets with relative-scale camera annotations, such as RealEstate10K.
Structure-from-motion revisited
J. L. Schönberger and J.-M. Frahm · 2016
Earlier work this paper cites.
Stereo magnification: Learning view synthesis using multiplane images
T. Zhou, R. Tucker, J. Flynn, G. Fyffe, and N. Snavely · 2018
Earlier work this paper cites.
Raft: Recurrent all-pairs field transforms for optical flow
Z. Teed and J. Deng · 2020
Earlier work this paper cites.
Improving image generation with better captions
J. Betker, G. Goh, L. Jing, T. Brooks, J. Wang, L. Li, L. Ouyang, J. Zhuang, J. Lee, Y. Guo, et al · 2023
Earlier work this paper cites.
Metric3d v2: A versatile monocular geometric foundation model for zero-shot metric depth and surface normal estimation
M. Hu, W. Yin, C. Zhang, Z. Cai, X. Long, H. Chen, K. Wang, G. Yu, C. Shen, and S. Shen · 2024
Earlier work this paper cites.
Vbench: Comprehensive benchmark suite for video generative models
Z. Huang, Y. He, J. Yu, F. Zhang, C. Si, Y. Jiang, Y. Zhang, T. Wu, Q. Jin, N. Chanpaisit, et al · 2024
Cited alongside, same era.
Miradata: A large-scale video dataset with long durations and structured captions
X. Ju, Y. Gao, Z. Zhang, Z. Yuan, X. Wang, A. Zeng, Y. Xiong, Q. Xu, and Y. Shan · 2024
Cited alongside, same era.
Cotracker: It is better to track together
N. Karaev, I. Rocco, B. Graham, N. Neverova, A. Vedaldi, and C. Rupprecht · 2024
Cited alongside, same era.
Dl3dv-10k: A large-scale scene dataset for deep learning-based 3d vision
L. Ling, Y. Sheng, Z. Tu, W. Zhao, C. Xin, K. Wan, L. Yu, Q. Guo, Z. Yu, Y. Lu, et al · 2024
Cited alongside, same era.
Sam 2: Segment anything in images and videos
N. Ravi, V. Gabeur, Y.-T. Hu, R. Hu, C. Ryali, T. Ma, H. Khedr, R. Rädle, C. Rolland, L. Gustafson, et al · 2024
Later among the works it cites.
Q. Wang, Y. Shi, J. Ou, R. Chen, K. Lin, J. Wang, B. Jiang, H. Yang, M. Zheng, X. Tao, et al · 2024
Later among the works it cites.
Cogvideox: Text-to-video diffusion models with an expert transformer
Z. Yang, J. Teng, W. Zheng, M. Ding, S. Huang, J. Xu, Y. Yang, W. Hong, X. Zhang, G. Feng, et al · 2024
Later among the works it cites.
Monst3r: A simple approach for estimating geometry in the presence of motion
J. Zhang, C. Herrmann, J. Hur, V. Jampani, T. Darrell, F. Cole, D. Sun, and M.-H. Yang · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…