Fetching the paper…
Reading the bibliography…
Understanding, navigating, and exploring the 3D physical real world has long been a central challenge in the development of artificial intelligence.
Image quality assessment: from error visibility to structural similarity
Z. Wang, A. Bovik, H. Sheikh, and E. Simoncelli · 2004
Earlier work this paper cites.
Image quality metrics: Psnr vs. ssim
A. Horé and D. Ziou · 2010
Earlier work this paper cites.
The unreasonable effectiveness of deep features as a perceptual metric, 2018
R. Zhang, P. Isola, A. A. Efros, E. Shechtman, and O. Wang · 2018
Earlier work this paper cites.
Towards accurate generative models of video: A new metric and challenges, 2019
T. Unterthiner, S. van Steenkiste, K. Kurach, R. Marinier, M. Michalski, and S. Gelly · 2019
Earlier work this paper cites.
J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al · 2023
Earlier work this paper cites.
Stable video diffusion: Scaling latent video diffusion models to large datasets, 2023
A. Blattmann, T. Dockhorn, S. Kulal, D. Mendelevitch, M. Kilian, D. Lorenz, Y. Levi, Z. English, V. Voleti, A. Letts, V. Jampani, and R. Rombach · 2023
Earlier work this paper cites.
Stable zero123, 2023
StabilityAI · 2023
Earlier work this paper cites.
Diffusion with forward models: Solving stochastic inverse problems without direct supervision
A. Tewari, T. Yin, G. Cazenavette, S. Rezchikov, J. Tenenbaum, F. Durand, B. Freeman, and V. Sitzmann · 2023
Cited alongside, same era.
Flux.1-[dev] panorama lora (v2), 2024
J. Bilcke · 2024
Cited alongside, same era.
Closed-loop visuomotor control with generative expectation for robotic manipulation
Q. Bu, J. Zeng, L. Chen, Y. Yang, G. Zhou, J. Yan, P. Luo, H. Cui, Y. Ma, and H. Li · 2024
Cited alongside, same era.
Genie 2: A large-scale foundation world model, 2024
DeepMind · 2024
Cited alongside, same era.
Evidential active recognition: Intelligent and prudent open-world embodied perception
L. Fan, M. Liang, Y. Li, G. Hua, and Y. Wu · 2024
Cited alongside, same era.
Videopoet: A large language model for zero-shot video generation
T. Lu, T. Shu, A. Yuille, D. Khashabi, and J. Chen · 2024
Closest in time.
Video generation models as world simulators, 2024
OpenAI · 2024
Closest in time.
Triposr: Fast 3d object reconstruction from a single image
D. Tochilkin, D. Pankratz, Z. Liu, Z. Huang, A. Letts, Y. Li, D. Liang, C. Laforte, V. Jampani, and Y.-P. Cao · 2024
Closest in time.
Sv3d: Novel multi-view synthesis and 3d generation from a single image using latent video diffusion
V. Voleti, C.-H. Yao, M. Boss, A. Letts, D. Pankratz, D. Tochilkin, C. Laforte, R. Rombach, and V. Jampani · 2024
Closest in time.
Generating worlds, 2024
WorldLabs · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
D. Kondratyuk, L. Yu, X. Gu, J. Lezama, J. Huang, R. Hornung, H. Adam, H. Akbari, Y. Alon, V. Birodkar, et al · 2024
Cited alongside, same era.
Flux.1 [dev], 2024
B. F. Labs · 2024
Cited alongside, same era.
Video language planning
Y. Du, M. Yang, P. Florence, F. Xia, A. Wahid, B. Ichter, P. Sermanet, T. Yu, P. Abbeel, J. B. Tenenbaum, et al
Cited in the paper.
Learning universal policies via text-guided video generation
Y. Du, S. Yang, B. Dai, H. Dai, O. Nachum, J. Tenenbaum, D. Schuurmans, and P. Abbeel
Cited in the paper.
This&that: Language-gesture controlled video generation for robot planning
B. Wang, N. Sridhar, C. Feng, M. Van der Merwe, A. Fishman, N. Fazeli, and J. J. Park
Cited in the paper.
Dust3r: Geometric 3d vision made easy
S. Wang, V. Leroy, Y. Cabon, B. Chidlovskii, and J. Revaud
Cited in the paper.
S. Yang, J. Walker, J. Parker-Holder, Y. Du, J. Bruce, A. Barreto, P. Abbeel, and D. Schuurmans · 2024
Closest in time.
Wonderworld: Interactive 3d scene generation from a single image
H.-X. Yu, H. Duan, C. Herrmann, W. T. Freeman, and J. Wu · 2024
Closest in time.