Fetching the paper…
Reading the bibliography…
Vision foundation models trained on massive amounts of visual data have shown unprecedented reasoning and planning skills in open-world settings.
Retargetting motion to new characters
M. Gleicher · 1998
Earlier work this paper cites.
Unsupervised learning for physical interaction through video prediction, 2016
C. Finn, I. Goodfellow, and S. Levine · 2016
Earlier work this paper cites.
Neural 3d mesh renderer, 2017
H. Kato, Y. Ushiku, and T. Harada · 2017
Earlier work this paper cites.
Perspective transformer nets: Learning single-view 3d object reconstruction without 3d supervision, 2017
X. Yan, J. Yang, E. Yumer, Y. Guo, and H. Lee · 2017
Earlier work this paper cites.
Stochastic variational video prediction
M. Babaeizadeh, C. Finn, D. Erhan, R. H. Campbell, and S. Levine · 2017
Earlier work this paper cites.
Time-contrastive networks: Self-supervised learning from video, 2018
P. Sermanet, C. Lynch, Y. Chebotar, J. Hsu, E. Jang, S. Schaal, and S. Levine · 2018
Earlier work this paper cites.
Stochastic adversarial video prediction, 2018
A. X. Lee, R. Zhang, F. Ebert, P. Abbeel, C. Finn, and S. Levine · 2018
Earlier work this paper cites.
Soft rasterizer: A differentiable renderer for image-based 3d reasoning
S. Liu, T. Li, W. Chen, and H. Li · 2019
Earlier work this paper cites.
Differentiable surface splatting for point-based geometry processing
W. Yifan, F. Serena, S. Wu, C. Öztireli, and O. Sorkine-Hornung · 2019
Earlier work this paper cites.
Deepsdf: Learning continuous signed distance functions for shape representation, 2019
J. J. Park, P. Florence, J. Straub, R. Newcombe, and S. Lovegrove · 2019
Earlier work this paper cites.
Occupancy networks: Learning 3d reconstruction in function space, 2019
L. Mescheder, M. Oechsle, M. Niemeyer, S. Nowozin, and A. Geiger · 2019
Earlier work this paper cites.
Camera-to-robot pose estimation from a single image, 2020
T. E. Lee, J. Tremblay, T. To, J. Cheng, T. Mosier, O. Kroemer, D. Fox, and S. Birchfield · 2020
Earlier work this paper cites.
Learning graph-convolutional representations for point cloud denoising, 2020
F. Pistilli, G. Fracastoro, D. Valsesia, and E. Magli · 2020
Earlier work this paper cites.
Nerf: Representing scenes as neural radiance fields for view synthesis, 2020
B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Barron, R. Ramamoorthi, and R. Ng · 2020
Earlier work this paper cites.
Curl: Contrastive unsupervised representations for reinforcement learning, 2020
A. Srinivas, M. Laskin, and P. Abbeel · 2020
Earlier work this paper cites.
Full-body visual self-modeling of robot morphologies, 2021
B. Chen, R. Kwiatkowski, C. Vondrick, and H. Lipson · 2021
Earlier work this paper cites.
Single-view robot pose and joint angle estimation via render and compare, 2021
Y. Labbé, J. Carpentier, M. Aubry, and J. Sivic · 2021
Earlier work this paper cites.
Escaping plato’s cave: 3d shape from adversarial rendering, 2021
P. Henzler, N. Mitra, and T. Ritschel · 2021
Cited alongside, same era.
Learning the predictability of the future
D. Suris, R. Liu, and C. Vondrick · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision, 2021
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever · 2021
Cited alongside, same era.
Learning generalizable robotic reward functions from ”in-the-wild” human videos, 2021
A. S. Chen, S. Nair, and C. Finn · 2021
Cited alongside, same era.
On the origins of self-modeling, 2022
R. Kwiatkowski, Y. Hu, B. Chen, and H. Lipson · 2022
Cited alongside, same era.
Modularity through attention: Efficient training and transfer of language-conditioned policies for robot manipulation, 2022
Humans as light bulbs: 3d human reconstruction from thermal reflection
R. Liu and C. Vondrick · 2023
Later among the works it cites.
Md-splatting: Learning metric deformation from 4d gaussians in highly deformable scenes, 2023
B. P. Duisterhof, Z. Mandi, Y. Yao, J.-W. Liu, M. Z. Shou, S. Song, and J. Ichnowski · 2023
Later among the works it cites.
Masked world models for visual control
Y. Seo, D. Hafner, H. Liu, F. Liu, S. James, K. Lee, and P. Abbeel · 2023
Later among the works it cites.
Video prediction models as rewards for reinforcement learning, 2023
A. Escontrela, A. Adeniji, W. Yan, A. Jain, X. B. Peng, K. Goldberg, Y. Lee, D. Hafner, and P. Abbeel · 2023
Later among the works it cites.
I2vgen-xl: High-quality image-to-video synthesis via cascaded diffusion models, 2023
S. Zhang, J. Wang, Y. Zhang, K. Zhao, H. Yuan, Z. Qin, X. Wang, D. Zhao, and J. Zhou · 2023
Later among the works it cites.
Unisim: A neural closed-loop sensor simulator, 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Zhou, S. Sonawani, M. Phielipp, S. Stepputtis, and H. B. Amor · 2022
Cited alongside, same era.
R3m: A universal visual representation for robot manipulation, 2022
S. Nair, A. Rajeswaran, V. Kumar, C. Finn, and A. Gupta · 2022
Cited alongside, same era.
Imagen video: High definition video generation with diffusion models, 2022
J. Ho, W. Chan, C. Saharia, J. Whang, R. Gao, A. Gritsenko, D. P. Kingma, B. Poole, M. Norouzi, D. J. Fleet, and T. Salimans · 2022
Cited alongside, same era.
Liv: Language-image representations and rewards for robotic control, 2023
Y. J. Ma, W. Liang, V. Som, V. Kumar, A. Zhang, O. Bastani, and D. Jayaraman · 2023
Cited alongside, same era.
Voxposer: Composable 3d value maps for robotic manipulation with language models, 2023
W. Huang, C. Wang, R. Zhang, Y. Li, J. Wu, and L. Fei-Fei · 2023
Cited alongside, same era.
3d gaussian splatting for real-time radiance field rendering
B. Kerbl, G. Kopanas, T. Leimkühler, and G. Drettakis · 2023
Cited alongside, same era.
Smpl: A skinned multi-person linear model
M. Loper, N. Mahmood, J. Romero, G. Pons-Moll, and M. J. Black · 2023
Cited alongside, same era.
Z. Yang, Y. Chen, J. Wang, S. Manivasagam, W.-C. Ma, A. J. Yang, and R. Urtasun · 2023
Later among the works it cites.
Video language planning, 2023
Y. Du, M. Yang, P. Florence, F. Xia, A. Wahid, B. Ichter, P. Sermanet, T. Yu, P. Abbeel, J. B. Tenenbaum, L. Kaelbling, A. Zeng, and J. Tompson · 2023
Later among the works it cites.
Compositional foundation models for hierarchical planning, 2023
A. Ajay, S. Han, Y. Du, S. Li, A. Gupta, T. Jaakkola, J. Tenenbaum, L. Kaelbling, A. Srivastava, and P. Agrawal · 2023
Later among the works it cites.
Zero-shot robotic manipulation with pretrained image-editing diffusion models
K. Black, M. Nakamoto, P. Atreya, H. Walke, C. Finn, A. Kumar, and S. Levine · 2023
Later among the works it cites.
Egocentric visual self-modeling for autonomous robot dynamics prediction and adaptation, 2024
Y. Hu, B. Chen, and H. Lipson · 2024
Closest in time.
Video as the new language for real-world decision making, 2024
S. Yang, J. Walker, J. Parker-Holder, Y. Du, J. Bruce, A. Barreto, P. Abbeel, and D. Schuurmans · 2024
Closest in time.
Generative camera dolly: Extreme monocular dynamic novel view synthesis
B. Van Hoorick, R. Wu, E. Ozguroglu, K. Sargent, R. Liu, P. Tokmakov, A. Dave, C. Zheng, and C. Vondrick · 2024
Closest in time.
Learning human-to-humanoid real-time whole-body teleoperation, 2024
T. He, Z. Luo, W. Xiao, C. Zhang, K. Kitani, C. Liu, and G. Shi · 2024
Closest in time.
Physdreamer: Physics-based interaction with 3d objects via video generation, 2024
T. Zhang, H.-X. Yu, R. Wu, B. Y. Feng, C. Zheng, N. Snavely, J. Wu, and W. T. Freeman · 2024
Closest in time.
Physgaussian: Physics-integrated 3d gaussians for generative dynamics, 2024
T. Xie, Z. Zong, Y. Qiu, X. Li, Y. Feng, Y. Yang, and C. Jiang · 2024
Closest in time.
Video generation models as world simulators
T. Brooks, B. Peebles, C. Holmes, W. DePue, Y. Guo, L. Jing, D. Schnurr, J. Taylor, T. Luhman, E. Luhman, C. Ng, R. Wang, and A. Ramesh · 2024
Closest in time.