Fetching the paper…
Reading the bibliography…
Despite tremendous progress in dexterous manipulation, current visuomotor policies remain fundamentally limited by two challenges: they struggle to generalize under perceptual or behavioral distribution shifts, and their performance is constrained by the size of human demonstration data.
Alvinn: An autonomous land vehicle in a neural network
D. A. Pomerleau · 1988
Earlier work this paper cites.
A framework for behavioural cloning
M. Bain and C. Sammut · 1995
Earlier work this paper cites.
A tutorial on energy-based learning
Y. LeCun, S. Chopra, R. Hadsell, M. Ranzato, and F. Huang · 2006
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
S. Ross, G. Gordon, and D. Bagnell · 2011
Earlier work this paper cites.
Video (language) modeling: A baseline for generative models of natural videos. arxiv 2014
M. Ranzato, A. Szlam, J. Bruna, M. Mathieu, R. Collobert, and S. Chopra · 2014
Earlier work this paper cites.
Unsupervised learning for physical interaction through video prediction
C. Finn, I. Goodfellow, and S. Levine · 2016
Earlier work this paper cites.
Generating videos with scene dynamics
C. Vondrick, H. Pirsiavash, and A. Torralba · 2016
Earlier work this paper cites.
Deep imitation learning for complex manipulation tasks from virtual reality teleoperation
T. Zhang, Z. McCarthy, O. Jow, D. Lee, X. Chen, K. Goldberg, and P. Abbeel · 2018
Earlier work this paper cites.
Time-contrastive networks: Self-supervised learning from video
P. Sermanet, C. Lynch, Y. Chebotar, J. Hsu, E. Jang, S. Schaal, S. Levine, and G. Brain · 2018
Earlier work this paper cites.
Stochastic variational video prediction
M. Babaeizadeh, C. Finn, D. Erhan, R. H. Campbell, and S. Levine · 2018
Earlier work this paper cites.
Stochastic adversarial video prediction
A. X. Lee, R. Zhang, F. Ebert, P. Abbeel, C. Finn, and S. Levine · 2018
Earlier work this paper cites.
CCNet: Extracting high quality monolingual datasets from web crawl data
G. Wenzek, M.-A. Lachaux, A. Conneau, V. Chaudhary, F. Guzmán, A. Joulin, and E. Grave · 2019
Earlier work this paper cites.
Self-supervised correspondence in visuomotor policy learning
P. Florence, L. Manuelli, and R. Tedrake · 2019
Earlier work this paper cites.
Implicit generation and modeling with energy based models
Y. Du and I. Mordatch · 2019
Earlier work this paper cites.
CURL: Contrastive unsupervised representations for reinforcement learning
M. Laskin, A. Srinivas, and P. Abbeel · 2020
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever · 2021
Earlier work this paper cites.
Learning the predictability of the future
D. Suris, R. Liu, and C. Vondrick · 2021
Earlier work this paper cites.
Learning generalizable robotic reward functions from” in-the-wild” human videos
A. S. Chen, S. Nair, and C. Finn · 2021
Earlier work this paper cites.
RT-1: Robotics transformer for real-world control at scale
A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, J. Hsu, et al · 2022
Earlier work this paper cites.
Masked visual pre-training for motor control
T. Xiao, I. Radosavovic, T. Darrell, and J. Malik · 2022
Earlier work this paper cites.
Implicit behavioral cloning
P. Florence, C. Lynch, A. Zeng, O. A. Ramirez, A. Wahid, L. Downs, A. Wong, J. Lee, I. Mordatch, and J. Tompson · 2022
Cited alongside, same era.
R3M: A universal visual representation for robot manipulation
S. Nair, A. Rajeswaran, V. Kumar, C. Finn, and A. Gupta · 2022
Cited alongside, same era.
Imagen video: High definition video generation with diffusion models, 2022
J. Ho, W. Chan, C. Saharia, J. Whang, R. Gao, A. Gritsenko, D. P. Kingma, B. Poole, M. Norouzi, D. J. Fleet, and T. Salimans · 2022
Cited alongside, same era.
Diffusion policy: Visuomotor policy learning via action diffusion
C. Chi, S. Feng, Y. Du, Z. Xu, E. Cousineau, B. Burchfiel, and S. Song · 2023
Cited alongside, same era.
Stable video diffusion: Scaling latent video diffusion models to large datasets
A. Blattmann, T. Dockhorn, S. Kulal, D. Mendelevitch, M. Kilian, D. Lorenz, Y. Levi, Z. English, V. Voleti, A. Letts, et al · 2023
Cited alongside, same era.
Octo: An open-source generalist robot policy
O. M. Team, D. Ghosh, H. Walke, K. Pertsch, K. Black, O. Mees, S. Dasari, J. Hejna, T. Kreiman, C. Xu, et al · 2024
Later among the works it cites.
Video generation models as world simulators
T. Brooks, B. Peebles, C. Holmes, W. DePue, Y. Guo, L. Jing, D. Schnurr, J. Taylor, T. Luhman, E. Luhman, C. Ng, R. Wang, and A. Ramesh · 2024
Later among the works it cites.
Dreamitate: Real-world visuomotor policy learning via video generation
J. Liang, R. Liu, E. Ozguroglu, S. Sudhakar, A. Dave, P. Tokmakov, S. Song, and C. Vondrick · 2024
Later among the works it cites.
Video prediction policy: A generalist robot policy with predictive visual representations
Y. Hu, Y. Guo, P. Wang, X. Chen, Y.-J. Wang, J. Zhang, K. Sreenath, C. Lu, and J. Chen · 2024
Later among the works it cites.
VQ-BeT: Behavior generation with latent actions
S. Lee, Y. Wang, H. Etukuru, H. J. Kim, N. M. M. Shafiullah, and L. Pinto · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning universal policies via text-guided video generation
Y. Du, S. Yang, B. Dai, H. Dai, O. Nachum, J. Tenenbaum, D. Schuurmans, and P. Abbeel · 2023
Cited alongside, same era.
VoxPoser: Composable 3D value maps for robotic manipulation with language models
W. Huang, C. Wang, R. Zhang, Y. Li, J. Wu, and L. Fei-Fei · 2023
Cited alongside, same era.
Learning fine-grained bimanual manipulation with low-cost hardware
T. Z. Zhao, V. Kumar, S. Levine, and C. Finn · 2023
Cited alongside, same era.
Masked world models for visual control
Y. Seo, D. Hafner, H. Liu, F. Liu, S. James, K. Lee, and P. Abbeel · 2023
Cited alongside, same era.
Robot learning with masked visual pre-training
I. Radosavovic, T. Xiao, S. James, P. Abbeel, J. Malik, and T. Darrell · 2023
Cited alongside, same era.
LIV: Language-image representations and rewards for robotic control
Y. J. Ma, V. Kumar, A. Zhang, O. Bastani, and D. Jayaraman · 2023
Cited alongside, same era.
Video prediction models as rewards for reinforcement learning
A. Escontrela, A. Adeniji, W. Yan, A. Jain, X. B. Peng, K. Goldberg, Y. Lee, D. Hafner, and P. Abbeel · 2023
Cited alongside, same era.
C.-L. Cheang, G. Chen, Y. Jing, T. Kong, H. Li, Y. Li, Y. Liu, H. Wu, J. Xu, Y. Yang, et al · 2024
Later among the works it cites.
RoboCasa: Large-scale simulation of everyday tasks for generalist robots
S. Nasiriany, A. Maddukuri, L. Zhang, A. Parikh, A. Lo, A. Joshi, A. Mandlekar, and Y. Zhu · 2024
Later among the works it cites.
A dual process VLA: Efficient robotic manipulation leveraging VLM
B. Han, J. Kim, and J. Jang · 2024
Later among the works it cites.
3D diffuser actor: Policy diffusion with 3D scene representations
T.-W. Ke, N. Gkanatsios, and K. Fragkiadaki · 2024
Later among the works it cites.
3D diffusion policy
Y. Ze, G. Zhang, K. Zhang, C. Hu, M. Wang, and H. Xu · 2024
Later among the works it cites.
Scaling rectified flow transformers for high-resolution image synthesis
P. Esser, S. Kulal, A. Blattmann, R. Entezari, J. Müller, H. Saini, Y. Levi, D. Lorenz, A. Sauer, F. Boesel, et al · 2024
Later among the works it cites.
Fast ode-based sampling for diffusion models in around 5 steps
Z. Zhou, D. Chen, C. Wang, and C. Chen · 2024
Later among the works it cites.
Autoregressive image generation without vector quantization
T. Li, Y. Tian, H. Li, M. Deng, and K. He · 2024
Later among the works it cites.
Universal manipulation interface: In-the-wild robot teaching without in-the-wild robots
C. Chi, Z. Xu, C. Pan, E. Cousineau, B. Burchfiel, S. Feng, R. Tedrake, and S. Song · 2024
Later among the works it cites.
π 0 \pi_{0} : A vision-language-action flow model for general robot control
K. Black, N. Brown, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, L. Groom, K. Hausman, B. Ichter, et al · 2025
Closest in time.
A careful examination of large behavior models for multitask dexterous manipulation
J. Barreiros, A. Beaulieu, A. Bhat, R. Cory, E. Cousineau, H. Dai, C.-H. Fang, K. Hashimoto, M. Z. Irshad, M. Itkina, et al · 2025
Closest in time.
Prediction with action: Visual policy learning via joint denoising process
Y. Guo, Y. Hu, J. Zhang, Y.-J. Wang, X. Chen, C. Lu, and J. Chen · 2025
Closest in time.
Unified video action model
S. Li, Y. Gao, D. Sadigh, and S. Song · 2025
Closest in time.
Unified world models: Coupling video and action diffusion for pretraining on large robotic datasets
C. Zhu, R. Yu, S. Feng, B. Burchfiel, P. Shah, and A. Gupta · 2025
Closest in time.
Towards fusing point cloud and visual representations for imitation learning
A. Donat, X. Jia, X. Huang, A. Taranovic, D. Blessing, G. Li, H. Zhou, H. Zhang, R. Lioutikov, and G. Neumann · 2025
Closest in time.
GR00T N1: An open foundation model for generalist humanoid robots
J. Bjorck, F. Castañeda, N. Cherniadev, X. Da, R. Ding, L. Fan, Y. Fang, D. Fox, F. Hu, S. Huang, et al · 2025
Closest in time.