Fetching the paper…
Reading the bibliography…
Visual representations play a crucial role in developing generalist robotic policies.
The” something something” video database for learning and evaluating visual common sense
Goyal, R., Ebrahimi Kahou, S., Michalski, V., Materzynska, J., Westphal, S., Kim, H., Haenel, V., Fruend, I., Yianilos, P., Mueller-Freitag, M., et al · 2017
Earlier work this paper cites.
Towards accurate generative models of video: A new metric & challenges
Unterthiner, T., Van Steenkiste, S., Kurach, K., Marinier, R., Michalski, M., and Gelly, S · 2018
Earlier work this paper cites.
Unsupervised learning of visual features by contrasting cluster assignments
Caron, M., Misra, I., Mairal, J., Goyal, P., Bojanowski, P., and Joulin, A · 2020
Earlier work this paper cites.
A simple framework for contrastive learning of visual representations
Chen, T., Kornblith, S., Norouzi, M., and Hinton, G · 2020
Earlier work this paper cites.
Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning
Yu, T., Quillen, D., He, Z., Julian, R., Hausman, K., Finn, C., and Levine, S · 2020
Earlier work this paper cites.
Beit: Bert pre-training of image transformers
Bao, H., Dong, L., Piao, S., and Wei, F · 2021
Earlier work this paper cites.
An empirical study of training self-supervised vision transformers
Chen, X., Xie, S., and He, K · 2021
Earlier work this paper cites.
Bridge data: Boosting generalization of robotic skills with cross-domain datasets
Ebert, F., Yang, Y., Schmeckpeper, K., Bucher, B., Georgakis, G., Daniilidis, K., Finn, C., and Levine, S · 2021
Earlier work this paper cites.
Perceiver: General perception with iterative attention
Jaegle, A., Gimeno, F., Brock, A., Vinyals, O., Zisserman, A., and Carreira, J · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Earlier work this paper cites.
Flamingo: a visual language model for few-shot learning
Alayrac, J.-B., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., et al · 2022
Earlier work this paper cites.
Data2vec: A general framework for self-supervised learning in speech, vision and language
Baevski, A., Hsu, W.-N., Xu, Q., Babu, A., Gu, J., and Auli, M · 2022
Earlier work this paper cites.
Rt-1: Robotics transformer for real-world control at scale
Brohan, A., Brown, N., Carbajal, J., Chebotar, Y., Dabis, J., Finn, C., Gopalakrishnan, K., Hausman, K., Herzog, A., Hsu, J., et al · 2022
Earlier work this paper cites.
Ego4d: Around the world in 3,000 hours of egocentric video
Grauman, K., Westbury, A., Byrne, E., Chavis, Z., Furnari, A., Girdhar, R., Hamburger, J., Jiang, H., Liu, M., Liu, X., et al · 2022
Earlier work this paper cites.
Masked autoencoders are scalable vision learners
He, K., Chen, X., Xie, S., Li, Y., Dollár, P., and Girshick, R · 2022
Earlier work this paper cites.
Video diffusion models
Ho, J., Salimans, T., Gritsenko, A., Chan, W., Norouzi, M., and Fleet, D. J · 2022
Earlier work this paper cites.
Cogvideo: Large-scale pretraining for text-to-video generation via transformers
Hong, W., Ding, M., Zheng, W., Liu, X., and Tang, J · 2022
Earlier work this paper cites.
Vip: Towards universal visual reward and representation via value-implicit pre-training
Ma, Y. J., Sodhani, S., Jayaraman, D., Bastani, O., Kumar, V., and Zhang, A · 2022
Cited alongside, same era.
Calvin: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks
Mees, O., Hermann, L., Rosete-Beas, E., and Burgard, W · 2022
Cited alongside, same era.
R3m: A universal visual representation for robot manipulation
Nair, S., Rajeswaran, A., Kumar, V., Finn, C., and Gupta, A · 2022
Cited alongside, same era.
The unsurprising effectiveness of pre-trained vision models for control
Parisi, S., Rajeswaran, A., Purushwalkam, S., and Gupta, A · 2022
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B · 2022
Cited alongside, same era.
Real-world robot learning with masked visual pre-training
Radosavovic, I., Xiao, T., James, S., Abbeel, P., Malik, J., and Darrell, T · 2023
Later among the works it cites.
Denoising diffusion autoencoders are unified self-supervised learners
Xiang, W., Yang, H., Huang, D., and Wang, Y · 2023
Later among the works it cites.
Offline visual representation learning for embodied navigation
Yadav, K., Ramrakhya, R., Majumdar, A., Berges, V.-P., Kuhar, S., Batra, D., Baevski, A., and Maksymets, O · 2023
Later among the works it cites.
Gen2act: Human video generation in novel scenarios enables generalizable robot manipulation
Bharadhwaj, H., Dwibedi, D., Gupta, A., Tulsiani, S., Doersch, C., Xiao, T., Shah, D., Xia, F., Sadigh, D., and Kirmani, S · 2024
Closest in time.
Video generation models as world simulators
Brooks, T., Peebles, B., Holmes, C., DePue, W., Guo, Y., Jing, L., Schnurr, D., Taylor, J., Luhman, T., Luhman, E., Ng, C., Wang, R., and Ramesh, A · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zero-shot robotic manipulation with pretrained image-editing diffusion models
Black, K., Nakamoto, M., Atreya, P., Walke, H., Finn, C., Kumar, A., and Levine, S · 2023
Cited alongside, same era.
Rt-2: Vision-language-action models transfer web knowledge to robotic control
Brohan, A., Brown, N., Carbajal, J., Chebotar, Y., Chen, X., Choromanski, K., Ding, T., Driess, D., Dubey, A., Finn, C., et al · 2023
Cited alongside, same era.
Instructpix2pix: Learning to follow image editing instructions
Brooks, T., Holynski, A., and Efros, A. A · 2023
Cited alongside, same era.
Control-a-video: Controllable text-to-video generation with diffusion models
Chen, W., Ji, Y., Wu, J., Wu, H., Xie, P., Li, J., Xia, X., Xiao, X., and Lin, L · 2023
Cited alongside, same era.
Diffusion policy: Visuomotor policy learning via action diffusion
Chi, C., Feng, S., Du, Y., Xu, Z., Cousineau, E., Burchfiel, B., and Song, S · 2023
Cited alongside, same era.
Seer: Language instructed video prediction with latent diffusion models
Gu, X., Wen, C., Ye, W., Song, J., and Gao, Y · 2023
Cited alongside, same era.
Language-driven representation learning for robotics
Karamcheti, S., Nair, S., Chen, A. S., Kollar, T., Finn, C., Sadigh, D., and Liang, P · 2023
Cited alongside, same era.
Learning universal policies via text-guided video generation
Du, Y., Yang, S., Dai, B., Dai, H., Nachum, O., Tenenbaum, J., Schuurmans, D., and Abbeel, P · 2024
Closest in time.
Prediction with action: Visual policy learning via joint denoising process
Guo, Y., Hu, Y., Zhang, J., Wang, Y.-J., Chen, X., Lu, C., and Chen, J · 2024
Closest in time.
Pre-trained text-to-image diffusion models are versatile representation learners for control
Gupta, G., Yadav, K., Gal, Y., Batra, D., Kira, Z., Lu, C., and Rudner, T. G · 2024
Closest in time.
Robouniview: Visual-language model with unified view representation for robotic manipulation
Liu, F., Yan, F., Zheng, L., Feng, C., Huang, Y., and Ma, L · 2024
Closest in time.
Diffusion hyperfeatures: Searching through time and space for semantic correspondence
Luo, G., Dunlap, L., Park, D. H., Holynski, A., and Darrell, T · 2024
Closest in time.
Multimodal diffusion transformer: Learning versatile behavior from multimodal goals, 2024
Reuss, M., Ömer Erdinç Yağmurlu, Wenzel, F., and Lioutikov, R · 2024
Closest in time.
Octo: An open-source generalist robot policy
Team, O. M., Ghosh, D., Walke, H., Pertsch, K., Black, K., Mees, O., Dasari, S., Hejna, J., Kreiman, T., Xu, C., et al · 2024
Closest in time.
Vidman: Exploiting implicit dynamics from video diffusion model for effective robot manipulation
Wen, Y., Lin, J., Zhu, Y., Han, J., Xu, H., Zhao, S., and Liang, X · 2024
Closest in time.
Cogvideox: Text-to-video diffusion models with an expert transformer
Yang, Z., Teng, J., Zheng, W., Ding, M., Huang, S., Xu, J., Yang, Y., Hong, W., Zhang, X., Feng, G., et al · 2024
Closest in time.
Latent action pretraining from videos
Ye, S., Jang, J., Jeon, B., Joo, S., Yang, J., Peng, B., Mandlekar, A., Tan, R., Chao, Y.-W., Lin, B. Y., et al · 2024
Closest in time.
Improving vision-language-action model with online reinforcement learning
Guo, Y., Zhang, J., Chen, X., Ji, X., Wang, Y.-J., Hu, Y., and Chen, J · 2025
Closest in time.