Fetching the paper…
Reading the bibliography…
Learning for manipulation requires using policies that have access to rich sensory information such as point clouds or RGB images.
The farthest point strategy for progressive image sampling
Eldar, Y., Lindenbaum, M., Porat, M., and Zeevi, Y. Y · 1994
Earlier work this paper cites.
Denoising diffusion implicit models
Song, J., Meng, C., and Ermon, S · 2010
Earlier work this paper cites.
A connection between score matching and denoising autoencoders
Vincent, P · 2011
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Film: Visual reasoning with a general conditioning layer
Perez, E., Strub, F., De Vries, H., Dumoulin, V., and Courville, A · 2018
Earlier work this paper cites.
Goal-conditioned imitation learning
Ding, Y., Florensa, C., Abbeel, P., and Phielipp, M · 2019
Earlier work this paper cites.
Language-conditioned imitation learning for robot manipulation tasks
Stepputtis, S., Campbell, J., Phielipp, M., Lee, S., Baral, C., and Ben Amor, H · 2020
Earlier work this paper cites.
What matters in learning from offline human demonstrations for robot manipulation
Mandlekar, A., Xu, D., Wong, J., Nasiriany, S., Wang, C., Kulkarni, R., Fei-Fei, L., Savarese, S., Zhu, Y., and Martín-Martín, R · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision, 2021
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I · 2021
Earlier work this paper cites.
Rt-1: Robotics transformer for real-world control at scale
Brohan, A., Brown, N., Carbajal, J., Chebotar, Y., Dabis, J., Finn, C., Gopalakrishnan, K., Hausman, K., Herzog, A., Hsu, J., et al · 2022
Earlier work this paper cites.
Objaverse: A universe of annotated 3d objects
Deitke, M., Schwenk, D., Salvador, J., Weihs, L., Michel, O., VanderBilt, E., Schmidt, L., Ehsani, K., Kembhavi, A., and Farhadi, A · 2022
Earlier work this paper cites.
Elucidating the design space of diffusion-based generative models
Karras, T., Aittala, M., Aila, T., and Laine, S · 2022
Cited alongside, same era.
Scalable diffusion models with transformers
Peebles, W. and Xie, S · 2022
Cited alongside, same era.
Diffusion policy: Visuomotor policy learning via action diffusion
Chi, C., Feng, S., Du, Y., Xu, Z., Cousineau, E., Burchfiel, B., and Song, S · 2023
Cited alongside, same era.
Act3d: 3d feature field transformers for multi-task robotic manipulation
Gervet, T., Xian, Z., Gkanatsios, N., and Fragkiadaki, K · 2023
Cited alongside, same era.
Vision-language foundation models as effective robot imitators
Li, X., Liu, M., Zhang, H., Yu, C., Xu, J., Wu, H., Cheang, C., Jing, Y., Zhang, W., Liu, H., et al · 2023
Cited alongside, same era.
Sugar: Pre-training 3d visual representations for robotics
Chen, S., Garcia, R., Laptev, I., and Schmid, C · 2024
Later among the works it cites.
3d diffuser actor: Policy diffusion with 3d scene representations
Ke, T.-W., Gkanatsios, N., and Fragkiadaki, K · 2024
Later among the works it cites.
Openvla: An open-source vision-language-action model
Kim, M., Pertsch, K., Karamcheti, S., Xiao, T., Balakrishna, A., Nair, S., Rafailov, R., Foster, E., Lam, G., Sanketi, P., Vuong, Q., Kollar, T., Burchfiel, B., Tedrake, R., Sadigh, D., Levine, S., Liang, P., and Finn, C · 2024
Later among the works it cites.
Rdt-1b: a diffusion foundation model for bimanual manipulation
Liu, S., Wu, L., Li, B., Tan, H., Chen, H., Wang, Z., Xu, K., Su, H., and Zhu, J · 2024
Later among the works it cites.
Robocasa: Large-scale simulation of everyday tasks for generalist robots
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Perceiver-actor: A multi-task transformer for robotic manipulation
Shridhar, M., Manuelli, L., and Fox, D · 2023
Cited alongside, same era.
Learning fine-grained bimanual manipulation with low-cost hardware
Zhao, T. Z., Kumar, V., Levine, S., and Finn, C · 2023
Cited alongside, same era.
Roboagent: Generalization and efficiency in robot manipulation via semantic augmentations and action chunking
Bharadhwaj, H., Vakil, J., Sharma, M., Gupta, A., Tulsiani, S., and Kumar, V · 2024
Cited alongside, same era.
pi_0: A vision-language-action flow model for general robot control
Black, K., Brown, N., Driess, D., Esmail, A., Equi, M., Finn, C., Fusai, N., Groom, L., Hausman, K., Ichter, B., et al · 2024
Cited alongside, same era.
Gr-2: A generative video-language-action model with web-scale knowledge for robot manipulation
Cheang, C.-L., Chen, G., Jing, Y., Kong, T., Li, H., Li, Y., Liu, Y., Wu, H., Xu, J., Yang, Y., et al · 2024
Cited alongside, same era.
Pointnet: Deep learning on point sets for 3d classification and segmentation
Qi, C. R., Su, H., Mo, K., and Guibas, L. J
Cited in the paper.
Pointnet++: Deep hierarchical feature learning on point sets in a metric space
Qi, C. R., Yi, L., Su, H., and Guibas, L. J
Cited in the paper.
Nasiriany, S., Maddukuri, A., Zhang, L., Parikh, A., Lo, A., Joshi, A., Mandlekar, A., and Zhu, Y · 2024
Later among the works it cites.
Point cloud models improve visual robustness in robotic learners
Peri, S., Lee, I., Kim, C., Fuxin, L., Hermans, T., and Lee, S · 2024
Later among the works it cites.
Multimodal diffusion transformer: Learning versatile behavior from multimodal goals
Reuss, M., Yağmurlu, Ö. E., Wenzel, F., and Lioutikov, R · 2024
Later among the works it cites.
Ze, Y., Zhang, G., Zhang, K., Hu, C., Wang, M., and Xu, H · 2024
Later among the works it cites.
Point cloud matters: Rethinking the impact of different observation spaces on robot learning
Zhu, H., Wang, Y., Huang, D., Ye, W., Ouyang, W., and He, T · 2024
Later among the works it cites.
Gr-mg: Leveraging partially-annotated data via multi-modal goal-conditioned policy
Li, P., Wu, H., Huang, Y., Cheang, C., Wang, L., and Kong, T · 2025
Closest in time.