Fetching the paper…
Reading the bibliography…
End-to-end robot manipulation policies offer significant potential for enabling embodied agents to understand and interact with the world.
Affordance detection of tool parts from geometric features
A. Myers, C. L. Teo, C. Fermüller, and Y. Aloimonos · 2015
Earlier work this paper cites.
Trust region policy optimization
J. Schulman, S. Levine, P. Abbeel, M. Jordan, and P. Moritz · 2015
Earlier work this paper cites.
Deep spatial autoencoders for visuomotor learning
C. Finn, X. Y. Tan, Y. Duan, T. Darrell, S. Levine, and P. Abbeel · 2016
Earlier work this paper cites.
High precision grasp pose detection in dense clutter
M. Gualtieri, A. Ten Pas, K. Saenko, and R. Platt · 2016
Earlier work this paper cites.
End-to-end training of deep visuomotor policies
S. Levine, C. Finn, T. Darrell, and P. Abbeel · 2016
Earlier work this paper cites.
Learning contact-rich manipulation skills with guided policy search
S. Levine, N. Wagener, and P. Abbeel · 2016
Earlier work this paper cites.
Optimal path planning using rrt* based approaches: a survey and future directions
I. Noreen, A. Khan, and Z. Habib · 2016
Earlier work this paper cites.
Stable reinforcement learning with autoencoders for tactile and visual data
H. Van Hoof, N. Chen, M. Karl, P. van der Smagt, and J. Peters · 2016
Earlier work this paper cites.
Affordance detection for task-specific grasping using deep learning
M. Kokic, J. A. Stork, J. A. Haustein, and D. Kragic · 2017
Earlier work this paper cites.
What is an affordance? 40 years later
F. Osiurak, Y. Rossetti, and A. Badets · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Earlier work this paper cites.
Multi-view self-supervised deep learning for 6d pose estimation in the amazon picking challenge
A. Zeng, K.-T. Yu, S. Song, D. Suo, E. Walker, A. Rodriguez, and J. Xiao · 2017
Earlier work this paper cites.
Closing the loop for robotic grasping: A real-time, generative grasp synthesis approach
D. Morrison, P. Corke, and J. Leitner · 2018
Earlier work this paper cites.
Category-level 6d object pose recovery in depth images
C. Sahin and T.-K. Kim · 2018
Earlier work this paper cites.
Deep object pose estimation for semantic robotic grasping of household objects
J. Tremblay, T. To, B. Sundaralingam, Y. Xiang, D. Fox, and S. Birchfield · 2018
Earlier work this paper cites.
Robotic pick-and-place of novel objects in clutter with multi-affordance grasping and cross-domain image matching
A. Zeng, S. Song, K.-T. Yu, E. Donlon, F. R. Hogan, M. Bauza, D. Ma, O. Taylor, M. Liu, E. Romo, et al · 2018
Earlier work this paper cites.
Reinforcement and imitation learning for diverse visuomotor skills
Y. Zhu, Z. Wang, J. Merel, A. Rusu, T. Erez, S. Cabi, S. Tunyasuvunakool, J. Kramár, R. Hadsell, N. de Freitas, et al · 2018
Cited alongside, same era.
kpam-sc: Generalizable manipulation planning using keypoint affordance and shape completion
W. Gao and R. Tedrake · 2019
Cited alongside, same era.
Learning ambidextrous robot grasping policies
J. Mahler, M. Matl, V. Satish, M. Danielczuk, B. DeRose, S. McKinley, and K. Goldberg · 2019
Cited alongside, same era.
Keto: Learning keypoint representations for tool manipulation
Z. Qin, K. Fang, Y. Zhu, L. Fei-Fei, and S. Savarese · 2019
Cited alongside, same era.
Vision-based robotic grasping from object localization, object pose estimation to grasp estimation for parallel grippers: a review
G. Du, K. Wang, S. Lian, and K. Zhao · 2021
Cited alongside, same era.
Learning fine-grained bimanual manipulation with low-cost hardware
T. Z. Zhao, V. Kumar, S. Levine, and C. Finn · 2023
Later among the works it cites.
Pi0: A vision-language-action flow model for general robot control
K. Black, N. Brown, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, L. Groom, K. Hausman, B. Ichter, et al · 2024
Later among the works it cites.
Kilobot: A programming language for deploying perception-guided industrial manipulators at scale
W. Gao, J. Wang, X. Zhu, J. Zhong, Y. Shen, and Y. Ding · 2024
Later among the works it cites.
Rekep: Spatio-temporal reasoning of relational keypoint constraints for robotic manipulation
W. Huang, C. Wang, Y. Li, R. Zhang, and L. Fei-Fei · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
kpam 2.0: Feedback control for category-level robotic manipulation
W. Gao and R. Tedrake · 2021
Cited alongside, same era.
Manipulation planning for object re-orientation based on semantic segmentation keypoint detection
C.-C. Wong, L.-Y. Yeh, C.-C. Liu, C.-Y. Tsai, and H. Aoyama · 2021
Cited alongside, same era.
Code as policies: Language model programs for embodied control
J. Liang, W. Huang, F. Xia, P. Xu, K. Hausman, B. Ichter, P. Florence, and A. Zeng · 2022
Cited alongside, same era.
J. Urain, N. Funk, J. Peters, and G. Chalvatzaki · 2022
Cited alongside, same era.
J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al · 2023
Cited alongside, same era.
J. Bai, S. Bai, Y. Chu, Z. Cui, K. Dang, X. Deng, Y. Fan, W. Ge, Y. Han, F. Huang, et al · 2023
Cited alongside, same era.
Diffusion policy: Visuomotor policy learning via action diffusion
C. Chi, Z. Xu, S. Feng, E. Cousineau, Y. Du, B. Burchfiel, R. Tedrake, and S. Song · 2023
Cited alongside, same era.
T.-W. Ke, N. Gkanatsios, and K. Fragkiadaki · 2024
Later among the works it cites.
Openvla: An open-source vision-language-action model
M. J. Kim, K. Pertsch, S. Karamcheti, T. Xiao, A. Balakrishna, S. Nair, R. Rafailov, E. Foster, G. Lam, P. Sanketi, et al · 2024
Later among the works it cites.
Laso: Language-guided affordance segmentation on 3d object
Y. Li, N. Zhao, J. Xiao, C. Feng, X. Wang, and T.-s. Chua · 2024
Later among the works it cites.
Improved baselines with visual instruction tuning
H. Liu, C. Li, Y. Li, and Y. J. Lee · 2024
Later among the works it cites.
Rdt-1b: a diffusion foundation model for bimanual manipulation
S. Liu, L. Wu, B. Li, H. Tan, H. Chen, Z. Wang, K. Xu, H. Su, and J. Zhu · 2024
Later among the works it cites.
Groma: Localized visual tokenization for grounding multimodal large language models
C. Ma, Y. Jiang, J. Wu, Z. Yuan, and X. Qi · 2024
Later among the works it cites.
Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution
P. Wang, S. Bai, S. Tan, S. Wang, Z. Fan, J. Bai, K. Chen, X. Liu, J. Wang, W. Ge, et al · 2024
Later among the works it cites.
Robotic control via embodied chain-of-thought reasoning
M. Zawalski, W. Chen, K. Pertsch, O. Mees, C. Finn, and S. Levine · 2024
Later among the works it cites.
Omnimanip: Towards general robotic manipulation via object-centric interaction primitives as spatial constraints
M. Pan, J. Zhang, T. Wu, Y. Zhao, W. Gao, and H. Dong · 2025
Closest in time.
Foundationgrasp: Generalizable task-oriented grasping with foundation models
C. Tang, D. Huang, W. Dong, R. Xu, and H. Zhang · 2025
Closest in time.
Uad: Unsupervised affordance distillation for generalization in robotic manipulation
Y. Tang, W. Huang, Y. Wang, C. Li, R. Yuan, R. Zhang, J. Wu, and L. Fei-Fei · 2025
Closest in time.
Skil: Semantic keypoint imitation learning for generalizable data-efficient manipulation
S. Wang, J. You, Y. Hu, J. Li, and Y. Gao · 2025
Closest in time.