Fetching the paper…
Reading the bibliography…
This work proposes a retrieve-and-transfer framework for zero-shot robotic manipulation, dubbed RAM, featuring generalizability across various objects, environments, and embodiments.
Random sample consensus: a paradigm for model fitting with applications to image analysis and automated cartography
M. A. Fischler and R. C. Bolles · 1981
Earlier work this paper cites.
Least squares quantization in pcm
S. Lloyd · 1982
Earlier work this paper cites.
Reducing the barrier to entry of complex robotic software: a moveit! case study
D. Coleman, I. Sucan, S. Chitta, and N. Correll · 2014
Earlier work this paper cites.
The ycb object and model set: Towards common benchmarks for manipulation research
B. Calli, A. Singh, A. Walsman, S. Srinivasa, P. Abbeel, and A. M. Dollar · 2015
Earlier work this paper cites.
One-shot visual imitation learning via meta-learning, 2017
C. Finn, T. Yu, T. Zhang, P. Abbeel, and S. Levine · 2017
Earlier work this paper cites.
Scaling egocentric vision: The epic-kitchens dataset
D. Damen, H. Doughty, G. M. Farinella, S. Fidler, A. Furnari, E. Kazakos, D. Moltisanti, J. Munro, T. Perrett, W. Price, et al · 2018
Earlier work this paper cites.
Rlbench: The robot learning benchmark & learning environment
S. James, Z. Ma, D. R. Arrojo, and A. J. Davison · 2020
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
P. Lewis, E. Perez, A. Piktus, F. Petroni, V. Karpukhin, N. Goyal, H. Küttler, M. Lewis, W.-t. Yih, T. Rocktäschel, et al · 2020
Earlier work this paper cites.
Where2act: From pixels to actions for articulated 3d objects
K. Mo, L. J. Guibas, M. Mukadam, A. Gupta, and S. Tulsiani · 2021
Earlier work this paper cites.
Graspness discovery in clutters for fast and accurate grasp detection
C. Wang, H.-S. Fang, M. Gou, H. Fang, J. Gao, and C. Lu · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Earlier work this paper cites.
Asformer: Transformer for action segmentation
F. Yi, H. Wen, and T. Jiang · 2021
Earlier work this paper cites.
Emerging properties in self-supervised vision transformers
M. Caron, H. Touvron, I. Misra, H. Jégou, J. Mairal, P. Bojanowski, and A. Joulin · 2021
Earlier work this paper cites.
Isaac gym: High performance gpu-based physics simulation for robot learning, 2021
V. Makoviychuk, L. Wawrzyniak, Y. Guo, M. Lu, K. Storey, M. Macklin, D. Hoeller, N. Rudin, A. Allshire, A. Handa, and G. State · 2021
Earlier work this paper cites.
Calvin: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks
O. Mees, L. Hermann, E. Rosete-Beas, and W. Burgard · 2022
Earlier work this paper cites.
Learning affordance grounding from exocentric images
H. Luo, W. Zhai, J. Zhang, Y. Cao, and D. Tao · 2022
Earlier work this paper cites.
Ego4d: Around the world in 3,000 hours of egocentric video
K. Grauman, A. Westbury, E. Byrne, Z. Chavis, A. Furnari, R. Girdhar, J. Hamburger, H. Jiang, M. Liu, X. Liu, et al · 2022
Earlier work this paper cites.
Hoi4d: A 4d egocentric dataset for category-level human-object interaction
Y. Liu, Y. Liu, C. Jiang, K. Lyu, W. Wan, H. Shen, B. Liang, Z. Fu, H. Wang, and L. Yi · 2022
Earlier work this paper cites.
Joint hand motion and interaction hotspots prediction from egocentric videos
S. Liu, S. Tripathi, S. Majumdar, and X. Wang · 2022
Earlier work this paper cites.
VAT-mart: Learning visual action trajectory proposals for manipulating 3d ARTiculated objects
R. Wu, Y. Zhao, K. Mo, Z. Guo, Y. Wang, T. Wu, Q. Fan, X. Chen, L. Guibas, and H. Dong · 2022
Earlier work this paper cites.
Adaafford: Learning to adapt manipulation affordance for 3d articulated objects via few-shot interactions
Y. Wang, R. Wu, K. Mo, J. Ke, Q. Fan, L. J. Guibas, and H. Dong · 2022
Earlier work this paper cites.
Towards more generalizable one-shot visual imitation learning
Z. Mandi, F. Liu, K. Lee, and P. Abbeel · 2022
Earlier work this paper cites.
Human-to-robot imitation in the wild
S. Bahl, A. Gupta, and D. Pathak · 2022
Earlier work this paper cites.
High-resolution image synthesis with latent diffusion models
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer · 2022
Earlier work this paper cites.
Fine-grained egocentric hand-object segmentation: Dataset, model, and applications
L. Zhang, S. Zhou, S. Stent, and J. Shi · 2022
Earlier work this paper cites.
Open x-embodiment: Robotic learning datasets and rt-x models
A. Padalkar, A. Pooley, A. Jain, A. Bewley, A. Herzog, A. Irpan, A. Khazatsky, A. Rai, A. Singh, A. Brohan, et al · 2023
Cited alongside, same era.
Y. Xu, W. Wan, J. Zhang, H. Liu, Z. Shan, H. Shen, R. Wang, H. Geng, Y. Weng, J. Chen, et al · 2023
Cited alongside, same era.
W. Wan, H. Geng, Y. Liu, Z. Shan, Y. Yang, L. Yi, and H. Wang · 2023
Cited alongside, same era.
Affordances from human videos as a versatile representation for robotics
S. Bahl, R. Mendonca, L. Chen, U. Jain, and D. Pathak · 2023
Cited alongside, same era.
Anygrasp: Robust and efficient grasp perception in spatial and temporal domains
Telling left from right: Identifying geometry-aware semantic correspondence
J. Zhang, C. Herrmann, J. Hur, E. Chen, V. Jampani, D. Sun, and M.-H. Yang · 2023
Later among the works it cites.
Manipllm: Embodied multimodal large language model for object-centric robotic manipulation
X. Li, M. Zhang, Y. Geng, H. Geng, Y. Long, Y. Shen, R. Zhang, J. Liu, and H. Dong · 2023
Later among the works it cites.
Stopnet: Multiview-based 6-dof suction detection for transparent objects on production lines
Y. Kuang, Q. Han, D. Li, Q. Dai, L. Ding, D. Sun, H. Zhao, and H. Wang · 2023
Later among the works it cites.
Gamma: Graspability-aware mobile manipulation policy learning based on online grasping pose fusion
J. Zhang, N. Gireesh, J. Wang, X. Fang, C. Xu, W. Chen, L. Dai, and H. Wang · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
H.-S. Fang, C. Wang, H. Fang, M. Gou, J. Liu, H. Yan, W. Liu, Y. Xie, and C. Lu · 2023
Cited alongside, same era.
Curobo: Parallelized collision-free robot motion generation
B. Sundaralingam, S. K. S. Hari, A. Fishman, C. Garrett, K. Van Wyk, V. Blukis, A. Millane, H. Oleynikova, A. Handa, F. Ramos, et al · 2023
Cited alongside, same era.
End-to-end affordance learning for robotic manipulation
Y. Geng, B. An, H. Geng, Y. Chen, Y. Yang, and H. Dong · 2023
Cited alongside, same era.
Voxposer: Composable 3d value maps for robotic manipulation with language models
W. Huang, C. Wang, R. Zhang, Y. Li, J. Wu, and L. Fei-Fei · 2023
Cited alongside, same era.
Look before you leap: Unveiling the power of gpt-4v in robotic vision-language planning
Y. Hu, F. Lin, T. Zhang, L. Yi, and Y. Gao · 2023
Cited alongside, same era.
Sage: Bridging semantic and actionable parts for generalizable articulated-object manipulation under language instructions, 2023
H. Geng, S. Wei, C. Deng, B. Shen, H. Wang, and L. Guibas · 2023
Cited alongside, same era.
Make a donut: Language-guided hierarchical emd-space planning for zero-shot deformable object manipulation, 2023
Y. You, B. Shen, C. Deng, H. Geng, H. Wang, and L. Guibas · 2023
Cited alongside, same era.
R. Gong, J. Huang, Y. Zhao, H. Geng, X. Gao, Q. Wu, W. Ai, Z. Zhou, D. Terzopoulos, S.-C. Zhu, B. Jia, and S. Huang · 2023
Cited alongside, same era.
The dawn of lmms: Preliminary explorations with gpt-4v(ision), 2023
Z. Yang, L. Li, K. Lin, J. Wang, C.-C. Lin, Z. Liu, and L. Wang · 2023
Later among the works it cites.
Droid: A large-scale in-the-wild robot manipulation dataset
A. Khazatsky, K. Pertsch, S. Nair, A. Balakrishna, S. Dasari, S. Karamcheti, S. Nasiriany, M. K. Srirama, L. Y. Chen, K. Ellis, et al · 2024
Closest in time.
Sora: A review on background, technology, limitations, and opportunities of large vision models, 2024
Y. Liu, K. Zhang, Y. Li, Z. Yan, C. Gao, R. Chen, Z. Yuan, Y. Huang, H. Sun, J. Gao, L. He, and L. Sun · 2024
Closest in time.
Rt-sketch: Goal-conditioned imitation learning from hand-drawn sketches, 2024
P. Sundaresan, Q. Vuong, J. Gu, P. Xu, T. Xiao, S. Kirmani, T. Yu, M. Stark, A. Jain, K. Hausman, D. Sadigh, J. Bohg, and S. Schaal · 2024
Closest in time.
General flow as foundation affordance for scalable robot learning, 2024
C. Yuan, C. Wen, T. Zhang, and Y. Gao · 2024
Closest in time.
Moka: Open-vocabulary robotic manipulation through mark-based visual prompting
F. Liu, K. Fang, P. Abbeel, and S. Levine · 2024
Closest in time.
Copa: General robotic manipulation through spatial constraints of parts with foundation models, 2024
H. Huang, F. Lin, Y. Hu, S. Wang, and Y. Gao · 2024
Closest in time.
Open6DOR: Benchmarking open-instruction 6-dof object rearrangement and a VLM-based approach
Y. Ding, H. Geng, C. Xu, X. Fang, J. Zhang, S. Wei, Q. Dai, Z. Zhang, and H. Wang · 2024
Closest in time.
Ag2manip: Learning novel manipulation skills with agent-agnostic visual and action representations
P. Li, T. Liu, Y. Li, M. Han, H. Geng, S. Wang, Y. Zhu, S.-C. Zhu, and S. Huang · 2024
Closest in time.
Pivot: Iterative visual prompting elicits actionable knowledge for vlms
S. Nasiriany, F. Xia, W. Yu, T. Xiao, J. Liang, I. Dasgupta, A. Xie, D. Driess, A. Wahid, Z. Xu, et al · 2024
Closest in time.
Mobile aloha: Learning bimanual mobile manipulation with low-cost whole-body teleoperation
Z. Fu, T. Z. Zhao, and C. Finn · 2024
Closest in time.
Roboclip: One demonstration is enough to learn robot policies
S. Sontakke, J. Zhang, S. Arnold, K. Pertsch, E. Bıyık, D. Sadigh, C. Finn, and L. Itti · 2024
Closest in time.
Octo: An open-source generalist robot policy
O. M. Team, D. Ghosh, H. Walke, K. Pertsch, K. Black, O. Mees, S. Dasari, J. Hejna, T. Kreiman, C. Xu, et al · 2024
Closest in time.
Dinobot: Robot manipulation via retrieval and alignment with vision foundation models
N. Di Palo and E. Johns · 2024
Closest in time.
One-shot imitation learning with invariance matching for robotic manipulation
X. Zhang and A. Boularias · 2024
Closest in time.
Zero-shot imitation policy via search in demonstration dataset
F. Malato, F. Leopold, A. Melnik, and V. Hautamäki · 2024
Closest in time.
Y. Ju, K. Hu, G. Zhang, G. Zhang, M. Jiang, and H. Xu · 2024
Closest in time.
A tale of two features: Stable diffusion complements dino for zero-shot semantic correspondence
J. Zhang, C. Herrmann, J. Hur, L. Polania Cabrera, V. Jampani, D. Sun, and M.-H. Yang · 2024
Closest in time.
Probing the 3d awareness of visual foundation models, 2024
M. E. Banani, A. Raj, K.-K. Maninis, A. Kar, Y. Li, M. Rubinstein, D. Sun, L. Guibas, J. Johnson, and V. Jampani · 2024
Closest in time.
Openfmnav: Towards open-set zero-shot object navigation via vision-language foundation models
Y. Kuang, H. Lin, and M. Jiang · 2024
Closest in time.
Grounded sam: Assembling open-world models for diverse visual tasks, 2024
T. Ren, S. Liu, A. Zeng, J. Lin, K. Li, H. Cao, J. Chen, X. Huang, Y. Chen, F. Yan, Z. Zeng, H. Zhang, F. Li, J. Yang, H. Li, Q. Jiang, and L. Zhang · 2024
Closest in time.