Fetching the paper…
Reading the bibliography…
We address the challenge of acquiring real-world manipulation skills with a scalable framework.
The theory of affordances
J. J. Gibson · 1977
Earlier work this paper cites.
Method for registration of 3-d shapes
P. J. Besl and N. D. McKay · 1992
Earlier work this paper cites.
Auto-encoding variational bayes
D. P. Kingma and M. Welling · 2013
Earlier work this paper cites.
beta-vae: Learning basic visual concepts with a constrained variational framework
I. Higgins, L. Matthey, A. Pal, C. Burgess, X. Glorot, M. Botvinick, S. Mohamed, and A. Lerchner · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Pointnet++: Deep hierarchical feature learning on point sets in a metric space
C. R. Qi, L. Yi, H. Su, and L. J. Guibas · 2017
Earlier work this paper cites.
First-person hand action benchmark with rgb-d videos and 3d hand pose annotations
G. Garcia-Hernando, S. Yuan, S. Baek, and T.-K. Kim · 2018
Earlier work this paper cites.
Scaling egocentric vision: The epic-kitchens dataset
D. Damen, H. Doughty, G. M. Farinella, S. Fidler, A. Furnari, E. Kazakos, D. Moltisanti, J. Munro, T. Perrett, W. Price, and M. Wray · 2018
Earlier work this paper cites.
Motion perception in reinforcement learning with dynamic objects
A. Amiranashvili, A. Dosovitskiy, V. Koltun, and T. Brox · 2018
Earlier work this paper cites.
Partnet: A large-scale benchmark for fine-grained and hierarchical part-level 3d object understanding
K. Mo, S. Zhu, A. X. Chang, L. Yi, S. Tripathi, L. J. Guibas, and H. Su · 2019
Earlier work this paper cites.
Pytorch image models
R. Wightman · 2019
Earlier work this paper cites.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Earlier work this paper cites.
Understanding human hands in contact at internet scale
D. Shan, J. Geng, M. Shu, and D. F. Fouhey · 2020
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al · 2020
Earlier work this paper cites.
Mt-opt: Continuous multi-task robotic reinforcement learning at scale
D. Kalashnikov, J. Varley, Y. Chebotar, B. Swanson, R. Jonschkowski, C. Finn, S. Levine, and K. Hausman · 2021
Earlier work this paper cites.
Vat-mart: Learning visual action trajectory proposals for manipulating 3d articulated objects
R. Wu, Y. Zhao, K. Mo, Z. Guo, Y. Wang, T. Wu, Q. Fan, X. Chen, L. Guibas, and H. Dong · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Earlier work this paper cites.
Screwnet: Category-independent articulation model estimation from depth images using screw theory
A. Jain, R. Lioutikov, C. Chuck, and S. Niekum · 2021
Earlier work this paper cites.
Tactile-rl for insertion: Generalization to objects of unknown geometry
S. Dong, D. K. Jha, D. Romeres, S. Kim, D. Nikovski, and A. Rodriguez · 2021
Earlier work this paper cites.
Emergent abilities of large language models
J. Wei, Y. Tay, R. Bommasani, C. Raffel, B. Zoph, S. Borgeaud, D. Yogatama, M. Bosma, D. Zhou, D. Metzler, et al · 2022
Earlier work this paper cites.
Flowbot3d: Learning 3d articulation flow to manipulate articulated objects
B. Eisner, H. Zhang, and D. Held · 2022
Earlier work this paper cites.
Hoi4d: A 4d egocentric dataset for category-level human-object interaction
Y. Liu, Y. Liu, C. Jiang, K. Lyu, W. Wan, H. Shen, B. Liang, Z. Fu, H. Wang, and L. Yi · 2022
Earlier work this paper cites.
Pointnext: Revisiting pointnet++ with improved training and scaling strategies
G. Qian, Y. Li, H. Peng, J. Mai, H. Hammoud, M. Elhoseiny, and B. Ghanem · 2022
Earlier work this paper cites.
R3m: A universal visual representation for robot manipulation
S. Nair, A. Rajeswaran, V. Kumar, C. Finn, and A. Gupta · 2022
Earlier work this paper cites.
Joint hand motion and interaction hotspots prediction from egocentric videos
S. Liu, S. Tripathi, S. Majumdar, and X. Wang · 2022
Earlier work this paper cites.
Viola: Imitation learning for vision-based manipulation with object proposal priors
Y. Zhu, A. Joshi, P. Stone, and Y. Zhu · 2022
Earlier work this paper cites.
Egocentric prediction of action target in 3d
Y. Li, Z. Cao, A. Liang, B. Liang, L. Chen, H. Zhao, and C. Feng · 2022
Earlier work this paper cites.
Vip: Towards universal visual reward and representation via value-implicit pre-training
Y. J. Ma, S. Sodhani, D. Jayaraman, O. Bastani, V. Kumar, and A. Zhang · 2022
Earlier work this paper cites.
Masked visual pre-training for motor control
T. Xiao, I. Radosavovic, T. Darrell, and J. Malik · 2022
Earlier work this paper cites.
On pre-training for visuo-motor control: Revisiting a learning-from-scratch baseline
N. Hansen, Z. Yuan, Y. Ze, T. Mu, A. Rajeswaran, H. Su, H. Xu, and X. Wang · 2022
Cited alongside, same era.
O2o-afford: Annotation-free large-scale object-object affordance learning
K. Mo, Y. Qin, F. Xiang, H. Su, and L. Guibas · 2022
Cited alongside, same era.
Adaafford: Learning to adapt manipulation affordance for 3d articulated objects via few-shot interactions
Y. Wang, R. Wu, K. Mo, J. Ke, Q. Fan, L. J. Guibas, and H. Dong · 2022
Cited alongside, same era.
Human-to-robot imitation in the wild
S. Bahl, A. Gupta, and D. Pathak · 2022
Cited alongside, same era.
Fabricflownet: Bimanual cloth manipulation with a flow-based policy
T. Weng, S. M. Bajracharya, Y. Wang, K. Agrawal, and D. Held · 2022
Cited alongside, same era.
Can pre-trained text-to-image models generate visual goals for reinforcement learning?
J. Gao, K. Hu, G. Xu, and H. Xu · 2023
Later among the works it cites.
Video prediction models as rewards for reinforcement learning
A. Escontrela, A. Adeniji, W. Yan, A. Jain, X. B. Peng, K. Goldberg, Y. Lee, D. Hafner, and P. Abbeel · 2023
Later among the works it cites.
Y. Du, M. Yang, P. Florence, F. Xia, A. Wahid, B. Ichter, P. Sermanet, T. Yu, P. Abbeel, J. B. Tenenbaum, et al · 2023
Later among the works it cites.
Towards generalizable zero-shot manipulation via translating human interaction plans
H. Bharadhwaj, A. Gupta, V. Kumar, and S. Tulsiani · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Acid: Action-conditional implicit visual dynamics for deformable object manipulation
B. Shen, Z. Jiang, C. Choy, L. J. Guibas, S. Savarese, A. Anandkumar, and Y. Zhu · 2022
Cited alongside, same era.
Fine-grained egocentric hand-object segmentation: Dataset, model, and applications
L. Zhang, S. Zhou, S. Stent, and J. Shi · 2022
Cited alongside, same era.
Inner monologue: Embodied reasoning through planning with language models
W. Huang, F. Xia, T. Xiao, H. Chan, J. Liang, P. Florence, A. Zeng, J. Tompson, I. Mordatch, Y. Chebotar, et al · 2022
Cited alongside, same era.
Rt-1: Robotics transformer for real-world control at scale
A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, J. Hsu, et al · 2022
Cited alongside, same era.
Ego4d: Around the world in 3,000 hours of egocentric video
K. Grauman, A. Westbury, E. Byrne, Z. Chavis, A. Furnari, R. Girdhar, J. Hamburger, H. Jiang, M. Liu, X. Liu, et al · 2022
Cited alongside, same era.
Open x-embodiment: Robotic learning datasets and rt-x models
A. Padalkar, A. Pooley, A. Jain, A. Bewley, A. Herzog, A. Irpan, A. Khazatsky, A. Rai, A. Singh, A. Brohan, et al · 2023
Cited alongside, same era.
Toolflownet: Robotic manipulation with tools via predicting tool flow from point clouds
D. Seita, Y. Wang, S. J. Shetty, E. Y. Li, Z. Erickson, and D. Held · 2023
Cited alongside, same era.
H. Geng, H. Xu, C. Zhao, C. Xu, L. Yi, S. Huang, and H. Wang · 2023
Later among the works it cites.
Structured world models from human videos
R. Mendonca, S. Bahl, and D. Pathak · 2023
Later among the works it cites.
Zero-shot robot manipulation from passive human videos
H. Bharadhwaj, A. Gupta, S. Tulsiani, and V. Kumar · 2023
Later among the works it cites.
Robotap: Tracking arbitrary points for few-shot visual imitation
M. Vecerik, C. Doersch, Y. Yang, T. Davchev, Y. Aytar, G. Zhou, R. Hadsell, L. Agapito, and J. Scholz · 2023
Later among the works it cites.
Any-point trajectory modeling for policy learning
C. Wen, X. Lin, J. So, K. Chen, Q. Dou, Y. Gao, and P. Abbeel · 2023
Later among the works it cites.
Grounding dino: Marrying dino with grounded pre-training for open-set object detection
S. Liu, Z. Zeng, T. Ren, F. Li, H. Zhang, J. Yang, C. Li, J. Yang, H. Su, J. Zhu, et al · 2023
Later among the works it cites.
Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives
K. Grauman, A. Westbury, L. Torresani, K. Kitani, J. Malik, T. Afouras, K. Ashutosh, V. Baiyya, S. Bansal, B. Boote, et al · 2023
Later among the works it cites.
Do as i can, not as i say: Grounding language in robotic affordances
A. Brohan, Y. Chebotar, C. Finn, K. Hausman, A. Herzog, D. Ho, J. Ibarz, A. Irpan, E. Jang, R. Julian, et al · 2023
Later among the works it cites.
Scaling up and distilling down: Language-guided robot skill acquisition
H. Ha, P. Florence, and S. Song · 2023
Later among the works it cites.
Code as policies: Language model programs for embodied control
J. Liang, W. Huang, F. Xia, P. Xu, K. Hausman, B. Ichter, P. Florence, and A. Zeng · 2023
Later among the works it cites.
Palm-e: An embodied multimodal language model
D. Driess, F. Xia, M. S. Sajjadi, C. Lynch, A. Chowdhery, B. Ichter, A. Wahid, J. Tompson, Q. Vuong, T. Yu, et al · 2023
Later among the works it cites.
Look before you leap: Unveiling the power of gpt-4v in robotic vision-language planning
Y. Hu, F. Lin, T. Zhang, L. Yi, and Y. Gao · 2023
Later among the works it cites.
Voxposer: Composable 3d value maps for robotic manipulation with language models
W. Huang, C. Wang, R. Zhang, Y. Li, J. Wu, and L. Fei-Fei · 2023
Later among the works it cites.
Rt-2: Vision-language-action models transfer web knowledge to robotic control
A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, X. Chen, K. Choromanski, T. Ding, D. Driess, A. Dubey, C. Finn, et al · 2023
Later among the works it cites.
H. Bharadhwaj, J. Vakil, M. Sharma, A. Gupta, S. Tulsiani, and V. Kumar · 2023
Later among the works it cites.
Octo: An open-source generalist robot policy
Octo Model Team, D. Ghosh, H. Walke, K. Pertsch, K. Black, O. Mees, S. Dasari, J. Hejna, C. Xu, J. Luo, T. Kreiman, Y. Tan, D. Sadigh, C. Finn, and S. Levine · 2023
Later among the works it cites.
Zoedepth: Zero-shot transfer by combining relative and metric depth
S. F. Bhat, R. Birkl, D. Wofk, P. Wonka, and M. Müller · 2023
Later among the works it cites.
Diffusion policy: Visuomotor policy learning via action diffusion
C. Chi, S. Feng, Y. Du, Z. Xu, E. Cousineau, B. Burchfiel, and S. Song · 2023
Later among the works it cites.
Perceiver-actor: A multi-task transformer for robotic manipulation
M. Shridhar, L. Manuelli, and D. Fox · 2023
Later among the works it cites.
3d implicit transporter for temporally consistent keypoint discovery
C. Zhong, Y. Zheng, Y. Zheng, H. Zhao, L. Yi, X. Mu, L. Wang, P. Li, G. Zhou, C. Yang, et al · 2023
Later among the works it cites.
Moka: Open-vocabulary robotic manipulation through mark-based visual prompting
F. Liu, K. Fang, P. Abbeel, and S. Levine · 2024
Closest in time.
Track2act: Predicting point tracks from internet videos enables diverse zero-shot robot manipulation
H. Bharadhwaj, R. Mottaghi, A. Gupta, and S. Tulsiani · 2024
Closest in time.
Droid: A large-scale in-the-wild robot manipulation dataset
A. Khazatsky, K. Pertsch, S. Nair, A. Balakrishna, S. Dasari, S. Karamcheti, S. Nasiriany, M. K. Srirama, L. Y. Chen, K. Ellis, et al · 2024
Closest in time.
S. Huang, I. Ponomarenko, Z. Jiang, X. Li, X. Hu, P. Gao, H. Li, and H. Dong · 2024
Closest in time.