Fetching the paper…
Reading the bibliography…
Learning manipulation skills from human demonstration videos offers a promising path toward generalizable and interpretable robotic intelligence-particularly through the lens of actionable affordances.
The theory of affordances
J. Gibson · 1977
Earlier work this paper cites.
Superquadrics for segmenting and modeling range data
A. Leonardis, A. Jaklic, and F. Solina · 1997
Earlier work this paper cites.
Surf: Speeded up robust features
H. Bay, T. Tuytelaars, and L. Van Gool · 2006
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
V-rep: A versatile and scalable robot simulation framework
E. Rohmer, S. P. Singh, and M. Freese · 2013
Earlier work this paper cites.
Color-based skin segmentation: An evaluation of the state of the art
F. Saxen and A. Al-Hamadi · 2014
Earlier work this paper cites.
Affordance detection of tool parts from geometric features
A. Myers, C. L. Teo, C. Fermüller, and Y. Aloimonos · 2015
Earlier work this paper cites.
Attribute based affordance detection from human-object interaction images
M. Hassan and A. Dharmaratne · 2016
Earlier work this paper cites.
The kinetics human action video dataset
W. Kay, J. Carreira, K. Simonyan, B. Zhang, C. Hillier, S. Vijayanarasimhan, F. Viola, T. Green, T. Back, P. Natsev, et al · 2017
Earlier work this paper cites.
Weakly supervised affordance detection
J. Sawatzky, A. Srikantha, and J. Gall · 2017
Earlier work this paper cites.
Object-based affordances detection with convolutional neural networks and dense conditional random fields
A. Nguyen, D. Kanoulas, D. G. Caldwell, and N. G. Tsagarakis · 2017
Earlier work this paper cites.
Focal loss for dense object detection
T. Lin · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
I. Loshchilov and F. Hutter · 2017
Earlier work this paper cites.
Demo2vec: Reasoning object affordances from online videos
K. Fang, T.-L. Wu, D. Yang, S. Savarese, and J. J. Lim · 2018
Earlier work this paper cites.
Learning to act properly: Predicting and explaining affordances from images
C.-Y. Chuang, J. Li, A. Torralba, and S. Fidler · 2018
Earlier work this paper cites.
Superquadrics revisited: Learning 3d shape parsing beyond cuboids
D. Paschalidou, A. O. Ulusoy, and A. Geiger · 2019
Earlier work this paper cites.
Pyrep: Bringing v-rep to deep robot learning
S. James, M. Freese, and A. J. Davison · 2019
Earlier work this paper cites.
Rlbench: The robot learning benchmark & learning environment
S. James, Z. Ma, D. R. Arrojo, and A. J. Davison · 2020
Earlier work this paper cites.
Graspnet-1billion: A large-scale benchmark for general object grasping
H.-S. Fang, C. Wang, M. Gou, and C. Lu · 2020
Earlier work this paper cites.
Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters
J. Rasley, S. Rajbhandari, O. Ruwase, and Y. He · 2020
Earlier work this paper cites.
H. Luo, W. Zhai, J. Zhang, Y. Cao, and D. Tao · 2021
Earlier work this paper cites.
Affordance transfer learning for human-object interaction detection
Z. Hou, B. Yu, Y. Qiao, X. Peng, and D. Tao · 2021
Earlier work this paper cites.
Where2act: From pixels to actions for articulated 3d objects
K. Mo, L. J. Guibas, M. Mukadam, A. Gupta, and S. Tulsiani · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Earlier work this paper cites.
Nerf: Representing scenes as neural radiance fields for view synthesis
B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Barron, R. Ramamoorthi, and R. Ng · 2021
Earlier work this paper cites.
Isaac gym: High performance gpu-based physics simulation for robot learning
V. Makoviychuk, L. Wawrzyniak, Y. Guo, M. Lu, K. Storey, M. Macklin, D. Hoeller, N. Rudin, A. Allshire, A. Handa, et al · 2021
Earlier work this paper cites.
Ego4d: Around the world in 3,000 hours of egocentric video
K. Grauman, A. Westbury, E. Byrne, Z. Chavis, A. Furnari, R. Girdhar, J. Hamburger, H. Jiang, M. Liu, X. Liu, et al · 2022
Earlier work this paper cites.
R3m: A universal visual representation for robot manipulation
S. Nair, A. Rajeswaran, V. Kumar, C. Finn, and A. Gupta · 2022
Cited alongside, same era.
Masked visual pre-training for motor control
T. Xiao, I. Radosavovic, T. Darrell, and J. Malik · 2022
Cited alongside, same era.
Learning affordance grounding from exocentric images
H. Luo, W. Zhai, J. Zhang, Y. Cao, and D. Tao · 2022
Cited alongside, same era.
Learning affordance grounding from exocentric images
H. Luo, W. Zhai, J. Zhang, Y. Cao, and D. Tao · 2022
Cited alongside, same era.
Do as i can, not as i say: Grounding language in robotic affordances
M. Ahn, A. Brohan, N. Brown, Y. Chebotar, O. Cortes, B. David, C. Finn, C. Fu, K. Gopalakrishnan, K. Hausman, et al · 2022
Cited alongside, same era.
Affordances from human videos as a versatile representation for robotics
S. Bahl, R. Mendonca, L. Chen, U. Jain, and D. Pathak · 2023
Later among the works it cites.
Rvt: Robotic view transformer for 3d object manipulation
A. Goyal, J. Xu, Y. Guo, V. Blukis, Y.-W. Chao, and D. Fox · 2023
Later among the works it cites.
Lisa++: An improved baseline for reasoning segmentation with large language model
S. Yang, T. Qu, X. Lai, Z. Tian, B. Peng, S. Liu, and J. Jia · 2023
Later among the works it cites.
Segment anything
A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y. Lo, et al · 2023
Later among the works it cites.
Spawnnet: Learning generalizable visuomotor skills from pre-trained network
X. Lin, J. So, S. Mahalingam, F. Liu, and P. Abbeel · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens-100
D. Damen, H. Doughty, G. M. Farinella, A. Furnari, E. Kazakos, J. Ma, D. Moltisanti, J. Munro, T. Perrett, W. Price, et al · 2022
Cited alongside, same era.
Joint hand motion and interaction hotspots prediction from egocentric videos
S. Liu, S. Tripathi, S. Majumdar, and X. Wang · 2022
Cited alongside, same era.
Perceiver-Actor: A multi-task transformer for robotic manipulation
M. Shridhar, L. Manuelli, and D. Fox · 2022
Cited alongside, same era.
Language embedded radiance fields for zero-shot task-oriented grasping
A. Rashid, S. Sharma, C. M. Kim, J. Kerr, L. Y. Chen, A. Kanazawa, and K. Goldberg · 2023
Cited alongside, same era.
Voxposer: Composable 3d value maps for robotic manipulation with language models
W. Huang, C. Wang, R. Zhang, Y. Li, J. Wu, and L. Fei-Fei · 2023
Cited alongside, same era.
Handal: A dataset of real-world manipulable object categories with pose annotations, affordances, and reconstructions
A. Guo, B. Wen, J. Yuan, J. Tremblay, S. Tyree, J. Smith, and S. Birchfield · 2023
Cited alongside, same era.
Understanding 3d object interaction from a single image
S. Qian and D. F. Fouhey · 2023
Cited alongside, same era.
J. Zhou, T. Ma, K.-Y. Lin, Z. Wang, R. Qiu, and J. Liang · 2024
Later among the works it cites.
Y. Ju, K. Hu, G. Zhang, G. Zhang, M. Jiang, and H. Xu · 2024
Later among the works it cites.
Ram: Retrieval-based affordance transfer for generalizable zero-shot robotic manipulation
Y. Kuang, J. Ye, H. Geng, J. Mao, C. Deng, L. Guibas, H. Wang, and Y. Wang · 2024
Later among the works it cites.
Affordancellm: Grounding affordance from vision language models
S. Qian, W. Chen, M. Bai, X. Zhou, Z. Tu, and L. E. Li · 2024
Later among the works it cites.
Glover: Generalizable open-vocabulary affordance reasoning for task-oriented grasping
T. Ma, Z. Wang, J. Zhou, M. Wang, and J. Liang · 2024
Later among the works it cites.
Gaussiangrasper: 3d language gaussian splatting for open-vocabulary robotic grasping
Y. Zheng, X. Chen, Y. Zheng, S. Gu, R. Yang, B. Jin, P. Li, C. Zhong, Z. Wang, L. Liu, et al · 2024
Later among the works it cites.
Learning generalizable feature fields for mobile manipulation
R.-Z. Qiu, Y. Hu, G. Yang, Y. Song, Y. Fu, J. Ye, J. Mu, R. Yang, N. Atanasov, S. Scherer, et al · 2024
Later among the works it cites.
Learning precise affordances from egocentric videos for robotic manipulation
G. Li, N. Tsagkas, J. Song, R. Mon-Williams, S. Vijayakumar, K. Shao, and L. Sevilla-Lara · 2024
Later among the works it cites.
Manigaussian: Dynamic gaussian splatting for multi-task robotic manipulation
G. Lu, S. Zhang, Z. Wang, C. Liu, J. Lu, and Y. Tang · 2024
Later among the works it cites.
Where2explore: Few-shot affordance learning for unseen novel categories of articulated objects
C. Ning, R. Wu, H. Lu, K. Mo, and H. Dong · 2024
Later among the works it cites.
Learning environment-aware affordance for 3d articulated object manipulation under occlusions
R. Wu, K. Cheng, Y. Zhao, C. Ning, G. Zhan, and H. Dong · 2024
Later among the works it cites.
Rekep: Spatio-temporal reasoning of relational keypoint constraints for robotic manipulation
W. Huang, C. Wang, Y. Li, R. Zhang, and L. Fei-Fei · 2024
Later among the works it cites.
Dinobot: Robot manipulation via retrieval and alignment with vision foundation models
N. Di Palo and E. Johns · 2024
Later among the works it cites.
Ok-robot: What really matters in integrating open-knowledge models for robotics
P. Liu, Y. Orru, C. Paxton, N. M. M. Shafiullah, and L. Pinto · 2024
Later among the works it cites.
Arm-constrained curriculum learning for loco-manipulation of a wheel-legged robot
Z. Wang, Y. Jia, L. Shi, H. Wang, H. Zhao, X. Li, J. Zhou, J. Ma, and G. Zhou · 2024
Later among the works it cites.
Openvla: An open-source vision-language-action model
M. J. Kim, K. Pertsch, S. Karamcheti, T. Xiao, A. Balakrishna, S. Nair, R. Rafailov, E. Foster, G. Lam, P. Sanketi, et al · 2024
Later among the works it cites.
Octo: An open-source generalist robot policy
O. M. Team, D. Ghosh, H. Walke, K. Pertsch, K. Black, O. Mees, S. Dasari, J. Hejna, T. Kreiman, C. Xu, et al · 2024
Later among the works it cites.
Lisa: Reasoning segmentation via large language model
X. Lai, Z. Tian, Y. Chen, Y. Li, Y. Yuan, S. Liu, and J. Jia · 2024
Later among the works it cites.
Foundationpose: Unified 6d pose estimation and tracking of novel objects
B. Wen, W. Yang, J. Kautz, and S. Birchfield · 2024
Later among the works it cites.
Rvt-2: Learning precise manipulation from few demonstrations
A. Goyal, V. Blukis, J. Xu, Y. Guo, Y.-W. Chao, and D. Fox · 2024
Later among the works it cites.
Contrastive imitation learning for language-guided multi-task robotic manipulation
T. Ma, J. Zhou, Z. Wang, R. Qiu, and J. Liang · 2024
Later among the works it cites.
S. Bai, K. Chen, X. Liu, J. Wang, W. Ge, S. Song, K. Dang, P. Wang, S. Wang, J. Tang, et al · 2025
Closest in time.