Fetching the paper…
Reading the bibliography…
Affordance, defined as the potential actions that an object offers, is crucial for embodied AI agents.
Gibson, J.J.: The ecological approach to visual perception. (1979)
1979
Earlier work this paper cites.
Communications of the ACM 24
Fischler, M.A., Bolles, R.C.: Random sample consensus: a paradigm for model fitting with applications to image analysis and automated cartography · 1981
Earlier work this paper cites.
Foundations and Trends® in Computer Graphics and Vision 2
Szeliski, R., et al.: Image alignment and stitching: A tutorial · 2007
Earlier work this paper cites.
Computer vision and image understanding 110
Bay, H., Ess, A., Tuytelaars, T., Van Gool, L.: Speeded-up robust features (surf) · 2008
Earlier work this paper cites.
In: CVPR (2009)
Deng, J., Dong, W., Socher, R., Li, L.J., Li, K., Fei-Fei, L.: Imagenet: A large-scale hierarchical image database · 2009
Earlier work this paper cites.
The International journal of robotics research 32
Koppula, H.S., Gupta, R., Saxena, A.: Learning human activities and object affordances from rgb-d videos · 2013
Earlier work this paper cites.
ICRA (2015)
Myers, A., Teo, C.L., Fermüller, C., Aloimonos, Y.: Affordance detection of tool parts from geometric features · 2015
Earlier work this paper cites.
In: CVPR (2016)
He, K., Zhang, X., Ren, S., Sun, J.: Deep residual learning for image recognition · 2016
Earlier work this paper cites.
In: 2016 fourth international conference on 3D vision (3DV), pp. 565–571. Ieee (2016)
Milletari, F., Navab, N., Ahmadi, S.A.: V-net: Fully convolutional neural networks for volumetric medical image segmentation · 2016
Earlier work this paper cites.
In: 2016 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pp. 2765–2770. IEEE (2016)
Nguyen, A., Kanoulas, D., Caldwell, D.G., Tsagarakis, N.G.: Detecting object affordances with convolutional neural networks · 2016
Earlier work this paper cites.
arXiv preprint arXiv:1712.07576 (2017)
Chuang, C.Y., Li, J., Torralba, A., Fidler, S.: Learning to act properly: Predicting and explaining affordances from images · 2017
Earlier work this paper cites.
In: Proceedings of the IEEE international conference on computer vision, pp. 2980–2988 (2017)
Lin, T.Y., Goyal, P., Girshick, R., He, K., Dollár, P.: Focal loss for dense object detection · 2017
Earlier work this paper cites.
arXiv preprint arXiv:1711.05101 (2017)
Loshchilov, I., Hutter, F.: Decoupled weight decay regularization · 2017
Earlier work this paper cites.
In: IROS (2017)
Nguyen, A., Kanoulas, D., Caldwell, D.G., Tsagarakis, N.G.: Object-based affordances detection with convolutional neural networks and dense conditional random fields · 2017
Earlier work this paper cites.
In: ICCVW (2017)
Sawatzky, J., Gall, J.: Adaptive binarization for weakly supervised affordance segmentation · 2017
Earlier work this paper cites.
CVPR (2017)
Sawatzky, J., Srikantha, A., Gall, J.: Weakly supervised affordance detection · 2017
Earlier work this paper cites.
In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 2881–2890 (2017)
Zhao, H., Shi, J., Qi, X., Wang, X., Jia, J.: Pyramid scene parsing network · 2017
Earlier work this paper cites.
In: Proceedings of the European conference on computer vision (ECCV), pp. 801–818 (2018)
Chen, L.C., Zhu, Y., Papandreou, G., Schroff, F., Adam, H.: Encoder-decoder with atrous separable convolution for semantic image segmentation · 2018
Earlier work this paper cites.
In: CVPR (2018)
Chuang, C.Y., Li, J., Torralba, A., Fidler, S.: Learning to act properly: Predicting and explaining affordances from images · 2018
Earlier work this paper cites.
In: European Conference on Computer Vision (ECCV) (2018)
Damen, D., Doughty, H., Farinella, G.M., Fidler, S., Furnari, A., Kazakos, E., Moltisanti, D., Munro, J., Perrett, T., Price, W., Wray, M.: Scaling egocentric vision: The epic-kitchens dataset · 2018
Earlier work this paper cites.
ICRA (2018)
Do, T.T., Nguyen, A., Reid, I.: Affordancenet: An end-to-end deep learning approach for object affordance detection · 2018
Earlier work this paper cites.
CVPR (2018)
Fang, K., Wu, T.L., Yang, D., Savarese, S., Lim, J.J.: Demo2Vec: Reasoning Object Affordances from Online Videos · 2018
Earlier work this paper cites.
In: Proceedings of the European Conference on Computer Vision (ECCV) (2018)
Li, Y., Liu, M., Rehg, J.M.: In the eye of beholder: Joint learning of gaze and actions in first person video · 2018
Earlier work this paper cites.
In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 8688–8697 (2019)
Nagarajan, T., Feichtenhofer, C., Grauman, K.: Grounded human-object interaction hotspots from video · 2019
Earlier work this paper cites.
In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 9226–9235 (2019)
Oh, S.W., Lee, J.Y., Xu, N., Kim, S.J.: Video object segmentation using space-time memory networks · 2019
Earlier work this paper cites.
IEEE Transactions on Pattern Analysis & Machine Intelligence (01), 1–1 (2020)
Damen, D., Doughty, H., Farinella, G., Fidler, S., Furnari, A., Kazakos, E., Moltisanti, D., Munro, J., Perrett, T., Price, W., et al.: The epic-kitchens dataset: Collection, challenges and baselines · 2020
Earlier work this paper cites.
In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 11,444–11,453 (2020)
Fang, H.S., Wang, C., Gou, M., Lu, C.: Graspnet-1billion: A large-scale benchmark for general object grasping · 2020
Earlier work this paper cites.
In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 9869–9878 (2020)
Shan, D., Geng, J., Shu, M., Fouhey, D.F.: Understanding human hands in contact at internet scale · 2020
Earlier work this paper cites.
In: ICCV (2021)
Caron, M., Touvron, H., Misra, I., Jégou, H., Mairal, J., Bojanowski, P., Joulin, A.: Emerging properties in self-supervised vision transformers · 2021
Cited alongside, same era.
In: CVPR (2021)
Deng, S., Xu, X., Wu, C., Chen, K., Jia, K.: 3d affordancenet: A benchmark for visual object affordance understanding · 2021
Cited alongside, same era.
ICLR (2021)
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., Houlsby, N.: An image is worth 16x16 words: Transformers for image recognition at scale · 2021
Cited alongside, same era.
arXiv preprint arXiv:2110.07058 (2021)
Grauman, K., Westbury, A., Byrne, E., Chavis, Z., Furnari, A., Girdhar, R., Hamburger, J., Jiang, H., Liu, M., Liu, X., et al.: Ego4d: Around the world in 3,000 hours of egocentric video · 2021
Cited alongside, same era.
arXiv preprint arXiv:2106.09685 (2021)
Hu, E.J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., Chen, W.: Lora: Low-rank adaptation of large language models · 2021
In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 10,922–10,931 (2023)
Li, G., Jampani, V., Sun, D., Sevilla-Lara, L.: Locate: Localize and transfer object parts for weakly supervised affordance grounding · 2023
Later among the works it cites.
arXiv preprint arXiv:2303.05499 (2023)
Liu, S., Zeng, Z., Ren, T., Li, F., Zhang, H., Yang, J., Li, C., Yang, J., Su, H., Zhu, J., et al.: Grounding dino: Marrying dino with grounded pre-training for open-set object detection · 2023
Later among the works it cites.
arXiv preprint arXiv:2303.00871 (2023)
Mur-Labadia, L., Martinez-Cantin, R., Guerrero, J.J.: Bayesian deep learning for affordance segmentation in images · 2023
Later among the works it cites.
In: 2023 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pp. 5692–5698. IEEE (2023)
Nguyen, T., Vu, M.N., Vuong, A., Nguyen, D., Vo, T., Le, N., Nguyen, A.: Open-vocabulary affordance detection in 3d point clouds · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
In: International conference on machine learning, pp. 8748–8763. PMLR (2021)
Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al.: Learning transferable visual models from natural language supervision · 2021
Cited alongside, same era.
ICCV (2021)
Ranftl, R., Bochkovskiy, A., Koltun, V.: Vision transformers for dense prediction · 2021
Cited alongside, same era.
2021 IEEE International Conference on Robotics and Automation (ICRA) (2021)
Sundermeyer, M., Mousavian, A., Triebel, R., Fox, D.: Contact-graspnet: Efficient 6-dof grasp generation in cluttered scenes · 2021
Cited alongside, same era.
IEEE Access 9
Thermos, S., Potamianos, G., Daras, P.: Joint object affordance reasoning and segmentation in rgb-d videos · 2021
Cited alongside, same era.
Advances in Neural Information Processing Systems 34
Xie, E., Wang, W., Yu, Z., Anandkumar, A., Alvarez, J.M., Luo, P.: Segformer: Simple and efficient design for semantic segmentation with transformers · 2021
Cited alongside, same era.
ECCVW What is Motion For (2022)
Amir, S., Gandelsman, Y., Bagon, S., Dekel, T.: Deep vit features as dense visual descriptors · 2022
Cited alongside, same era.
arXiv preprint arXiv:2205.08534 (2022)
Chen, Z., Duan, Y., Wang, W., He, J., Lu, T., Dai, J., Qiao, Y.: Vision transformer adapter for dense predictions · 2022
Cited alongside, same era.
URL https://cdn.openai.com/papers/GPTV%20System%20Card.pdf
OpenAI: Gpt-4v(ision) system card (2023) · 2023
Later among the works it cites.
arXiv preprint arXiv:2304.07193 (2023)
Oquab, M., Darcet, T., Moutakanni, T., Vo, H., Szafraniec, M., Khalidov, V., Fernandez, P., Haziza, D., Massa, F., El-Nouby, A., et al.: Dinov2: Learning robust visual features without supervision · 2023
Later among the works it cites.
In: 7th Annual Conference on Robot Learning (2023)
Rashid, A., Sharma, S., Kim, C.M., Kerr, J., Chen, L.Y., Kanazawa, A., Goldberg, K.: Language embedded radiance fields for zero-shot task-oriented grasping · 2023
Later among the works it cites.
In: CoRL (2023)
Shen, W., Yang, G., Yu, A., Wong, J., Kaelbling, L.P., Isola, P.: Distilled feature fields enable few-shot language-guided manipulation · 2023
Later among the works it cites.
arXiv preprint arXiv:2305.11173 (2023)
Sun, P., Chen, S., Zhu, C., Xiao, F., Luo, P., Xie, S., Yan, Z.: Going denser with open-vocabulary part segmentation · 2023
Later among the works it cites.
Wang, Y., Li, Z., Zhang, M., Driggs-Campbell, K., Wu, J., Fei-Fei, L., Li, Y.: D 3 fields: Dynamic 3d descriptor fields for zero-shot generalizable robotic manipulation · 2023
Later among the works it cites.
arXiv preprint arXiv:2310.05107 (2023)
Wei, M., Yue, X., Zhang, W., Kong, S., Liu, X., Pang, J.: Ov-parts: Towards open-vocabulary part segmentation · 2023
Later among the works it cites.
Zhou, Z., Lei, Y., Zhang, B., Liu, L., Liu, Y.: Zegclip: Towards adapting clip for zero-shot semantic segmentation · 2023
Later among the works it cites.
arXiv preprint arXiv:2402.13181 (2024)
Di Palo, N., Johns, E.: Dinobot: Robot manipulation via retrieval and alignment with vision foundation models · 2024
Closest in time.
IEEE Robotics and Automation Letters (2024)
Di Palo, N., Johns, E.: On the effectiveness of retrieval, alignment, and replay in manipulation · 2024
Closest in time.
arXiv preprint arXiv:2403.08248 (2024)
Huang, H., Lin, F., Hu, Y., Wang, S., Gao, Y.: Copa: General robotic manipulation through spatial constraints of parts with foundation models · 2024
Closest in time.
arXiv preprint arXiv:2401.07487 (2024)
Ju, Y., Hu, K., Zhang, G., Zhang, G., Jiang, M., Xu, H.: Robo-abc: Affordance generalization beyond categories via semantic correspondence for robot manipulation · 2024
Closest in time.
In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2024)
Ke, B., Obukhov, A., Huang, S., Metzger, N., Daudt, R.C., Schindler, K.: Repurposing diffusion-based image generators for monocular depth estimation · 2024
Closest in time.
In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 3086–3096 (2024)
Li, G., Sun, D., Sevilla-Lara, L., Jampani, V.: One-shot open affordance learning with foundation models · 2024
Closest in time.
arXiv preprint arXiv:2406.05951 (2024)
van Oort, T., Miller, D., Browne, W.N., Marticorena, N., Haviland, J., Suenderhauf, N.: Open-vocabulary part-based grasping · 2024
Closest in time.
In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, pp. 7587–7597 (2024)
Qian, S., Chen, W., Bai, M., Zhou, X., Tu, Z., Li, L.E.: Affordancellm: Grounding affordance from vision language models · 2024
Closest in time.
arXiv preprint arXiv:2401.14159 (2024)
Ren, T., Liu, S., Zeng, A., Lin, J., Li, K., Cao, H., Chen, J., Huang, X., Chen, Y., Yan, F., et al.: Grounded sam: Assembling open-world models for diverse visual tasks · 2024
Closest in time.
arXiv preprint arXiv:2404.11000 (2024)
Tong, E., Opipari, A., Lewis, S., Zeng, Z., Jenkins, O.C.: Oval-prompt: Open-vocabulary affordance localization for robot manipulation through llm affordance-grounding · 2024
Closest in time.
In: IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) (2024)
Tsagkas, N., Rome, J., Ramamoorthy, S., Mac Aodha, O., Lu, C.X.: Click to grasp: Zero-shot precise manipulation via visual diffusion descriptors · 2024
Closest in time.
In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 16,111–16,121 (2024)
Xiong, Y., Varadarajan, B., Wu, L., Xiang, X., Xiao, F., Zhu, C., Dai, X., Wang, D., Sun, F., Iandola, F., et al.: Efficientsam: Leveraged masked image pretraining for efficient segment anything · 2024
Closest in time.
In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 10,371–10,381 (2024)
Yang, L., Kang, B., Huang, Z., Xu, X., Feng, J., Zhao, H.: Depth anything: Unleashing the power of large-scale unlabeled data · 2024
Closest in time.
arXiv preprint arXiv:2404.02523 (2024)
Yoshida, T., Kurita, S., Nishimura, T., Mori, S.: Text-driven affordance learning from egocentric vision · 2024
Closest in time.
In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 445–456 (2024)
Zhan, X., Yang, L., Zhao, Y., Mao, K., Xu, H., Lin, Z., Li, K., Lu, C.: Oakink2: A dataset of bimanual hands-object manipulation in complex task completion · 2024
Closest in time.