Fetching the paper…
Reading the bibliography…
Affordance grounding aims to locate objects' "action possibilities" regions, which is an essential step toward embodied intelligence.
Gibson JJ (1977) The theory of affordances. Hilldale
1977
Earlier work this paper cites.
Swain MJ, Ballard DH (1991) Color indexing. International Journal of Computer Vision (IJCV) 7(1):11–32
1991
Earlier work this paper cites.
Lee DD, Seung HS (2000) Algorithms for non-negative matrix factorization. In: NIPS
2000
Earlier work this paper cites.
Rizzolatti G, Craighero L (2004) The mirror-neuron system. Annu Rev Neurosci 27:169–192
2004
Earlier work this paper cites.
Peters RJ, Iyer A, Itti L, Koch C (2005) Components of bottom-up gaze allocation in natural images. Vision research 45(18):2397–2416
2005
Earlier work this paper cites.
Stark M, Lies P, Zillich M, Wyatt J, Schiele B (2008) Functional object class detection based on learned affordance cues. In: International conference on computer vision systems, Springer, pp 435–444
2008
Earlier work this paper cites.
Kolda TG, Bader BW (2009) Tensor decompositions and applications. SIAM review 51(3):455–500
2009
Earlier work this paper cites.
Grabner H, Gall J, Van Gool L (2011) What makes a chair a chair? In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), IEEE, pp 1529–1536
2011
Earlier work this paper cites.
Kjellström H, Romero J, Kragić D (2011) Visual object-action recognition: Inferring object affordances from human demonstration. Computer Vision and Image Understanding 115(1):81–90
2011
Earlier work this paper cites.
Judd T, Durand F, Torralba A (2012) A benchmark of computational models of saliency to predict human fixations
2012
Earlier work this paper cites.
Soomro K, Zamir AR, Shah M (2012) Ucf101: A dataset of 101 human actions classes from videos in the wild. arXiv preprint arXiv:12120402
2012
Earlier work this paper cites.
Koppula HS, Gupta R, Saxena A (2013) Learning human activities and object affordances from rgb-d videos. The International Journal of Robotics Research 32(8):951–970
2013
Earlier work this paper cites.
Koppula HS, Saxena A (2014) Physically grounded spatio-temporal object affordances. In: European Conference on Computer Vision (ECCV), Springer, pp 831–847
2014
Earlier work this paper cites.
Lin TY, Maire M, Belongie S, Hays J, Perona P, Ramanan D, Dollár P, Zitnick CL (2014) Microsoft coco: Common objects in context. In: Proceedings of the European Conference on Computer Vision (ECCV), Springer, pp 740–755
2014
Earlier work this paper cites.
Simonyan K, Zisserman A (2014) Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:14091556
2014
Earlier work this paper cites.
Soran B, Farhadi A, Shapiro L (2014) Action recognition in the presence of one egocentric and multiple static cameras. In: Asian Conference on Computer Vision, Springer, pp 178–193
2014
Earlier work this paper cites.
Bylinskii Z, Judd T, Borji A, Itti L, Durand F, Oliva A, Torralba A (2015) Mit saliency benchmark
2015
Earlier work this paper cites.
Fouhey DF, Wang X, Gupta A (2015) In defense of the direct perception of affordances. arXiv preprint arXiv:150501085
2015
Earlier work this paper cites.
Hinton G, Vinyals O, Dean J (2015) Distilling the knowledge in a neural network. arXiv preprint arXiv:150302531
2015
Earlier work this paper cites.
Myers A, Teo CL, Fermüller C, Aloimonos Y (2015) Affordance detection of tool parts from geometric features. In: 2015 IEEE International Conference on Robotics and Automation (ICRA), IEEE, pp 1374–1381
2015
Earlier work this paper cites.
Nagarajan T, Grauman K (2020) Learning affordance landscapes for interaction exploration in 3d environments. Advances in Neural Information Processing Systems 33:2005–2015
2015
Earlier work this paper cites.
He K, Zhang X, Ren S, Sun J (2016) Deep residual learning for image recognition. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp 770–778
2016
Earlier work this paper cites.
Kümmerer M, Wallis TS, Bethge M (2016) Deepgaze ii: Reading fixations from deep features trained on object recognition. arXiv preprint arXiv:161001563
2016
Earlier work this paper cites.
Lau M, Dev K, Shi W, Dorsey J, Rushmeier H (2016) Tactile mesh saliency. ACM Transactions on Graphics (TOG) 35(4):1–11
2016
Earlier work this paper cites.
Nguyen A, Kanoulas D, Caldwell DG, Tsagarakis NG (2016) Detecting object affordances with convolutional neural networks. In: 2016 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), IEEE, pp 2765–2770
2016
Cited alongside, same era.
Srikantha A, Gall J (2016) Weakly supervised learning of affordances. arXiv preprint arXiv:160502964
2016
Cited alongside, same era.
Zhou B, Khosla A, Lapedriza A, Oliva A, Torralba A (2016) Learning deep features for discriminative localization. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp 2921–2929
2016
Cited alongside, same era.
Bohg J, Hausman K, Sankaran B, Brock O, Kragic D, Schaal S, Sukhatme GS (2017) Interactive perception: Leveraging action in perception and perception in action. IEEE Transactions on Robotics 33(6):1273–1291
2017
Cited alongside, same era.
Mi J, Liang H, Katsakis N, Tang S, Li Q, Zhang C, Zhang J (2020) Intention-related natural language grounding via object affordance detection and intention semantic extraction. Frontiers in Neurorobotics p 26
2020
Later among the works it cites.
Zhang J, Tao D (2020) Empowering things with intelligence: a survey of the progress, challenges, and opportunities in artificial intelligence of things. IEEE Internet of Things Journal 8(10):7789–7817
2020
Later among the works it cites.
Zhao X, Cao Y, Kang Y (2020) Object affordance detection with relationship-aware network. Neural Computing and Applications 32(18):14321–14333
2020
Later among the works it cites.
Deng S, Xu X, Wu C, Chen K, Jia K (2021) 3d affordancenet: A benchmark for visual object affordance understanding. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp 1778–1787
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Fan C, Lee J, Xu M, Singh KK, Yong JL (2017) Identifying first-person camera wearers in third-person videos. In: 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
2017
Cited alongside, same era.
Lakani SR, Rodríguez-Sánchez AJ, Piater J (2017) Can affordances guide object decomposition into semantically meaningful parts? In: 2017 IEEE Winter Conference on Applications of Computer Vision (WACV), IEEE, pp 82–90
2017
Cited alongside, same era.
Nguyen A, Kanoulas D, Caldwell DG, Tsagarakis NG (2017) Object-based affordances detection with convolutional neural networks and dense conditional random fields. In: 2017 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), IEEE, pp 5908–5915
2017
Cited alongside, same era.
Sawatzky J, Gall J (2017) Adaptive binarization for weakly supervised affordance segmentation. In: Proceedings of the IEEE International Conference on Computer Vision Workshops, pp 1383–1391
2017
Cited alongside, same era.
Sawatzky J, Srikantha A, Gall J (2017) Weakly supervised affordance detection. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
2017
Cited alongside, same era.
Tang Y, Tian Y, Lu J, Feng J, Zhou J (2017) Action recognition in rgb-d egocentric videos. In: 2017 IEEE International Conference on Image Processing (ICIP), IEEE, pp 3410–3414
2017
Cited alongside, same era.
Bylinskii Z, Judd T, Oliva A, Torralba A, Durand F (2018) What do different evaluation metrics tell us about saliency models? IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI) 41(3):740–757
2018
Cited alongside, same era.
Chao YW, Liu Y, Liu X, Zeng H, Deng J (2018) Learning to detect human-object interactions. In: 2018 IEEE Winter Conference on Applications of Computer Vision (WACV), IEEE, pp 381–389
2018
Cited alongside, same era.
Gao W, Wan F, Pan X, Peng Z, Tian Q, Han Z, Zhou B, Ye Q (2021) Ts-cam: Token semantic coupled attention map for weakly supervised object localization. In: Proceedings of the IEEE International Conference on Computer Vision (ICCV), pp 2886–2895
2021
Later among the works it cites.
Geng Z, Guo MH, Chen H, Li X, Wei K, Lin Z (2021) Is attention better than matrix decomposition? arXiv preprint arXiv:210904553
2021
Later among the works it cites.
Hassanin M, Khan S, Tahtali M (2021) Visual affordance and function understanding: A survey. ACM Computing Surveys (CSUR) 54(3):1–35
2021
Later among the works it cites.
Li Y, Nagarajan T, Xiong B, Grauman K (2021) Ego-exo: Transferring visual representations from third-person to first-person videos. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp 6943–6953
2021
Later among the works it cites.
Liu Z, Lin Y, Cao Y, Hu H, Wei Y, Zhang Z, Lin S, Guo B (2021) Swin transformer: Hierarchical vision transformer using shifted windows. In: Proceedings of the IEEE/CVF international conference on computer vision, pp 10012–10022
2021
Later among the works it cites.
Mandikal P, Grauman K (2021) Learning dexterous grasping with object-centric visual affordances. In: 2021 IEEE International Conference on Robotics and Automation (ICRA), IEEE, pp 6169–6176
2021
Later among the works it cites.
Pan X, Gao Y, Lin Z, Tang F, Dong W, Yuan H, Huang F, Xu C (2021) Unveiling the potential of structure preserving for weakly supervised object localization. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp 11642–11651
2021
Later among the works it cites.
Wu P, Zhai W, Cao Y (2021) Background activation suppression for weakly supervised object localization. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
2021
Later among the works it cites.
Yang Y, Ni Z, Gao M, Zhang J, Tao D (2021) Collaborative pushing and grasping of tightly stacked objects via deep reinforcement learning. IEEE/CAA Journal of Automatica Sinica 9(1):135–145
2021
Later among the works it cites.
Grauman K, Westbury A, Byrne E, Chavis Z, Furnari A, Girdhar R, Hamburger J, Jiang H, Liu M, Liu X, et al. (2022) Ego4d: Around the world in 3,000 hours of egocentric video. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp 18995–19012
2022
Closest in time.
Liu S, Tripathi S, Majumdar S, Wang X (2022) Joint hand motion and interaction hotspots prediction from egocentric videos. arXiv preprint arXiv:220401696
2022
Closest in time.
Luo H, Zhai W, Zhang J, Cao Y, Tao D (2022) Learning affordance grounding from exocentric images. arXiv preprint arXiv:220309905
2022
Closest in time.
Lv Y, Zhang J, Dai Y, Li A, Barnes N, Fan DP (2022) Towards deeper understanding of camouflaged object detection. arXiv preprint arXiv:220511333
2022
Closest in time.
Ramesh A, Dhariwal P, Nichol A, Chu C, Chen M (2022) Hierarchical text-conditional image generation with clip latents. arXiv preprint arXiv:220406125
2022
Closest in time.
Zhai W, Luo H, Zhang J, Cao Y, Tao D (2022) One-shot object affordance detection in the wild. International Journal of Computer Vision (IJCV)
2022
Closest in time.
2022
Closest in time.
Kirillov A, Mintun E, Ravi N, Mao H, Rolland C, Gustafson L, Xiao T, Whitehead S, Berg AC, Lo WY, et al. (2023) Segment anything. arXiv preprint arXiv:230402643
2023
Closest in time.
Shen Y, Song K, Tan X, Li D, Lu W, Zhuang Y (2023) Hugginggpt: Solving ai tasks with chatgpt and its friends in huggingface. arXiv preprint arXiv:230317580
2023
Closest in time.
Zhang Q, Xu Y, Zhang J, Tao D (2023) Vitaev2: Vision transformer advanced by exploring inductive bias for image recognition and beyond. International Journal of Computer Vision (IJCV) pp 1–22
2023
Closest in time.