Fetching the paper…
Reading the bibliography…
Comprehending natural language instructions is a critical skill for robots to cooperate effectively with humans.
S. Hinterstoisser, V. Lepetit, S. Ilic, S. Holzer, G. Bradski, K. Konolige, and N. Navab, “Model based training, detection and pose estimation of texture-less 3d objects in heavily cluttered scenes,” in Asian conference on computer vision . Springer, 2012, pp. 548–562
2012
Earlier work this paper cites.
S. Kazemzadeh, V. Ordonez, M. Matten, and T. Berg, “Referitgame: Referring to objects in photographs of natural scenes,” in Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP) , 2014, pp. 787–798
2014
Earlier work this paper cites.
J. Mao, J. Huang, A. Toshev, O. Camburu, A. L. Yuille, and K. Murphy, “Generation and comprehension of unambiguous object descriptions,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 11–20
2016
Earlier work this paper cites.
L. Yu, P. Poirson, S. Yang, A. C. Berg, and T. L. Berg, “Modeling context in referring expressions,” in European Conference on Computer Vision . Springer, 2016, pp. 69–85
2016
Earlier work this paper cites.
T. Hodaň, J. Matas, and Š. Obdržálek, “On evaluation of 6d object pose estimation,” in European Conference on Computer Vision . Springer, 2016, pp. 606–619
2016
Earlier work this paper cites.
M. Rad and V. Lepetit, “Bb8: A scalable, accurate, robust to partial occlusion method for predicting the 3d poses of challenging objects without using depth,” in Proceedings of the IEEE international conference on computer vision , 2017, pp. 3828–3836
2017
Earlier work this paper cites.
A. Zeng, K.-T. Yu, S. Song, D. Suo, E. Walker, A. Rodriguez, and J. Xiao, “Multi-view self-supervised deep learning for 6d pose estimation in the amazon picking challenge,” in 2017 IEEE international conference on robotics and automation (ICRA) . IEEE, 2017, pp. 1386–1383
2017
Earlier work this paper cites.
J. Johnson, B. Hariharan, L. Van Der Maaten, L. Fei-Fei, C. Lawrence Zitnick, and R. Girshick, “Clevr: A diagnostic dataset for compositional language and elementary visual reasoning,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2017, pp. 2901–2910
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
C. Hennersperger, B. Fuerst, S. Virga, O. Zettinig, B. Frisch, T. Neff, and N. Navab, “Towards mri-based autonomous robotic us acquisitions: a first feasibility study,” IEEE transactions on medical imaging , vol. 36, no. 2, pp. 538–548, 2017
2017
Earlier work this paper cites.
J. Hatori, Y. Kikuchi, S. Kobayashi, K. Takahashi, Y. Tsuboi, Y. Unno, W. Ko, and J. Tan, “Interactively picking real-world objects with unconstrained spoken language instructions,” in 2018 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2018, pp. 3774–3781
2018
Earlier work this paper cites.
M. Shridhar and D. Hsu, “Interactive visual grounding of referring expressions for human-robot interaction,” Robotics: Science and Systems (RSS) , 2018
2018
Earlier work this paper cites.
Y. Xiang, T. Schmidt, V. Narayanan, and D. Fox, “Posecnn: A convolutional neural network for 6d object pose estimation in cluttered scenes,” in Robotics: Science and Systems (RSS) , 2018
2018
Earlier work this paper cites.
Y. Li, G. Wang, X. Ji, Y. Xiang, and D. Fox, “Deepim: Deep iterative matching for 6d pose estimation,” in Proceedings of the European Conference on Computer Vision (ECCV) , 2018, pp. 683–698
2018
Earlier work this paper cites.
J. Tremblay, T. To, B. Sundaralingam, Y. Xiang, D. Fox, and S. Birchfield, “Deep object pose estimation for semantic robotic grasping of household objects,” in Conference on Robot Learning (CoRL) , 2018
2018
Earlier work this paper cites.
L. Yu, Z. Lin, X. Shen, J. Yang, X. Lu, M. Bansal, and T. L. Berg, “Mattnet: Modular attention network for referring expression comprehension,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2018, pp. 1307–1315
2018
Earlier work this paper cites.
2018
Cited alongside, same era.
H. Ahn, S. Choi, N. Kim, G. Cha, and S. Oh, “Interactive text2pickup networks for natural language-based human–robot collaboration,” IEEE Robotics and Automation Letters , vol. 3, no. 4, pp. 3308–3315, 2018
2018
Cited alongside, same era.
S. Peng, Y. Liu, Q. Huang, X. Zhou, and H. Bao, “Pvnet: Pixel-wise voting network for 6dof pose estimation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2019, pp. 4561–4570
2019
Cited alongside, same era.
Z. Li, G. Wang, and X. Ji, “Cdpn: Coordinates-based disentangled pose network for real-time rgb-based 6-dof object pose estimation,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2019, pp. 7678–7687
2019
Cited alongside, same era.
Y. Chen, R. Xu, Y. Lin, and P. A. Vela, “A joint network for grasp detection conditioned on natural language commands,” in 2021 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2021, pp. 4576–4582
2021
Later among the works it cites.
H. Zhang, Y. Lu, C. Yu, D. Hsu, X. La, and N. Zheng, “Invigorate: Interactive visual grounding and grasping in clutter,” Robotics: Science and Systems (RSS) , 2021
2021
Later among the works it cites.
H. Ahn, O. Kwon, K. Kim, J. Jeong, H. Jun, H. Lee, D. Lee, and S. Oh, “Visually grounding language instruction for history-dependent manipulation,” in 2022 International Conference on Robotics and Automation (ICRA) . IEEE, 2022, pp. 675–682
2022
Later among the works it cites.
C. Cheang, H. Lin, Y. Fu, and X. Xue, “Learning 6-dof object poses to grasp category-level objects by language instructions,” in 2022 International Conference on Robotics and Automation (ICRA) , 2022, pp. 8476–8482
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Litvak, A. Biess, and A. Bar-Hillel, “Learning pose estimation for high-precision robotic assembly using simulated depth images,” in 2019 International Conference on Robotics and Automation (ICRA) . IEEE, 2019, pp. 3521–3527
2019
Cited alongside, same era.
R. Liu, C. Liu, Y. Bai, and A. L. Yuille, “Clevr-ref+: Diagnosing visual reasoning with referring expressions,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2019, pp. 4185–4194
2019
Cited alongside, same era.
2019
Cited alongside, same era.
Y. Labbé, J. Carpentier, M. Aubry, and J. Sivic, “Cosypose: Consistent multi-view multi-object 6d pose estimation,” in European Conference on Computer Vision . Springer, 2020, pp. 574–591
2020
Cited alongside, same era.
Y. Hu, P. Fua, W. Wang, and M. Salzmann, “Single-stage 6d object pose estimation,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2020, pp. 2930–2939
2020
Cited alongside, same era.
X. Deng, Y. Xiang, A. Mousavian, C. Eppner, T. Bretl, and D. Fox, “Self-supervised 6d object pose estimation for robot manipulation,” in 2020 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2020, pp. 3665–3671
2020
Cited alongside, same era.
S. Stevšić, S. Christen, and O. Hilliges, “Learning to assemble: Estimating 6d poses for robotic object-object manipulation,” IEEE Robotics and Automation Letters , vol. 5, no. 2, pp. 1159–1166, 2020
2020
Cited alongside, same era.
M. Denninger, M. Sundermeyer, D. Winkelbauer, D. Olefir, T. Hodan, Y. Zidan, M. Elbadrawy, M. Knauer, H. Katam, and A. Lodhi, “Blenderproc: Reducing the reality gap with photorealistic rendering,” in International Conference on Robotics: Sciene and Systems, RSS 2020 , 2020
2020
Cited alongside, same era.
R. L. Haugaard and A. G. Buch, “Surfemb: Dense and continuous correspondence distributions for object pose estimation with learnt surface embeddings,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 6749–6758
2022
Later among the works it cites.
K. Chen, R. Cao, S. James, Y. Li, Y.-H. Liu, P. Abbeel, and Q. Dou, “Sim-to-real 6d object pose estimation via iterative self-training for robotic bin-picking,” Proceedings of the European Conference on Computer Vision (ECCV) , 2022
2022
Later among the works it cites.
B. Wen, W. Lian, K. Bekris, and S. Schaal, “Catgrasp: Learning category-level task-relevant grasping in clutter from simulation,” in 2022 International Conference on Robotics and Automation (ICRA) . IEEE, 2022, pp. 6401–6408
2022
Later among the works it cites.
Y. Di, R. Zhang, Z. Lou, F. Manhardt, X. Ji, N. Navab, and F. Tombari, “Gpv-pose: Category-level object pose estimation via geometry-guided point-wise voting,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 6781–6791
2022
Later among the works it cites.
R. Zhang, Y. Di, Z. Lou, F. Manhardt, F. Tombari, and X. Ji, “Rbp-pose: Residual bounding box projection for category-level pose estimation,” in European Conference on Computer Vision . Springer, 2022, pp. 655–672
2022
Later among the works it cites.
R. Zhang, Y. Di, F. Manhardt, F. Tombari, and X. Ji, “Ssp-pose: Symmetry-aware shape prior deformation for direct category-level object pose estimation,” in 2022 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) . IEEE, 2022, pp. 7452–7459
2022
Later among the works it cites.
2022
Later among the works it cites.
G. Zhai, D. Huang, S.-C. Wu, H. Jung, Y. Di, F. Manhardt, F. Tombari, N. Navab, and B. Busam, “Monograspnet: 6-dof grasping with a single rgb image,” in 2023 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2023, pp. 1708–1714
2023
Closest in time.
2023
Closest in time.
M. Zaccaria, F. Manhardt, Y. Di, F. Tombari, J. Aleotti, and M. Giorgini, “Self-supervised category-level 6d object pose estimation with optical flow consistency,” IEEE Robotics and Automation Letters , vol. 8, no. 5, pp. 2510–2517, 2023
2023
Closest in time.
Y. Su, Y. Di, G. Zhai, F. Manhardt, J. Rambach, B. Busam, D. Stricker, and F. Tombari, “Opa-3d: Occlusion-aware pixel-wise aggregation for monocular 3d object detection,” IEEE Robotics and Automation Letters , vol. 8, no. 3, pp. 1327–1334, 2023
2023
Closest in time.