Fetching the paper…
Reading the bibliography…
We study the task of language-conditioned pick and place in clutter, where a robot should grasp a target object in open clutter and move it to a specified place.
S. Yenamandra, A. Ramachandran, K. Yadav, A. S. Wang, M. Khanna, T. Gervet, T.-Y. Yang, V. Jain, A. Clegg, J. M. Turner et al. , “Homerobot: Open-vocabulary mobile manipulation,” in Conference on Robot Learning . PMLR, 2023, pp. 1975–2011
2011
Earlier work this paper cites.
J. E. King, M. Cognetti, and S. S. Srinivasa, “Rearrangement planning using object-centric and robot-centric action spaces,” in 2016 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2016, pp. 3940–3947
2016
Earlier work this paper cites.
J. L. Schönberger and J.-M. Frahm, “Structure-from-motion revisited,” in Conference on Computer Vision and Pattern Recognition (CVPR) , 2016
2016
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
K. Fang, Y. Bai, S. Hinterstoisser, S. Savarese, and M. Kalakrishnan, “Multi-task domain adaptation for deep learning of instance grasping from simulation,” in 2018 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2018, pp. 3516–3523
2018
Earlier work this paper cites.
M. Danielczuk, J. Mahler, C. Correa, and K. Goldberg, “Linear push policies to increase grasp access for robot bin picking,” in 2018 IEEE 14th international conference on automation science and engineering (CASE) . IEEE, 2018, pp. 1249–1256
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
2019
Earlier work this paper cites.
A. H. Qureshi, A. Mousavian, C. Paxton, M. C. Yip, and D. Fox, “Nerp: Neural rearrangement planning for unknown objects,” in Robotics: Science and Systems (RSS) , 2020
2020
Earlier work this paper cites.
H.-S. Fang, C. Wang, M. Gou, and C. Lu, “Graspnet-1billion: A large-scale benchmark for general object grasping,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2020, pp. 11 444–11 453
2020
Earlier work this paper cites.
A. Murali, A. Mousavian, C. Eppner, C. Paxton, and D. Fox, “6-dof grasping for target-driven object manipulation in clutter,” in 2020 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2020, pp. 6232–6238
2020
Earlier work this paper cites.
A. Kurenkov, J. Taglic, R. Kulkarni, M. Dominguez-Kuhne, A. Garg, R. Martín-Martín, and S. Savarese, “Visuomotor mechanical search: Learning to retrieve target objects in clutter,” in 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) . IEEE, 2020, pp. 8408–8414
2020
Earlier work this paper cites.
Y. Yang, H. Liang, and C. Choi, “A deep learning approach to grasping the invisible,” IEEE Robotics and Automation Letters , vol. 5, no. 2, pp. 2232–2239, 2020
2020
Earlier work this paper cites.
B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Barron, R. Ramamoorthi, and R. Ng, “Nerf: Representing scenes as neural radiance fields for view synthesis,” in Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part I , 2020, pp. 405–421
2020
Earlier work this paper cites.
K. Xu, H. Yu, Q. Lai, Y. Wang, and R. Xiong, “Efficient learning of goal-oriented push-grasping synergy in clutter,” IEEE Robotics and Automation Letters , vol. 6, no. 4, pp. 6337–6344, 2021
2021
Earlier work this paper cites.
A. Zeng, P. Florence, J. Tompson, S. Welker, J. Chien, M. Attarian, T. Armstrong, I. Krasin, D. Duong, V. Sindhwani et al. , “Transporter networks: Rearranging the visual world for robotic manipulation,” in Conference on Robot Learning . PMLR, 2021, pp. 726–747
2021
Earlier work this paper cites.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al. , “Learning transferable visual models from natural language supervision,” in International Conference on Machine Learning . PMLR, 2021, pp. 8748–8763
2021
Earlier work this paper cites.
Z. Jiang, Y. Zhu, M. Svetlik, K. Fang, and Y. Zhu, “Synergies between affordance and geometry: 6-dof grasp detection via implicit representations,” 2021
2021
Earlier work this paper cites.
E. Coumans and Y. Bai, “Pybullet, a python module for physics simulation for games, robotics and machine learning,” http://pybullet.org , 2016–2021
2021
Earlier work this paper cites.
A. Goyal, A. Mousavian, C. Paxton, Y.-W. Chao, B. Okorn, J. Deng, and D. Fox, “Ifor: Iterative flow minimization for robotic object rearrangement,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 14 787–14 797
2022
Earlier work this paper cites.
M. Shridhar, L. Manuelli, and D. Fox, “Cliport: What and where pathways for robotic manipulation,” in Conference on Robot Learning . PMLR, 2022, pp. 894–906
2022
Earlier work this paper cites.
——, “Perceiver-actor: A multi-task transformer for robotic manipulation,” in 6th Annual Conference on Robot Learning , 2022
2022
Earlier work this paper cites.
W. Goodwin, S. Vaze, I. Havoutis, and I. Posner, “Semantically grounded object matching for robust robotic scene rearrangement,” in 2022 International Conference on Robotics and Automation (ICRA) . IEEE, 2022, pp. 11 138–11 144
2022
Earlier work this paper cites.
M. Ahn, A. Brohan, N. Brown, Y. Chebotar, O. Cortes, B. David, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog et al. , “Do as i can, not as i say: Grounding language in robotic affordances,” in Conference on Robot Learning , 2022
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
C. Zhou, C. C. Loy, and B. Dai, “Extract free dense labels from clip,” in European Conference on Computer Vision , 2022, pp. 696–712
2022
Earlier work this paper cites.
K. Xu, H. Yu, R. Huang, D. Guo, Y. Wang, and R. Xiong, “Efficient object manipulation to an arbitrary goal pose: Learning-based anytime prioritized planning,” in 2022 International Conference on Robotics and Automation (ICRA) . IEEE, 2022, pp. 7277–7283
2022
Cited alongside, same era.
H. Tian, C. Song, C. Wang, X. Zhang, and J. Pan, “Sampling-based planning for retrieving near-cylindrical objects in cluttered scenes using hierarchical graphs,” IEEE Transactions on Robotics , vol. 39, no. 1, pp. 165–182, 2022
2022
Cited alongside, same era.
A. Zeng, S. Song, K.-T. Yu, E. Donlon, F. R. Hogan, M. Bauza, D. Ma, O. Taylor, M. Liu, E. Romo et al. , “Robotic pick-and-place of novel objects in clutter with multi-affordance grasping and cross-domain image matching,” The International Journal of Robotics Research , vol. 41, no. 7, pp. 690–705, 2022
2022
Cited alongside, same era.
A. Simeonov, Y. Du, A. Tagliasacchi, J. B. Tenenbaum, A. Rodriguez, P. Agrawal, and V. Sitzmann, “Neural descriptor fields: Se (3)-equivariant object representations for manipulation,” in 2022 International Conference on Robotics and Automation (ICRA) . IEEE, 2022, pp. 6394–6400
K. Xu, Z. Zhou, J. Wu, H. Lu, R. Xiong, and Y. Wang, “Grasp, see and place: Efficient unknown object rearrangement with policy structure prior,” IEEE Transactions on Robotics , 2024
2024
Later among the works it cites.
A. Goyal, V. Blukis, J. Xu, Y. Guo, Y.-W. Chao, and D. Fox, “Rvt-2: Learning precise manipulation from few demonstrations,” in Robotics: Science and Systems (RSS) , 2024
2024
Later among the works it cites.
M. Ji, R.-Z. Qiu, X. Zou, and X. Wang, “Graspsplats: Efficient manipulation with 3d feature splatting,” in 8th Annual Conference on Robot Learning , 2024
2024
Later among the works it cites.
O. Shorinwa, J. Tucker, A. Smith, A. Swann, T. Chen, R. Firoozi, M. D. Kennedy, and M. Schwager, “Splat-mover: multi-stage, open-vocabulary robotic manipulation via editable gaussian splatting,” in 8th Annual Conference on Robot Learning , 2024
2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2022
Cited alongside, same era.
B. Tang and G. S. Sukhatme, “Selective object rearrangement in clutter,” in Conference on Robot Learning . PMLR, 2023, pp. 1001–1010
2023
Cited alongside, same era.
Y. Jiang, A. Gupta, Z. Zhang, G. Wang, Y. Dou, Y. Chen, L. Fei-Fei, A. Anandkumar, Y. Zhu, and L. Fan, “Vima: robot manipulation with multimodal prompts,” in Proceedings of the 40th International Conference on Machine Learning , 2023, pp. 14 975–15 022
2023
Cited alongside, same era.
Y. Ze, G. Yan, Y.-H. Wu, A. Macaluso, Y. Ge, J. Ye, N. Hansen, L. E. Li, and X. Wang, “Gnfactor: Multi-task real robot learning with generalizable neural feature fields,” in Conference on Robot Learning . PMLR, 2023, pp. 284–301
2023
Cited alongside, same era.
T. Gervet, Z. Xian, N. Gkanatsios, and K. Fragkiadaki, “Act3d: 3d feature field transformers for multi-task robotic manipulation,” in 7th Annual Conference on Robot Learning , 2023
2023
Cited alongside, same era.
K. M. Jatavallabhula, A. Kuwajerwala, Q. Gu, M. Omama, T. Chen, A. Maalouf, S. Li, G. S. Iyer, S. Saryazdi, N. V. Keetha et al. , “Conceptfusion: Open-set multimodal 3d mapping,” in Robotics: Science and Systems (RSS) , 2023
2023
Cited alongside, same era.
A. Rashid, S. Sharma, C. M. Kim, J. Kerr, L. Y. Chen, A. Kanazawa, and K. Goldberg, “Language embedded radiance fields for zero-shot task-oriented grasping,” in 7th Annual Conference on Robot Learning , 2023
2023
Cited alongside, same era.
W. Huang, C. Wang, R. Zhang, Y. Li, J. Wu, and L. Fei-Fei, “Voxposer: Composable 3d value maps for robotic manipulation with language models,” in Conference on Robot Learning . PMLR, 2023, pp. 540–562
2023
Cited alongside, same era.
K. Xu, R. Chen, S. Zhao, Z. Li, H. Yu, C. Chen, Y. Wang, and R. Xiong, “Failure-aware policy learning for self-assessable robotics tasks,” in 2023 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2023, pp. 9544–9550
2023
Cited alongside, same era.
Y. Zheng, X. Chen, Y. Zheng, S. Gu, R. Yang, B. Jin, P. Li, C. Zhong, Z. Wang, L. Liu et al. , “Gaussiangrasper: 3d language gaussian splatting for open-vocabulary robotic grasping,” IEEE Robotics and Automation Letters , 2024
2024
Later among the works it cites.
Y. Wang, M. Zhang, Z. Li, T. Kelestemur, K. R. Driggs-Campbell, J. Wu, L. Fei-Fei, and Y. Li, “D 3 fields: Dynamic 3d descriptor fields for zero-shot generalizable rearrangement,” in 8th Annual Conference on Robot Learning , 2024
2024
Later among the works it cites.
Y. Qian, X. Zhu, O. Biza, S. Jiang, L. Zhao, H. Huang, Y. Qi, and R. Platt, “Thinkgrasp: A vision-language system for strategic part grasping in clutter,” in 8th Annual Conference on Robot Learning , 2024
2024
Later among the works it cites.
G. Zhai, X. Cai, D. Huang, Y. Di, F. Manhardt, F. Tombari, N. Navab, and B. Busam, “Sg-bot: Object rearrangement via coarse-to-fine robotic imagination on scene graphs,” in 2024 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2024, pp. 4303–4310
2024
Later among the works it cites.
A. D. Vuong, M. N. Vu, H. Le, B. Huang, H. T. T. Binh, T. Vo, A. Kugi, and A. Nguyen, “Grasp-anything: Large-scale grasp dataset from foundation models,” in 2024 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2024, pp. 14 030–14 037
2024
Later among the works it cites.
2024
Later among the works it cites.
M. Oquab, T. Darcet, T. Moutakanni, H. Vo, M. Szafraniec, V. Khalidov, P. Fernandez, D. Haziza, F. Massa, A. El-Nouby et al. , “Dinov2: Learning robust visual features without supervision,” Transactions on Machine Learning Research Journal , pp. 1–31, 2024
2024
Later among the works it cites.
Y. Yang, H. Yu, X. Lou, Y. Liu, and C. Choi, “Attribute-based robotic grasping with data-efficient adaptation,” IEEE Transactions on Robotics , vol. 40, pp. 1566–1579, 2024
2024
Later among the works it cites.
T.-W. Ke, N. Gkanatsios, and K. Fragkiadaki, “3d diffuser actor: Policy diffusion with 3d scene representations,” in 8th Annual Conference on Robot Learning , 2024
2024
Later among the works it cites.
Y. Deng, J. Wang, J. Zhao, J. Dou, Y. Yang, and Y. Yue, “Openobj: Open-vocabulary object-level neural radiance fields with fine-grained understanding,” IEEE Robotics and Automation Letters , 2024
2024
Later among the works it cites.
S. H. Vemprala, R. Bonatti, A. Bucker, and A. Kapoor, “Chatgpt for robotics: Design principles and model abilities,” IEEE Access , 2024
2024
Later among the works it cites.
N. Wake, A. Kanehira, K. Sasabuchi, J. Takamatsu, and K. Ikeuchi, “Gpt-4v (ision) for robotics: Multimodal task planning from human demonstration,” IEEE Robotics and Automation Letters , 2024
2024
Later among the works it cites.
Y. Hu, F. Lin, T. Zhang, L. Yi, and Y. Gao, “Look before you leap: Unveiling the power of gpt-4v in robotic vision-language planning,” in First Workshop on Vision-Language Models for Navigation and Manipulation at ICRA 2024
2024
Later among the works it cites.
A. O’Neill, A. Rehman, A. Maddukuri, A. Gupta, A. Padalkar, A. Lee, A. Pooley, A. Gupta, A. Mandlekar, A. Jain et al. , “Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0,” in 2024 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2024, pp. 6892–6903
2024
Later among the works it cites.
2024
Later among the works it cites.
M. J. Kim, K. Pertsch, S. Karamcheti, T. Xiao, A. Balakrishna, S. Nair, R. Rafailov, E. P. Foster, P. R. Sanketi, Q. Vuong et al. , “Openvla: An open-source vision-language-action model,” in 8th Annual Conference on Robot Learning , 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
J. Su, M. Ahmed, Y. Lu, S. Pan, W. Bo, and Y. Liu, “Roformer: Enhanced transformer with rotary position embedding,” Neurocomputing , vol. 568, p. 127063, 2024
2024
Later among the works it cites.
“Openai. gpt-4o: Openai’s multimodal vision-language system.” 2023. Accessed: 2024-06-05. [Online]. Available: https://openai.com/research/gpt-4o
2024
Later among the works it cites.
P. Gao, S. Geng, R. Zhang, T. Ma, R. Fang, Y. Zhang, H. Li, and Y. Qiao, “Clip-adapter: Better vision-language models with feature adapters,” International Journal of Computer Vision , vol. 132, no. 2, pp. 581–595, 2024
2024
Later among the works it cites.