Fetching the paper…
Reading the bibliography…
From rearranging objects on a table to putting groceries into shelves, robots must plan precise action points to perform tasks accurately and reliably.
Spatial mental models derived from survey and route descriptions
H. A. Taylor and B. Tversky · 1992
Earlier work this paper cites.
A new look at infant pointing
M. Tomasello, M. Carpenter, and U. Liszkowski · 2007
Earlier work this paper cites.
Thinking with sketches
B. Tversky and M. Suwa · 2009
Earlier work this paper cites.
Legibility and predictability of robot motion
A. D. Dragan, K. C. Lee, and S. S. Srinivasa · 2013
Earlier work this paper cites.
CoppeliaSim (formerly V-REP): a Versatile and Scalable Robot Simulation Framework
E. Rohmer, S. P. N. Singh, and M. Freese · 2013
Earlier work this paper cites.
Modeling context in referring expressions
L. Yu, P. Poirson, S. Yang, A. C. Berg, and T. L. Berg · 2016
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma, et al · 2017
Earlier work this paper cites.
Posecnn: A convolutional neural network for 6d object pose estimation in cluttered scenes
Y. Xiang, T. Schmidt, V. Narayanan, and D. Fox · 2017
Earlier work this paper cites.
Scenenet rgb-d: Can 5m synthetic images beat generic imagenet pre-training on indoor segmentation?
J. McCormac, A. Handa, S. Leutenegger, and A. J. Davison · 2017
Earlier work this paper cites.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Y. Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh · 2017
Earlier work this paper cites.
Affordancenet: An end-to-end deep learning approach for object affordance detection
T.-T. Do, A. Nguyen, and I. Reid · 2018
Earlier work this paper cites.
Dense object nets: Learning dense visual object descriptors by and for robotic manipulation
P. Florence, L. Manuelli, and R. Tedrake · 2018
Earlier work this paper cites.
Vizwiz grand challenge: Answering visual questions from blind people
D. Gurari, Q. Li, A. J. Stangl, A. Guo, C. Lin, K. Grauman, J. Luo, and J. P. Bigham · 2018
Earlier work this paper cites.
Gqa: A new dataset for real-world visual reasoning and compositional question answering
D. A. Hudson and C. D. Manning · 2019
Earlier work this paper cites.
Clevrer: Collision events for video representation and reasoning
K. Yi, C. Gan, Y. Li, P. Kohli, J. Wu, A. Torralba, and J. B. Tenenbaum · 2019
Earlier work this paper cites.
kpam: Keypoint affordances for category-level robotic manipulation
L. Manuelli, W. Gao, P. Florence, and R. Tedrake · 2019
Earlier work this paper cites.
Lvis: A dataset for large vocabulary instance segmentation
A. Gupta, P. Dollar, and R. Girshick · 2019
Earlier work this paper cites.
Habitat: A platform for embodied ai research
M. Savva, A. Kadian, O. Maksymets, Y. Zhao, E. Wijmans, B. Jain, J. Straub, J. Liu, V. Koltun, J. Malik, et al · 2019
Earlier work this paper cites.
Towards vqa models that can read
A. Singh, V. Natarjan, M. Shah, Y. Jiang, X. Chen, D. Parikh, and M. Rohrbach · 2019
Earlier work this paper cites.
D. Ding, F. Hill, A. Santoro, and M. Botvinick · 2020
Earlier work this paper cites.
Transporter networks: Rearranging the visual world for robotic manipulation
A. Zeng, P. Florence, J. Tompson, S. Welker, J. Chien, M. Attarian, T. Armstrong, I. Krasin, D. Duong, V. Sindhwani, et al · 2020
Earlier work this paper cites.
Keto: Learning keypoint representations for tool manipulation
Z. Qin, K. Fang, Y. Zhu, L. Fei-Fei, and S. Savarese · 2020
Cited alongside, same era.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Cited alongside, same era.
SAPIEN: A simulated part-based interactive environment
F. Xiang, Y. Qin, K. Mo, Y. Xia, H. Zhu, F. Liu, M. Liu, H. Jiang, Y. Yuan, H. Wang, L. Yi, A. X. Chang, L. J. Guibas, and H. Su · 2020
Cited alongside, same era.
Where2act: From pixels to actions for articulated 3d objects
K. Mo, L. J. Guibas, M. Mukadam, A. Gupta, and S. Tulsiani · 2021
Cited alongside, same era.
Contact-graspnet: Efficient 6-dof grasp generation in cluttered scenes
M. Sundermeyer, A. Mousavian, R. Triebel, and D. Fox · 2021
Cited alongside, same era.
Same object, different grasps: Data and semantic knowledge for task-oriented grasping
Code as policies: Language model programs for embodied control
J. Liang, W. Huang, F. Xia, P. Xu, K. Hausman, B. Ichter, P. Florence, and A. Zeng · 2023
Later among the works it cites.
Progprompt: Generating situated robot task plans using large language models
I. Singh, V. Blukis, A. Mousavian, A. Goyal, D. Xu, J. Tremblay, D. Fox, J. Thomason, and A. Garg · 2023
Later among the works it cites.
Rvt: Robotic view transformer for 3d object manipulation
A. Goyal, J. Xu, Y. Guo, V. Blukis, Y.-W. Chao, and D. Fox · 2023
Later among the works it cites.
Voxposer: Composable 3d value maps for robotic manipulation with language models
W. Huang, C. Wang, R. Zhang, Y. Li, J. Wu, and L. Fei-Fei · 2023
Later among the works it cites.
Rt-2: Vision-language-action models transfer web knowledge to robotic control
A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, X. Chen, K. Choromanski, T. Ding, D. Driess, A. Dubey, C. Finn, et al · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Murali, W. Liu, K. Marino, S. Chernova, and A. Gupta · 2021
Cited alongside, same era.
Synergies between affordance and geometry: 6-dof grasp detection via implicit representations
Z. Jiang, Y. Zhu, M. Svetlik, K. Fang, and Y. Zhu · 2021
Cited alongside, same era.
Manipulathor: A framework for visual object manipulation
K. Ehsani, W. Han, A. Herrasti, E. VanderBilt, L. Weihs, E. Kolve, A. Kembhavi, and R. Mottaghi · 2021
Cited alongside, same era.
Acronym: A large-scale grasp dataset based on simulation
C. Eppner, A. Mousavian, and D. Fox · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Cited alongside, same era.
Perceiver-actor: A multi-task transformer for robotic manipulation
M. Shridhar, L. Manuelli, and D. Fox · 2022
Cited alongside, same era.
Rt-1: Robotics transformer for real-world control at scale
A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, J. Hsu, et al · 2022
Cited alongside, same era.
Later among the works it cites.
Ar2-d2: Training a robot without a robot
J. Duan, Y. R. Wang, M. Shridhar, D. Fox, and R. Krishna · 2023
Later among the works it cites.
M2t2: Multi-task masked transformer for object-centric pick and place
W. Yuan, A. Murali, A. Mousavian, and D. Fox · 2023
Later among the works it cites.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, March 2023
W.-L. Chiang, Z. Li, Z. Lin, Y. Sheng, Z. Wu, H. Zhang, L. Zheng, S. Zhuang, Y. Zhuang, J. E. Gonzalez, I. Stoica, and E. P. Xing · 2023
Later among the works it cites.
Improved baselines with visual instruction tuning, 2023
H. Liu, C. Li, Y. Li, and Y. J. Lee · 2023
Later among the works it cites.
A. Murali, A. Mousavian, C. Eppner, A. Fishman, and D. Fox · 2023
Later among the works it cites.
Motion policy networks
A. Fishman, A. Murali, C. Eppner, B. Peele, B. Boots, and D. Fox · 2023
Later among the works it cites.
Vl-grasp: a 6-dof interactive grasp policy for language-oriented objects in cluttered indoor scenes
Y. Lu, Y. Fan, B. Deng, F. Liu, Y. Li, and S. Wang · 2023
Later among the works it cites.
Mme: A comprehensive evaluation benchmark for multimodal large language models
C. Fu, P. Chen, Y. Shen, Y. Qin, M. Zhang, X. Lin, J. Yang, X. Zheng, K. Li, X. Sun, Y. Wu, and R. Ji · 2023
Later among the works it cites.
Hello gpt-4o, May 2024
OpenAI · 2024
Closest in time.
Moka: Open-vocabulary robotic manipulation through mark-based visual prompting
F. Liu, K. Fang, P. Abbeel, and S. Levine · 2024
Closest in time.
Pivot: Iterative visual prompting elicits actionable knowledge for vlms
S. Nasiriany, F. Xia, W. Yu, T. Xiao, J. Liang, I. Dasgupta, A. Xie, D. Driess, A. Wahid, Z. Xu, et al · 2024
Closest in time.
Octo: An open-source generalist robot policy
O. M. Team, D. Ghosh, H. Walke, K. Pertsch, K. Black, O. Mees, S. Dasari, J. Hejna, T. Kreiman, C. Xu, et al · 2024
Closest in time.
Llava-next: Improved reasoning, ocr, and world knowledge, January 2024
H. Liu, C. Li, Y. Li, B. Li, Y. Zhang, S. Shen, and Y. J. Lee · 2024
Closest in time.
Spatialvlm: Endowing vision-language models with spatial reasoning capabilities
B. Chen, Z. Xu, S. Kirmani, B. Ichter, D. Driess, P. Florence, D. Sadigh, L. Guibas, and F. Xia · 2024
Closest in time.
Spacellava, 2024
S. Remyx AI (Mayorquin and T. Rodriguez · 2024
Closest in time.