Fetching the paper…
Reading the bibliography…
Understanding the environment and a robot's physical reachability is crucial for task execution.
Sam2 encodes the second methionine s-adenosyl transferase in saccharomyces cerevisiae: physiology and regulation of both enzymes
Dominique Thomas, Rodney Rothstein, Nathan Rosenberg, and Yolande Surdin-Kerjan · 1988
Earlier work this paper cites.
Capturing robot workspace structure: representing robot capabilities
Franziska Zacharias, Christoph Borst, and Gerd Hirzinger · 2007
Earlier work this paper cites.
Autonomous online generation of a motor representation of the workspace for intelligent whole-body reaching
Lorenzo Jamone, Martim Brandao, Lorenzo Natale, Kenji Hashimoto, Giulio Sandini, and Atsuo Takanishi · 2014
Earlier work this paper cites.
Scannet: Richly-annotated 3d reconstructions of indoor scenes
Angela Dai, Angel X Chang, Manolis Savva, Maciej Halber, Thomas Funkhouser, and Matthias Nießner · 2017
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy · 2020
Earlier work this paper cites.
Alfred: A benchmark for interpreting grounded instructions for everyday tasks
Mohit Shridhar, Jesse Thomason, Daniel Gordon, Yonatan Bisk, Winson Han, Roozbeh Mottaghi, Luke Zettlemoyer, and Dieter Fox · 2020
Earlier work this paper cites.
Encoding physical constraints in differentiable newton-euler algorithm
Giovanni Sutanto, Austin Wang, Yixin Lin, Mustafa Mukadam, Gaurav Sukhatme, Akshara Rai, and Franziska Meier · 2020
Earlier work this paper cites.
Revisiting embodiedqa: A simple baseline and beyond
Yu Wu, Lu Jiang, and Yi Yang · 2020
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Alec Radford, Kim, et al · 2021
Earlier work this paper cites.
Show me what you can do: Capability calibration on reachable workspace for human-robot collaboration
Xiaofeng Gao, Luyao Yuan, Tianmin Shu, Hongjing Lu, and Song-Chun Zhu · 2022
Earlier work this paper cites.
Simple open-vocabulary object detection with vision transformers, 2022
Matthias Minderer, Alexey Gritsenko, Austin Stone, et al · 2022
Earlier work this paper cites.
Workspace-based model predictive control for cable-driven robots
Chen Song and Darwin Lau · 2022
Earlier work this paper cites.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Cited alongside, same era.
Palm: Scaling language modeling with pathways
Aakanksha Chowdhery, Narang, et al · 2023
Cited alongside, same era.
Palm-e: An embodied multimodal language model
Danny Driess, Fei Xia, Mehdi SM Sajjadi, Corey Lynch, Aakanksha Chowdhery, Brian Ichter, Ayzaan Wahid, Jonathan Tompson, Quan Vuong, Tianhe Yu, et al · 2023
Cited alongside, same era.
Voxposer: Composable 3d value maps for robotic manipulation with language models
Wenlong Huang, Chen Wang, Ruohan Zhang, Yunzhu Li, Jiajun Wu, and Li Fei-Fei · 2023
Cited alongside, same era.
Code as policies: Language model programs for embodied control
Jacky Liang, Wenlong Huang, Fei Xia, Peng Xu, Karol Hausman, Brian Ichter, Pete Florence, and Andy Zeng · 2023
Instance-aware exploration-verification-exploitation for instance imagegoal navigation
Xiaohan Lei, Min Wang, Wengang Zhou, Li Li, and Houqiang Li · 2024
Later among the works it cites.
Llava-next-interleave: Tackling multi-image, video, and 3d in large multimodal models
Feng Li, Renrui Zhang, Hao Zhang, Yuanhan Zhang, Bo Li, Wei Li, Zejun Ma, and Chunyuan Li · 2024
Later among the works it cites.
Openeqa: Embodied question answering in the era of foundation models
Arjun Majumdar, Ajay, et al · 2024
Later among the works it cites.
Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0
Abby O’Neill, Rehman, et al · 2024
Later among the works it cites.
Robovqa: Multimodal long-horizon reasoning for robotics
Pierre Sermanet, Ding, et al · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Grounding dino: Marrying dino with grounded pre-training for open-set object detection
Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, Hao Zhang, Jie Yang, Qing Jiang, Chunyuan Li, Jianwei Yang, Hang Su, et al · 2023
Cited alongside, same era.
Sigmoid loss for language image pre-training, 2023
Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, and Lucas Beyer · 2023
Cited alongside, same era.
Spatialbot: Precise spatial understanding with vision language models
Wenxiao Cai, Yaroslav Ponomarenko, Jianhao Yuan, Xiaoqi Li, Wankou Yang, Hao Dong, and Bo Zhao · 2024
Cited alongside, same era.
Spoc: Imitating shortest paths in simulation enables effective navigation and manipulation in the real world
Kiana Ehsani, Tanmay Gupta, Rose Hendrix, Jordi Salvador, Luca Weihs, Kuo-Hao Zeng, Kunal Pratap Singh, Yejin Kim, Winson Han, Alvaro Herrasti, et al · 2024
Cited alongside, same era.
Multiply: A multisensory object-centric embodied large language model in 3d world
Yining Hong, Zishuo Zheng, Peihao Chen, Yian Wang, Junyan Li, and Chuang Gan · 2024
Cited alongside, same era.
Rekep: Spatio-temporal reasoning of relational keypoint constraints for robotic manipulation
Wenlong Huang, Chen Wang, Yunzhu Li, Ruohan Zhang, and Li Fei-Fei · 2024
Cited alongside, same era.
Lisa: Reasoning segmentation via large language model
Xin Lai, Zhuotao Tian, Yukang Chen, Yanwei Li, Yuhui Yuan, Shu Liu, and Jiaya Jia · 2024
Cited alongside, same era.
Qwen Team · 2024
Later among the works it cites.
Asap: Automated sequence planning for complex robotic assembly with physical feasibility
Yunsheng Tian, Karl DD Willis, Bassel Al Omari, Jieliang Luo, Pingchuan Ma, Yichen Li, Farhad Javid, Edward Gu, Joshua Jacob, Shinjiro Sueda, et al · 2024
Later among the works it cites.
Language-driven grasp detection
An Dinh Vuong, Minh Nhat Vu, Baoru Huang, Nghia Nguyen, Hieu Le, Thieu Vo, and Anh Nguyen · 2024
Later among the works it cites.
Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution
Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang, Zhihao Fan, Jinze Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, et al · 2024
Later among the works it cites.
Over-nav: Elevating iterative vision-and-language navigation with open-vocabulary detection and structured representation
Ganlong Zhao, Guanbin Li, Weikai Chen, and Yizhou Yu · 2024
Later among the works it cites.
3d-vla: A 3d vision-language-action generative world model
Haoyu Zhen, Xiaowen Qiu, Peihao Chen, Jincheng Yang, Xin Yan, Yilun Du, Yining Hong, and Chuang Gan · 2024
Later among the works it cites.
Mmbench: Is your multi-modal model an all-around player?
Yuan Liu, Haodong Duan, Yuanhan Zhang, Bo Li, Songyang Zhang, Wangbo Zhao, Yike Yuan, Jiaqi Wang, Conghui He, Ziwei Liu, et al · 2025
Closest in time.