Fetching the paper…
Reading the bibliography…
Robotic behavior synthesis, the problem of understanding multimodal inputs and generating precise physical control for robots, is an important part of Embodied AI.
A unified approach for motion and force control of robot manipulators: The operational space formulation
Khatib, O · 1987
Earlier work this paper cites.
The open motion planning library
Sucan, I. A., Moll, M., and Kavraki, L. E · 2012
Earlier work this paper cites.
Manipulation task simulation using ros and gazebo
Qian, W., Xia, Z., Xiong, J., Gan, Y., Guo, Y., Weng, S., Deng, H., Hu, Y., and Zhang, J · 2014
Earlier work this paper cites.
Benchmarking in manipulation research: The ycb object and model set and benchmarking protocols
Calli, B., Walsman, A., Singh, A., Srinivasa, S., Abbeel, P., and Dollar, A. M · 2015
Earlier work this paper cites.
Kinematic and dynamic modelling of ur5 manipulator
Kebria, P. M., Al-Wais, S., Abdi, H., and Nahavandi, S · 2016
Earlier work this paper cites.
Robot Operating System (ROS). , volume 1
Koubâa, A. et al · 2017
Earlier work this paper cites.
Pointnet++: Deep hierarchical feature learning on point sets in a metric space
Qi, C. R., Yi, L., Su, H., and Guibas, L. J · 2017
Earlier work this paper cites.
On evaluation of embodied navigation agents
Anderson, P., Chang, A., Chaplot, D. S., Dosovitskiy, A., Gupta, S., Koltun, V., Kosecka, J., Malik, J., Mottaghi, R., Savva, M., et al · 2018
Earlier work this paper cites.
Predicting gaze in egocentric video by learning task-dependent attention transition
Huang, Y., Cai, M., Li, Z., and Sato, Y · 2018
Earlier work this paper cites.
VirtualHome: Simulating household activities via programs
Puig, X., Ra, K., Boben, M., Li, J., Wang, T., Fidler, S., and Torralba, A · 2018
Earlier work this paper cites.
Open3d: A modern library for 3d data processing
Zhou, Q.-Y., Park, J., and Koltun, V · 2018
Earlier work this paper cites.
PartNet: A large-scale benchmark for fine-grained and hierarchical part-level 3D object understanding
Mo, K., Zhu, S., Chang, A. X., Yi, L., Tripathi, S., Guibas, L. J., and Su, H · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al · 2020
Earlier work this paper cites.
Sapien: A simulated part-based interactive environment
Xiang, F., Qin, Y., Mo, K., Xia, Y., Zhu, H., Liu, F., Liu, M., Jiang, H., Yuan, Y., Wang, H., et al · 2020
Earlier work this paper cites.
Habitat-matterport 3d dataset (hm3d): 1000 large-scale 3d environments for embodied ai
Ramakrishnan, S. K., Gokaslan, A., Wijmans, E., Maksymets, O., Clegg, A., Turner, J., Undersander, E., Galuba, W., Westbury, A., Chang, A. X., et al · 2021
Earlier work this paper cites.
Channel-wise attention-based network for self-supervised monocular depth estimation
Yan, J., Zhao, H., Bu, P., and Jin, Y · 2021
Earlier work this paper cites.
Do as i can and not as i say: Grounding language in robotic affordances
Ahn, M., Brohan, A., Brown, N., Chebotar, Y., Cortes, O., David, B., Finn, C., Fu, C., Gopalakrishnan, K., Hausman, K., Herzog, A., Ho, D., Hsu, J., Ibarz, J., Ichter, B., Irpan, A., Jang, E., Ruano, R. J., Jeffrey, K., Jesmonth, S., Joshi, N., Julian, R., Kalashnikov, D., Kuang, Y., Lee, K.-H., Levine, S., Lu, Y., Luu, L., Parada, C., Pastor, P., Quiambao, J., Rao, K., Rettinghouse, J., Reyes, D., Sermanet, P., Sievers, N., Tan, C., Toshev, A., Vanhoucke, V., Xia, F., Xiao, T., Xu, P., Xu, S., Yan, M., and Zeng, A · 2022
Earlier work this paper cites.
Flamingo: a visual language model for few-shot learning
Alayrac, J.-B., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., et al · 2022
Earlier work this paper cites.
Rt-1: Robotics transformer for real-world control at scale
Brohan, A., Brown, N., Carbajal, J., Chebotar, Y., Dabis, J., Finn, C., Gopalakrishnan, K., Hausman, K., Herzog, A., Hsu, J., et al · 2022
Earlier work this paper cites.
Pali: A jointly-scaled multilingual language-image model
Chen, X., Wang, X., Changpinyo, S., Piergiovanni, A., Padlewski, P., Salz, D., Goodman, S., Grycner, A., Mustafa, B., Beyer, L., et al · 2022
Earlier work this paper cites.
Google scanned objects: A high-quality dataset of 3d scanned household items
Downs, L., Francis, A., Koenig, N., Kinman, B., Hickman, R., Reymann, K., McHugh, T. B., and Vanhoucke, V · 2022
Cited alongside, same era.
Flowbot3d: Learning 3d articulation flow to manipulate articulated objects
Eisner, B., Zhang, H., and Held, D · 2022
Cited alongside, same era.
Eva: Exploring the limits of masked visual representation learning at scale
Fang, Y., Wang, W., Xie, B., Sun, Q., Wu, L., Wang, X., Huang, T., Wang, X., and Cao, Y · 2022
Cited alongside, same era.
Ego4d: Around the world in 3,000 hours of egocentric video
Grauman, K., Westbury, A., Byrne, E., Chavis, Z., Furnari, A., Girdhar, R., Hamburger, J., Jiang, H., Liu, M., Liu, X., et al · 2022
Cited alongside, same era.
The franka emika robot: A reference platform for robotics research and education
Haddadin, S., Parusel, S., Johannsmeier, L., Golz, S., Gabl, S., Walch, F., Sabaghian, M., Jähne, C., Hausperger, L., and Haddadin, S · 2022
Anygrasp: Robust and efficient grasp perception in spatial and temporal domains
Fang, H.-S., Wang, C., Fang, H., Gou, M., Liu, J., Yan, H., Liu, W., Xie, Y., and Lu, C · 2023
Later among the works it cites.
Physically grounded vision-language models for robotic manipulation
Gao, J., Sarkar, B., Xia, F., Xiao, T., Wu, J., Ichter, B., Majumdar, A., and Sadigh, D · 2023
Later among the works it cites.
Tree-planner: Efficient close-loop task planning with large language models
Hu, M., Mu, Y., Yu, X., Ding, M., Wu, S., Shao, W., Chen, Q., Wang, B., Qiao, Y., and Luo, P · 2023
Later among the works it cites.
Visual language maps for robot navigation
Huang, C., Mees, O., Zeng, A., and Burgard, W · 2023
Later among the works it cites.
Habitat synthetic scenes dataset (hssd-200): An analysis of 3d scene scale and realism tradeoffs for objectgoal navigation, 2023
Khanna, M., Mao, Y., Jiang, H., Haresh, S., Shacklett, B., Batra, D., Clegg, A., Undersander, E., Chang, A. X., and Savva, M · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Akb-48: A real-world articulated object knowledge base
Liu, L., Xu, W., Fu, H., Qian, S., Yu, Q., Han, Y., and Lu, C · 2022
Cited alongside, same era.
R3m: A universal visual representation for robot manipulation
Nair, S., Rajeswaran, A., Kumar, V., Finn, C., and Gupta, A · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Cited alongside, same era.
Sinlu: Sinu-sigmoidal linear unit
Paul, A., Bandyopadhyay, R., Yoon, J. H., Geem, Z. W., and Sarkar, R · 2022
Cited alongside, same era.
Planning with large language models via corrective re-prompting
Raman, S. S., Cohen, V., Rosen, E., Idrees, I., Paulius, D., and Tellex, S · 2022
Cited alongside, same era.
Universal manipulation policy network for articulated objects
Xu, Z., He, Z., and Song, S · 2022
Cited alongside, same era.
Socratic models: Composing zero-shot multimodal reasoning with language
Zeng, A., Attarian, M., Ichter, B., Choromanski, K., Wong, A., Welker, S., Tombari, F., Purohit, A., Ryoo, M., Sindhwani, V., et al · 2022
Cited alongside, same era.
Later among the works it cites.
Lisa: Reasoning segmentation via large language model
Lai, X., Tian, Z., Chen, Y., Li, Y., Yuan, Y., Liu, S., and Jia, J · 2023
Later among the works it cites.
Code as Policies: Language model programs for embodied control
Liang, J., Huang, W., Xia, F., Xu, P., Hausman, K., Ichter, B., Florence, P., and Zeng, A · 2023
Later among the works it cites.
Lan-grasp: Using large language models for semantic object grasping
Mirjalili, R., Krawez, M., Silenzi, S., Blei, Y., and Burgard, W · 2023
Later among the works it cites.
Embodiedgpt: Vision-language pre-training via embodied chain of thought
Mu, Y., Zhang, Q., Hu, M., Wang, W., Ding, M., Jin, J., Wang, B., Dai, J., Qiao, Y., and Luo, P · 2023
Later among the works it cites.
Open x-embodiment: Robotic learning datasets and rt-x models
Padalkar, A., Pooley, A., Jain, A., Bewley, A., Herzog, A., Irpan, A., Khazatsky, A., Rai, A., Singh, A., Brohan, A., et al · 2023
Later among the works it cites.
Kosmos-2: Grounding multimodal large language models to the world
Peng, Z., Wang, W., Dong, L., Hao, Y., Huang, S., Ma, S., and Wei, F · 2023
Later among the works it cites.
Sayplan: Grounding large language models using 3d scene graphs for scalable task planning
Rana, K., Haviland, J., Garg, S., Abou-Chakra, J., Reid, I., and Suenderhauf, N · 2023
Later among the works it cites.
ProgPrompt: Generating situated robot task plans using large language models
Singh, I., Blukis, V., Mousavian, A., Goyal, A., Xu, D., Tremblay, J., Fox, D., Thomason, J., and Garg, A · 2023
Later among the works it cites.
Llm-planner: Few-shot grounded planning for embodied agents with large language models
Song, C. H., Wu, J., Washington, C., Sadler, B. M., Chao, W.-L., and Su, Y · 2023
Later among the works it cites.
Generative pretraining in multimodality
Sun, Q., Yu, Q., Cui, Y., Zhang, F., Zhang, X., Wang, Y., Gao, H., Liu, J., Huang, T., and Wang, X · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al · 2023
Later among the works it cites.
Chatgpt for robotics: Design principles and model abilities
Vemprala, S., Bonatti, R., Bucker, A., and Kapoor, A · 2023
Later among the works it cites.
mplug-docowl: Modularized multimodal large language model for document understanding, 2023
Ye, J., Hu, A., Xu, H., Ye, Q., Yan, M., Dan, Y., Zhao, C., Xu, G., Li, C., Tian, J., Qi, Q., Zhang, J., and Huang, F · 2023
Later among the works it cites.
Plan4mc: Skill reinforcement learning and planning for open-world minecraft tasks
Yuan, H., Zhang, C., Wang, H., Xie, F., Cai, P., Dong, H., and Lu, Z · 2023
Later among the works it cites.
Svit: Scaling up visual instruction tuning
Zhao, B., Wu, B., and Huang, T · 2023
Later among the works it cites.
3d implicit transporter for temporally consistent keypoint discovery
Zhong, C., Zheng, Y., Zheng, Y., Zhao, H., Yi, L., Mu, X., Wang, L., Li, P., Zhou, G., Yang, C., et al · 2023
Later among the works it cites.