Fetching the paper…
Reading the bibliography…
Despite great strides in language-guided manipulation, existing work has been constrained to table-top settings.
Trajectories and keyframes for kinesthetic teaching: A human-robot interaction perspective
B. Akgun, M. Cakmak, J. W. Yoo, and A. L. Thomaz · 2012
Earlier work this paper cites.
Pointnet++: Deep hierarchical feature learning on point sets in a metric space
C. R. Qi, L. Yi, H. Su, and L. J. Guibas · 2017
Earlier work this paper cites.
J. Mahler, J. Liang, S. Niyaz, M. Laskey, R. Doan, X. Liu, J. A. Ojea, and K. Goldberg · 2017
Earlier work this paper cites.
Prospection: Interpretable plans from language by predicting the future
C. Paxton, Y. Bisk, J. Thomason, A. Byravan, and D. Foxl · 2019
Earlier work this paper cites.
““Good robot!”: Efficient reinforcement learning for multi-step visual tasks with sim to real transfer
A. Hundt, B. Killeen, N. Greene, H. Wu, H. Kwon, C. Paxton, and G. D. Hager · 2020
Earlier work this paper cites.
6-dof grasping for target-driven object manipulation in clutter
A. Murali, A. Mousavian, C. Eppner, C. Paxton, and D. Fox · 2020
Earlier work this paper cites.
Language models are few-shot learners
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei · 2020
Earlier work this paper cites.
Transporter networks: Rearranging the visual world for robotic manipulation
A. Zeng, P. Florence, J. Tompson, S. Welker, J. Chien, M. Attarian, T. Armstrong, I. Krasin, D. Duong, V. Sindhwani, et al · 2021
Earlier work this paper cites.
Where2act: From pixels to actions for articulated 3d objects
K. Mo, L. J. Guibas, M. Mukadam, A. Gupta, and S. Tulsiani · 2021
Earlier work this paper cites.
An end-to-end transformer model for 3d object detection
I. Misra, R. Girdhar, and A. Joulin · 2021
Earlier work this paper cites.
Vat-mart: Learning visual action trajectory proposals for manipulating 3d articulated objects
R. Wu, Y. Zhao, K. Mo, Z. Guo, Y. Wang, T. Wu, Q. Fan, X. Chen, L. Guibas, and H. Dong · 2021
Earlier work this paper cites.
Umpnet: Universal manipulation policy network for articulated objects
Z. Xu, Z. He, and S. Song · 2021
Earlier work this paper cites.
Unseen object instance segmentation for robotic environments
C. Xie, Y. Xiang, A. Mousavian, and D. Fox · 2021
Earlier work this paper cites.
Learning rgb-d feature embeddings for unseen object instance segmentation
Y. Xiang, C. Xie, A. Mousavian, and D. Fox · 2021
Earlier work this paper cites.
Contact-graspnet: Efficient 6-dof grasp generation in cluttered scenes
M. Sundermeyer, A. Mousavian, R. Triebel, and D. Fox · 2021
Earlier work this paper cites.
Nerp: Neural rearrangement planning for unknown objects
A. H. Qureshi, A. Mousavian, C. Paxton, M. C. Yip, and D. Fox · 2021
Earlier work this paper cites.
Film: Following instructions in language with modular methods
S. Y. Min, D. S. Chaplot, P. Ravikumar, Y. Bisk, and R. Salakhutdinov · 2021
Earlier work this paper cites.
Perceiver io: A general architecture for structured inputs & outputs
A. Jaegle, S. Borgeaud, J.-B. Alayrac, C. Doersch, C. Ionescu, D. Ding, S. Koppula, D. Zoran, A. Brock, E. Shelhamer, et al · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Cited alongside, same era.
Polymetis
Y. Lin, A. S. Wang, G. Sutanto, A. Rai, and F. Meier · 2021
Cited alongside, same era.
Lora: Low-rank adaptation of large language models, 2021
E. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, L. Wang, and W. Chen · 2021
Cited alongside, same era.
Rt-1: Robotics transformer for real-world control at scale
A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, J. Hsu, et al · 2022
Cited alongside, same era.
Grounding language with visual affordances over unstructured data
O. Mees, J. Borja-Diaz, and W. Burgard · 2022
Later among the works it cites.
A persistent spatial semantic representation for high-level natural language instruction execution
V. Blukis, C. Paxton, D. Fox, A. Garg, and Y. Artzi · 2022
Later among the works it cites.
Language models as zero-shot planners: Extracting actionable knowledge for embodied agents
W. Huang, P. Abbeel, D. Pathak, and I. Mordatch · 2022
Later among the works it cites.
The design of stretch: A compact, lightweight mobile manipulator for indoor human environments
C. C. Kemp, A. Edsinger, H. M. Clever, and B. Matulevich · 2022
Later among the works it cites.
Dextreme: Transfer of agile in-hand manipulation from simulation to reality
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Perceiver-actor: A multi-task transformer for robotic manipulation
M. Shridhar, L. Manuelli, and D. Fox · 2022
Cited alongside, same era.
Sornet: Spatial object-centric representations for sequential manipulation
W. Yuan, C. Paxton, K. Desingh, and D. Fox · 2022
Cited alongside, same era.
Cliport: What and where pathways for robotic manipulation
M. Shridhar, L. Manuelli, and D. Fox · 2022
Cited alongside, same era.
Visual language maps for robot navigation
C. Huang, O. Mees, A. Zeng, and W. Burgard · 2022
Cited alongside, same era.
Q-attention: Enabling efficient learning for vision-based robotic manipulation
S. James and A. J. Davison · 2022
Cited alongside, same era.
Do as i can, not as i say: Grounding language in robotic affordances
M. Ahn, A. Brohan, N. Brown, Y. Chebotar, O. Cortes, B. David, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, et al · 2022
Cited alongside, same era.
Instruction-driven history-aware policies for robotic manipulations
P.-L. Guhur, S. Chen, R. Garcia, M. Tapaswi, I. Laptev, and C. Schmid · 2022
Cited alongside, same era.
A. Handa, A. Allshire, V. Makoviychuk, A. Petrenko, R. Singh, J. Liu, D. Makoviichuk, K. Van Wyk, A. Zhurkevich, B. Sundaralingam, et al · 2022
Later among the works it cites.
Clip-fields: Weakly supervised semantic fields for robotic memory
N. M. M. Shafiullah, C. Paxton, L. Pinto, S. Chintala, and A. Szlam · 2022
Later among the works it cites.
Palm-e: An embodied multimodal language model
D. Driess, F. Xia, M. S. Sajjadi, C. Lynch, A. Chowdhery, B. Ichter, A. Wahid, J. Tompson, Q. Vuong, T. Yu, et al · 2023
Closest in time.
Text2motion: From natural language instructions to feasible plans
K. Lin, C. Agia, T. Migimatsu, M. Pavone, and J. Bohg · 2023
Closest in time.
Evaluating continual learning on a home robot, 2023
S. Powers, A. Gupta, and C. Paxton · 2023
Closest in time.
Llama: Open and efficient foundation language models
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, A. Rodriguez, A. Joulin, E. Grave, and G. Lample · 2023
Closest in time.
What learning algorithm is in-context learning? investigations with linear models, 2023
E. Akyürek, D. Schuurmans, J. Andreas, T. Ma, and D. Zhou · 2023
Closest in time.
Homerobot: Open-vocabulary mobile manipulation
S. Yenamandra, A. Ramachandran, K. Yadav, A. Wang, M. Khanna, T. Gervet, T.-Y. Yang, V. Jain, A. W. Clegg, J. Turner, et al · 2023
Closest in time.
In-Context Retrieval-Augmented Language Models, Jan. 2023
O. Ram, Y. Levine, I. Dalmedigos, D. Muhlgay, A. Shashua, K. Leyton-Brown, and Y. Shoham · 2023
Closest in time.
Larger language models do in-context learning differently, Mar. 2023
J. Wei, J. Wei, Y. Tay, D. Tran, A. Webson, Y. Lu, X. Chen, H. Liu, D. Huang, D. Zhou, and T. Ma · 2023
Closest in time.
The Framework Tax: Disparities Between Inference Efficiency in Research and Deployment
J. Fernandez, J. Kahn, C. Na, Y. Bisk, and E. Strubell · 2023
Closest in time.
Diffusion policy: Visuomotor policy learning via action diffusion
C. Chi, S. Feng, Y. Du, Z. Xu, E. Cousineau, B. Burchfiel, and S. Song · 2023
Closest in time.
Usa-net: Unified semantic and affordance representations for robot memory
B. Bolte, A. Wang, J. Yang, M. Mukadam, M. Kalakrishnan, and C. Paxton · 2023
Closest in time.