Long-horizon manipulation of unknown objects via task and motion planning with estimated affordances
A. Curtis, X. Fang, L. P. Kaelbling, T. Lozano-Pérez, and C. R. Garrett · 2022
Later among the works it cites.
The unsurprising effectiveness of pre-trained vision models for control
S. Parisi, A. Rajeswaran, S. Purushwalkam, and A. Gupta · 2022
Later among the works it cites.
R3m: A universal visual representation for robot manipulation
S. Nair, A. Rajeswaran, V. Kumar, C. Finn, and A. Gupta · 2022
Later among the works it cites.
Do as i can, not as i say: Grounding language in robotic affordances
M. Ahn, A. Brohan, N. Brown, Y. Chebotar, O. Cortes, B. David, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, et al · 2022
Later among the works it cites.
Perceiver-actor: A multi-task transformer for robotic manipulation
M. Shridhar, L. Manuelli, and D. Fox · 2022
Later among the works it cites.
Real-world robot learning with masked visual pre-training
I. Radosavovic, T. Xiao, S. James, P. Abbeel, J. Malik, and T. Darrell · 2022
Later among the works it cites.
What matters in language conditioned robotic imitation learning over unstructured data
O. Mees, L. Hermann, and W. Burgard · 2022
Later among the works it cites.
Instruction-following agents with jointly pre-trained vision-language models
Original
H. Liu, L. Lee, K. Lee, and P. Abbeel · 2022
Later among the works it cites.
Cliport: What and where pathways for robotic manipulation
M. Shridhar, L. Manuelli, and D. Fox · 2022
Later among the works it cites.
Simple open-vocabulary object detection with vision transformers
M. Minderer, A. A. Gritsenko, A. Stone, M. Neumann, D. Weissenborn, A. Dosovitskiy, A. Mahendran, A. Arnab, M. Dehghani, Z. Shen, X. Wang, X. Zhai, T. Kipf, and N. Houlsby · 2022
Later among the works it cites.
Open-vocabulary object detection via vision and language knowledge distillation
X. Gu, T.-Y. Lin, W. Kuo, and Y. Cui · 2022
Later among the works it cites.
Regionclip: Region-based language-image pretraining
Y. Zhong, J. Yang, P. Zhang, C. Li, N. Codella, L. H. Li, L. Zhou, X. Dai, L. Yuan, Y. Li, and J. Gao · 2022
Later among the works it cites.
VIP: Towards universal visual reward and representation via value-implicit pre-training
Y. J. Ma, S. Sodhani, D. Jayaraman, O. Bastani, V. Kumar, and A. Zhang · 2023
Closest in time.
RT-1: Robotics transformer for real-world control at scale
A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, J. Hsu, et al · 2023
Closest in time.
VIMA: General robot manipulation with multimodal prompts
Y. Jiang, A. Gupta, Z. Zhang, G. Wang, Y. Dou, Y. Chen, L. Fei-Fei, A. Anandkumar, Y. Zhu, and L. Fan · 2023
Closest in time.
Pali: A jointly-scaled multilingual language-image model
X. Chen, X. Wang, S. Changpinyo, A. Piergiovanni, P. Padlewski, D. Salz, S. Goodman, A. Grycner, B. Mustafa, L. Beyer, et al · 2023
Closest in time.
Cows on pasture: Baselines and benchmarks for language-driven zero-shot object navigation
S. Y. Gadre, M. Wortsman, G. Ilharco, L. Schmidt, and S. Song · 2023
Closest in time.