Fetching the paper…
Reading the bibliography…
A central challenge towards developing robots that can relate human language to their perception and actions is the scarcity of natural language annotations in diverse robot datasets.
A density-based algorithm for discovering clusters in large spatial databases with noise
M. Ester, H.-P. Kriegel, J. Sander, and X. Xu · 1996
Earlier work this paper cites.
Learning latent plans from play
C. Lynch, M. Khansari, T. Xiao, V. Kumar, J. Tompson, S. Levine, and P. Sermanet · 2020
Earlier work this paper cites.
CLIP2Video: Mastering Video-Text Retrieval via Image CLIP, June 2021
H. Fang, P. Xiong, L. Xu, and Y. Chen · 2021
Earlier work this paper cites.
Language Conditioned Imitation Learning over Unstructured Data, July 2021
C. Lynch and P. Sermanet · 2021
Earlier work this paper cites.
Calvin: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks
O. Mees, L. Hermann, E. Rosete-Beas, and W. Burgard · 2022
Earlier work this paper cites.
Image segmentation using text and image prompts
T. Lüddecke and A. Ecker · 2022
Earlier work this paper cites.
CLIP-Fields: Weakly Supervised Semantic Fields for Robotic Memory
N. M. M. Shafiullah, C. Paxton, L. Pinto, S. Chintala, and A. Szlam · 2022
Earlier work this paper cites.
Do as i can, not as i say: Grounding language in robotic affordances
M. Ahn, A. Brohan, N. Brown, Y. Chebotar, O. Cortes, B. David, C. Finn, C. Fu, K. Gopalakrishnan, K. Hausman, et al · 2022
Earlier work this paper cites.
Gmflow: Learning optical flow via global matching
H. Xu, J. Zhang, J. Cai, H. Rezatofighi, and D. Tao · 2022
Earlier work this paper cites.
X-Clip: End-to-end multi-grained contrastive learning for video-text retrieval
Y. Ma, G. Xu, X. Sun, M. Yan, J. Zhang, and R. Ji · 2022
Earlier work this paper cites.
Affordance learning from play for sample-efficient policy learning
J. Borja-Diaz, O. Mees, G. Kalweit, L. Hermann, J. Boedecker, and W. Burgard · 2022
Earlier work this paper cites.
Q-attention: Enabling efficient learning for vision-based robotic manipulation
S. James and A. J. Davison · 2022
Earlier work this paper cites.
Language models with image descriptors are strong few-shot video-language learners
Z. Wang, M. Li, R. Xu, L. Zhou, J. Lei, X. Lin, S. Wang, Z. Yang, C. Zhu, D. Hoiem, et al · 2022
Earlier work this paper cites.
Xmem: Long-term video object segmentation with an atkinson-shiffrin memory model
H. K. Cheng and A. G. Schwing · 2022
Earlier work this paper cites.
What matters in language conditioned robotic imitation learning over unstructured data
O. Mees, L. Hermann, and W. Burgard · 2022
Earlier work this paper cites.
Bridgedata v2: A dataset for robot learning at scale
H. R. Walke, K. Black, T. Z. Zhao, Q. Vuong, C. Zheng, P. Hansen-Estruch, A. W. He, V. Myers, M. J. Kim, M. Du, et al · 2023
Earlier work this paper cites.
Interactive language: Talking to robots in real time
C. Lynch, A. Wahid, J. Tompson, T. Ding, J. Betker, R. Baruch, T. Armstrong, and P. Florence · 2023
Earlier work this paper cites.
RT-1: robotics transformer for real-world control at scale
A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, J. Hsu, J. Ibarz, B. Ichter, A. Irpan, T. Jackson, S. Jesmonth, N. J. Joshi, R. Julian, D. Kalashnikov, Y. Kuang, I. Leal, K. Lee, S. Levine, Y. Lu, U. Malla, D. Manjunath, I. Mordatch, O. Nachum, C. Parada, J. Peralta, E. Perez, K. Pertsch, J. Quiambao, K. Rao, M. S. Ryoo, G. Salazar, P. R. Sanketi, K. Sayed, J. Singh, S. Sontakke, A. Stone, C. Tan, H. T. Tran, V. Vanhoucke, S. Vega, Q. Vuong, F. Xia, T. Xiao, P. Xu, S. Xu, T. Yu, and B. Zitkovich · 2023
Earlier work this paper cites.
Scaling Robot Learning with Semantically Imagined Experience, Feb. 2023
T. Yu, T. Xiao, A. Stone, J. Tompson, A. Brohan, S. Wang, J. Singh, C. Tan, D. M, J. Peralta, et al · 2023
Earlier work this paper cites.
Sigmoid loss for language image pre-training
X. Zhai, B. Mustafa, A. Kolesnikov, and L. Beyer · 2023
Earlier work this paper cites.
Open-vocabulary queryable scene representations for real world planning
B. Chen, F. Xia, B. Ichter, K. Rao, K. Gopalakrishnan, M. S. Ryoo, A. Stone, and D. Kappler · 2023
Earlier work this paper cites.
Visual language maps for robot navigation
C. Huang, O. Mees, A. Zeng, and W. Burgard · 2023
Earlier work this paper cites.
Audio visual language maps for robot navigation
C. Huang, O. Mees, A. Zeng, and W. Burgard · 2023
Cited alongside, same era.
Grounding dino: Marrying dino with grounded pre-training for open-set object detection, 2023
S. Liu, Z. Zeng, T. Ren, F. Li, H. Zhang, J. Yang, C. Li, J. Yang, H. Su, J. Zhu, et al · 2023
Cited alongside, same era.
Tracking anything with decoupled video segmentation
H. K. Cheng, S. W. Oh, B. Price, A. Schwing, and J.-Y. Lee · 2023
Cited alongside, same era.
Robots that ask for help: Uncertainty alignment for large language model planners
A. Z. Ren, A. Dixit, A. Bodrova, S. Singh, S. Tu, N. Brown, P. Xu, L. T. Takayama, F. Xia, J. Varley, et al · 2023
Cited alongside, same era.
REFLECT: Summarizing robot experiences for failure explanation and correction
Z. Liu, A. Bahety, and S. Song · 2023
Cited alongside, same era.
Gemini: A family of highly capable multimodal models, 2023
Latent plans for task-agnostic offline reinforcement learning
E. Rosete-Beas, O. Mees, G. Kalweit, J. Boedecker, and W. Burgard · 2023
Later among the works it cites.
Droid: A large-scale in-the-wild robot manipulation dataset
A. Khazatsky, K. Pertsch, S. Nair, A. Balakrishna, S. Dasari, S. Karamcheti, S. Nasiriany, M. K. Srirama, L. Y. Chen, K. Ellis, et al · 2024
Closest in time.
Multimodal diffusion transformer: Learning versatile behavior from multimodal goals
M. Reuss, Ö. E. Yağmurlu, F. Wenzel, and R. Lioutikov · 2024
Closest in time.
Ok-robot: What really matters in integrating open-knowledge models for robotics
P. Liu, Y. Orru, C. Paxton, N. M. M. Shafiullah, and L. Pinto · 2024
Closest in time.
Vision-language models provide promptable representations for reinforcement learning
W. Chen, O. Mees, A. Kumar, and S. Levine · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
G. Team · 2023
Cited alongside, same era.
Video-llava: Learning united visual representation by alignment before projection, 2023
B. Lin, Y. Ye, B. Zhu, J. Cui, M. Ning, P. Jin, and L. Yuan · 2023
Cited alongside, same era.
Octo: An open-source generalist robot policy
Octo Model Team, D. Ghosh, H. Walke, K. Pertsch, K. Black, O. Mees, S. Dasari, J. Hejna, C. Xu, J. Luo, et al · 2023
Cited alongside, same era.
Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V, Nov. 2023
J. Yang, H. Zhang, F. Li, X. Zou, C. Li, and J. Gao · 2023
Cited alongside, same era.
Grounding language with visual affordances over unstructured data, 2023
O. Mees, J. Borja-Diaz, and W. Burgard · 2023
Cited alongside, same era.
Perceiver-actor: A multi-task transformer for robotic manipulation
M. Shridhar, L. Manuelli, and D. Fox · 2023
Cited alongside, same era.
Vid2seq: Large-scale pretraining of a visual language model for dense video captioning
A. Yang, A. Nagrani, P. H. Seo, A. Miech, J. Pont-Tuset, I. Laptev, J. Sivic, and C. Schmid · 2023
Cited alongside, same era.
RobotVQA: Multimodal long-horizon reasoning for robotics
P. Sermanet, T. Ding, J. Zhao, F. Xia, D. Dwibedi, K. Gopalakrishnan, C. Chan, G. Dulac-Arnold, S. Maddineni, N. J. Joshi, et al · 2024
Closest in time.
Robotic control via embodied chain-of-thought reasoning
M. Zawalski, W. Chen, K. Pertsch, O. Mees, C. Finn, and S. Levine · 2024
Closest in time.
Lelan: Learning a language-conditioned navigation policy from in-the-wild video
N. Hirose, C. Glossop, A. Sridhar, D. Shah, O. Mees, and S. Levine · 2024
Closest in time.
Spatialvlm: Endowing vision-language models with spatial reasoning capabilities
B. Chen, Z. Xu, S. Kirmani, B. Ichter, D. Sadigh, L. Guibas, and F. Xia · 2024
Closest in time.
Universal visual decomposer: Long-horizon manipulation made easy
Z. Zhang, Y. Li, O. Bastani, A. Gupta, D. Jayaraman, Y. J. Ma, and L. Weihs · 2024
Closest in time.
Evaluating real-world robot manipulation policies in simulation
X. Li, K. Hsu, J. Gu, K. Pertsch, O. Mees, H. R. Walke, C. Fu, I. Lunawat, I. Sieh, S. Kirmani, et al · 2024
Closest in time.
Depth anything: Unleashing the power of large-scale unlabeled data
L. Yang, B. Kang, Z. Huang, X. Xu, J. Feng, and H. Zhao · 2024
Closest in time.
Mementos: A comprehensive benchmark for multimodal large language model reasoning over image sequences, 2024
X. Wang, Y. Zhou, X. Liu, H. Lu, Y. Xu, F. He, J. Yoon, T. Lu, G. Bertasius, M. Bansal, et al · 2024
Closest in time.
”Task Success” is not enough: Investigating the use of video-language models as behavior critics for catching undesirable agent behaviors, 2024
L. Guan, Y. Zhou, D. Liu, Y. Zha, H. B. Amor, and S. Kambhampati · 2024
Closest in time.
Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language Models, Feb. 2024
S. Karamcheti, S. Nair, A. Balakrishna, P. Liang, T. Kollar, and D. Sadigh · 2024
Closest in time.
Scaling open-vocabulary object detection
M. Minderer, A. Gritsenko, and N. Houlsby · 2024
Closest in time.
SPRINT: Scalable policy pre-training via language instruction relabeling
J. Zhang, K. Pertsch, J. Zhang, and J. J. Lim · 2024
Closest in time.
EfficientSAM: Leveraged masked image pretraining for efficient segment anything
Y. Xiong, B. Varadarajan, L. Wu, X. Xiang, F. Xiao, C. Zhu, X. Dai, D. Wang, F. Sun, F. Iandola, et al · 2024
Closest in time.
Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0
A. O’Neill, A. Rehman, A. Maddukuri, A. Gupta, A. Padalkar, A. Lee, A. Pooley, A. Gupta, A. Mandlekar, A. Jain, et al · 2024
Closest in time.
Physically grounded vision-language models for robotic manipulation
J. Gao, B. Sarkar, F. Xia, T. Xiao, J. Wu, B. Ichter, A. Majumdar, and D. Sadigh · 2024
Closest in time.
Roboclip: One demonstration is enough to learn robot policies
S. Sontakke, J. Zhang, S. Arnold, K. Pertsch, E. Bıyık, D. Sadigh, C. Finn, and L. Itti · 2024
Closest in time.