Fetching the paper…
Reading the bibliography…
Humans use different modalities, such as speech, text, images, videos, etc., to communicate their intent and goals with teammates.
The symbol grounding problem
S. Harnad · 1990
Earlier work this paper cites.
Dual coding theory and education
J. M. Clark and A. Paivio · 1991
Earlier work this paper cites.
Intersensory redundancy guides attentional selectivity and perceptual learning in infancy
L. E. Bahrick and R. Lickliter · 2000
Earlier work this paper cites.
Deep residual learning for image recognition, 2015
K. He, X. Zhang, S. Ren, and J. Sun · 2015
Earlier work this paper cites.
Sgdr: Stochastic gradient descent with warm restarts, 2017
I. Loshchilov and F. Hutter · 2017
Earlier work this paper cites.
Zero-shot visual imitation
D. Pathak, P. Mahmoudieh, G. Luo, P. Agrawal, D. Chen, Y. Shentu, E. Shelhamer, J. Malik, A. A. Efros, and T. Darrell · 2018
Earlier work this paper cites.
One-shot imitation from observing humans via domain-adaptive meta-learning
T. Yu, C. Finn, A. Xie, S. Dasari, T. Zhang, P. Abbeel, and S. Levine · 2018
Earlier work this paper cites.
Task-embedded control networks for few-shot imitation learning
S. James, M. Bloesch, and A. J. Davison · 2018
Earlier work this paper cites.
Vision-based multi-task manipulation for inexpensive robots using end-to-end learning from demonstration
R. Rahmatizadeh, P. Abolghasemi, L. Bölöni, and S. Levine · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
Earlier work this paper cites.
Mask r-cnn, 2018
K. He, G. Gkioxari, P. Dollár, and R. Girshick · 2018
Earlier work this paper cites.
Planning with goal-conditioned policies
S. Nasiriany, V. Pong, S. Lin, and S. Levine · 2019
Earlier work this paper cites.
Lxmert: Learning cross-modality encoder representations from transformers
H. Tan and M. Bansal · 2019
Earlier work this paper cites.
Making sense of vision and touch: Self-supervised learning of multimodal representations for contact-rich tasks, 2019
M. A. Lee, Y. Zhu, K. Srinivasan, P. Shah, S. Savarese, L. Fei-Fei, A. Garg, and J. Bohg · 2019
Earlier work this paper cites.
Connecting touch and vision via cross-modal prediction
Y. Li, J.-Y. Zhu, R. Tedrake, and A. Torralba · 2019
Earlier work this paper cites.
Decoupled weight decay regularization, 2019
I. Loshchilov and F. Hutter · 2019
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library, 2019
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, A. Desmaison, A. Köpf, E. Yang, Z. DeVito, M. Raison, A. Tejani, S. Chilamkurthy, B. Steiner, L. Fang, J. Bai, and S. Chintala · 2019
Earlier work this paper cites.
Language-conditioned imitation learning for robot manipulation tasks
S. Stepputtis, J. Campbell, M. Phielipp, S. Lee, C. Baral, and H. Ben Amor · 2020
Earlier work this paper cites.
Learning one-shot imitation from humans without humans
A. Bonardi, S. James, and A. J. Davison · 2020
Earlier work this paper cites.
Uniter: Universal image-text representation learning
Y.-C. Chen, L. Li, L. Yu, A. El Kholy, F. Ahmed, Z. Gan, Y. Cheng, and J. Liu · 2020
Earlier work this paper cites.
Univl: A unified video and language pre-training model for multimodal understanding and generation
H. Luo, L. Ji, B. Shi, H. Huang, N. Duan, T. Li, J. Li, T. Bharti, and M. Zhou · 2020
Earlier work this paper cites.
Language conditioned imitation learning over unstructured data
C. Lynch and P. Sermanet · 2020
Earlier work this paper cites.
Y. Bisk, A. Holtzman, J. Thomason, J. Andreas, Y. Bengio, J. Chai, M. Lapata, A. Lazaridou, J. May, A. Nisnevich, et al · 2020
Cited alongside, same era.
Vokenization: Improving language understanding with contextualized, visual-grounded supervision
H. Tan and M. Bansal · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer, 2020
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu · 2020
Cited alongside, same era.
Concept2robot: Learning manipulation concepts from instructions and human demonstrations
L. Shao, T. Migimatsu, Q. Zhang, K. Yang, and J. Bohg · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
The task specification problem
P. Agrawal · 2022
Later among the works it cites.
Multimae: Multi-modal multi-task masked autoencoders, 2022
R. Bachmann, D. Mizrahi, A. Atanov, and A. Zamir · 2022
Later among the works it cites.
Grounded language-image pre-training, 2022
L. H. Li, P. Zhang, H. Zhang, J. Yang, C. Li, Y. Zhong, L. Wang, L. Yuan, L. Zhang, J.-N. Hwang, K.-W. Chang, and J. Gao · 2022
Later among the works it cites.
Multimodal masked autoencoders learn transferable representations, 2022
X. Geng, H. Liu, L. Lee, D. Schuurmans, S. Levine, and P. Abbeel · 2022
Later among the works it cites.
Groupvit: Semantic segmentation emerges from text supervision, 2022
J. Xu, S. D. Mello, S. Liu, W. Byeon, T. Breuel, J. Kautz, and X. Wang · 2022
Later among the works it cites.
Bridge-prompt: Towards ordinal action understanding in instructional videos, 2022
M. Li, L. Chen, Y. Duan, Z. Hu, J. Feng, J. Zhou, and J. Lu · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Vlm: Task-agnostic video-language model pre-training for video understanding, 2021
H. Xu, G. Ghosh, P.-Y. Huang, P. Arora, M. Aminzadeh, C. Feichtenhofer, F. Metze, and L. Zettlemoyer · 2021
Cited alongside, same era.
Florence: A new foundation model for computer vision, 2021
L. Yuan, D. Chen, Y.-L. Chen, N. Codella, X. Dai, J. Gao, H. Hu, X. Huang, B. Li, C. Li, C. Liu, M. Liu, Z. Liu, Y. Lu, Y. Shi, L. Wang, J. Wang, B. Xiao, Z. Xiao, J. Yang, M. Zeng, L. Zhou, and P. Zhang · 2021
Cited alongside, same era.
Scaling up visual and vision-language representation learning with noisy text supervision, 2021
C. Jia, Y. Yang, Y. Xia, Y.-T. Chen, Z. Parekh, H. Pham, Q. V. Le, Y. Sung, Z. Li, and T. Duerig · 2021
Cited alongside, same era.
Regionclip: Region-based language-image pretraining, 2021
Y. Zhong, J. Yang, P. Zhang, C. Li, N. Codella, L. H. Li, L. Zhou, X. Dai, L. Yuan, Y. Li, and J. Gao · 2021
Cited alongside, same era.
Vidlankd: Improving language understanding via video-distilled knowledge transfer
Z. Tang, J. Cho, H. Tan, and M. Bansal · 2021
Cited alongside, same era.
Learning language-conditioned robot behavior from offline data and crowd-sourced annotation
S. Nair, E. Mitchell, K. Chen, S. Savarese, C. Finn, et al · 2022
Cited alongside, same era.
Rt-1: Robotics transformer for real-world control at scale
A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, J. Hsu, et al · 2022
Cited alongside, same era.
Later among the works it cites.
All in one: Exploring unified video-language pre-training, 2022
A. J. Wang, Y. Ge, R. Yan, Y. Ge, X. Lin, G. Cai, J. Wu, Y. Shan, X. Qie, and M. Z. Shou · 2022
Later among the works it cites.
Lavender: Unifying video-language understanding as masked language modeling, 2022
L. Li, Z. Gan, K. Lin, C.-C. Lin, Z. Liu, C. Liu, and L. Wang · 2022
Later among the works it cites.
Robust speech recognition via large-scale weak supervision
A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever · 2022
Later among the works it cites.
Simmim: A simple framework for masked image modeling
Z. Xie, Z. Zhang, Y. Cao, Y. Lin, J. Bao, Z. Yao, Q. Dai, and H. Hu · 2022
Later among the works it cites.
Viola: Imitation learning for vision-based manipulation with object proposal priors
Y. Zhu, A. Joshi, P. Stone, and Y. Zhu · 2022
Later among the works it cites.
R3m: A universal visual representation for robot manipulation, 2022
S. Nair, A. Rajeswaran, V. Kumar, C. Finn, and A. Gupta · 2022
Later among the works it cites.
Dexterity from touch: Self-supervised pre-training of tactile representations with robotic play, 2023
I. Guzey, B. Evans, S. Chintala, and L. Pinto · 2023
Closest in time.
Goal representations for instruction following: A semi-supervised language interface to control, 2023
V. Myers, A. He, K. Fang, H. Walke, P. Hansen-Estruch, C.-A. Cheng, M. Jalobeanu, A. Kolobov, A. Dragan, and S. Levine · 2023
Closest in time.
Vima: General robot manipulation with multimodal prompts, 2023
Y. Jiang, A. Gupta, Z. Zhang, G. Wang, Y. Dou, Y. Chen, L. Fei-Fei, A. Anandkumar, Y. Zhu, and L. Fan · 2023
Closest in time.
Ulip: Learning unified representation of language, image and point cloud for 3d understanding
L. Xue, M. Gao, C. Xing, R. Martín-Martín, J. Wu, C. Xiong, R. Xu, J. C. Niebles, and S. Savarese · 2023
Closest in time.
Instruction-driven history-aware policies for robotic manipulations
P.-L. Guhur, S. Chen, R. G. Pinel, M. Tapaswi, I. Laptev, and C. Schmid · 2023
Closest in time.
Vip: Towards universal visual reward and representation via value-implicit pre-training, 2023
Y. J. Ma, S. Sodhani, D. Jayaraman, O. Bastani, V. Kumar, and A. Zhang · 2023
Closest in time.
Imagebind: One embedding space to bind them all, 2023
R. Girdhar, A. El-Nouby, Z. Liu, M. Singh, K. V. Alwala, A. Joulin, and I. Misra · 2023
Closest in time.
Multimodality helps unimodality: Cross-modal few-shot learning with multimodal models, 2023
Z. Lin, S. Yu, Z. Kuang, D. Pathak, and D. Ramanan · 2023
Closest in time.
Libero: Benchmarking knowledge transfer for lifelong robot learning, 2023
B. Liu, Y. Zhu, C. Gao, Y. Feng, Q. Liu, Y. Zhu, and P. Stone · 2023
Closest in time.