Fetching the paper…
Reading the bibliography…
Recent vision-language-action (VLA) models rely on 2D inputs, lacking integration with the broader realm of the 3D physical world.
The effects of contextual scenes on the identification of objects
Palmer, S · 1975
Earlier work this paper cites.
Seeing and Visualizing: It’s Not What You Think
Pylyshyn, Z · 2003
Earlier work this paper cites.
Vision: A Computational Investigation into the Human Representation and Processing of Visual Information
Marr, D · 2010
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma, D. P. and Welling, M · 2013
Earlier work this paper cites.
Scannet: Richly-annotated 3d reconstructions of indoor scenes, 2017
Dai, A., Chang, A. X., Savva, M., Halber, M., Funkhouser, T., and Nießner, M · 2017
Earlier work this paper cites.
spaCy 2: Natural language understanding with Bloom embeddings, convolutional neural networks and incremental parsing
Honnibal, M. and Montani, I · 2017
Earlier work this paper cites.
Scaling egocentric vision: The epic-kitchens dataset
Damen, D., Doughty, H., Farinella, G. M., Fidler, S., Furnari, A., Kazakos, E., Moltisanti, D., Munro, J., Perrett, T., Price, W., et al · 2018
Earlier work this paper cites.
Scaling robot supervision to hundreds of hours with roboturk: Robotic manipulation dataset through human reasoning and dexterity
Mandlekar, A., Booher, J., Spero, M., Tung, A., Gupta, A., Zhu, Y., Garg, A., Savarese, S., and Fei-Fei, L · 2019
Earlier work this paper cites.
Rlbench: The robot learning benchmark & learning environment
James, S., Ma, Z., Arrojo, D. R., and Davison, A. J · 2020
Earlier work this paper cites.
Language conditioned imitation learning over unstructured data
Lynch, C. and Sermanet, P · 2020
Earlier work this paper cites.
Raft: Recurrent all-pairs field transforms for optical flow
Teed, Z. and Deng, J · 2020
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2021
Earlier work this paper cites.
Habitat-matterport 3d dataset (hm3d): 1000 large-scale 3d environments for embodied ai, 2021
Ramakrishnan, S. K., Gokaslan, A., Wijmans, E., Maksymets, O., Clegg, A., Turner, J., Undersander, E., Galuba, W., Westbury, A., Chang, A. X., Savva, M., Zhao, Y., and Batra, D · 2021
Earlier work this paper cites.
Playing with food: Learning food item representations through interactive exploration
Sawhney, A., Lee, S., Zhang, K., Veloso, M., and Kroemer, O · 2021
Earlier work this paper cites.
Lancon-learn: Learning with language to enable generalization in multi-task manipulation
Silva, A., Moorman, N., Silva, W., Zaidi, Z., Gopalan, N., and Gombolay, M · 2021
Earlier work this paper cites.
Flamingo: a visual language model for few-shot learning
Alayrac, J.-B., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., et al · 2022
Earlier work this paper cites.
Rt-1: Robotics transformer for real-world control at scale
Brohan, A., Brown, N., Carbajal, J., Chebotar, Y., Dabis, J., Finn, C., Gopalakrishnan, K., Hausman, K., Herzog, A., Hsu, J., et al · 2022
Earlier work this paper cites.
Objaverse: A universe of annotated 3d objects, 2022
Deitke, M., Schwenk, D., Salvador, J., Weihs, L., Michel, O., VanderBilt, E., Schmidt, L., Ehsani, K., Kembhavi, A., and Farhadi, A · 2022
Earlier work this paper cites.
Bc-z: Zero-shot task generalization with robotic imitation learning
Jang, E., Irpan, A., Khansari, M., Kappler, D., Ebert, F., Lynch, C., Levine, S., and Finn, C · 2022
Cited alongside, same era.
Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Li, J., Li, D., Xiong, C., and Hoi, S · 2022
Cited alongside, same era.
Hoi4d: A 4d egocentric dataset for category-level human-object interaction
Liu, Y., Liu, Y., Jiang, C., Lyu, K., Wan, W., Shen, H., Liang, B., Fu, Z., Wang, H., and Yi, L · 2022
Cited alongside, same era.
Calvin: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks
Mees, O., Hermann, L., Rosete-Beas, E., and Burgard, W · 2022
Cited alongside, same era.
Point-e: A system for generating 3d point clouds from complex prompts
Nichol, A., Jun, H., Dhariwal, P., Mishkin, P., and Chen, M · 2022
Cited alongside, same era.
UNIFIED-IO: A unified model for vision, language, and multi-modal tasks
Lu, J., Clark, C., Zellers, R., Mottaghi, R., and Kembhavi, A · 2023
Later among the works it cites.
Interactive language: Talking to robots in real time
Lynch, C., Wahid, A., Tompson, J., Ding, T., Betker, J., Baruch, R., Armstrong, T., and Florence, P · 2023
Later among the works it cites.
Grounding language with visual affordances over unstructured data
Mees, O., Borja-Diaz, J., and Burgard, W · 2023
Later among the works it cites.
Open x-embodiment: Robotic learning datasets and rt-x models
Padalkar, A., Pooley, A., Jain, A., Bewley, A., Herzog, A., Irpan, A., Khazatsky, A., Rai, A., Singh, A., Brohan, A., et al · 2023
Later among the works it cites.
Kosmos-2: Grounding multimodal large language models to the world
Peng, Z., Wang, W., Dong, L., Hao, Y., Huang, S., Ma, S., and Wei, F · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
High-resolution image synthesis with latent diffusion models
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B · 2022
Cited alongside, same era.
Zoedepth: Zero-shot transfer by combining relative and metric depth
Bhat, S. F., Birkl, R., Wofk, D., Wonka, P., and Müller, M · 2023
Cited alongside, same era.
Zero-shot robotic manipulation with pretrained image-editing diffusion models, 2023
Black, K., Nakamoto, M., Atreya, P., Walke, H., Finn, C., Kumar, A., and Levine, S · 2023
Cited alongside, same era.
Rt-2: Vision-language-action models transfer web knowledge to robotic control
Brohan, A., Brown, N., Carbajal, J., Chebotar, Y., Chen, X., Choromanski, K., Ding, T., Driess, D., Dubey, A., Finn, C., et al · 2023
Cited alongside, same era.
Instructpix2pix: Learning to follow image editing instructions, 2023
Brooks, T., Holynski, A., and Efros, A. A · 2023
Cited alongside, same era.
Clvr jaco play dataset, 2023
Dass, S., Yapeter, J., Zhang, J., Zhang, J., Pertsch, K., Nikolaidis, S., and Lim, J. J · 2023
Cited alongside, same era.
Dreamllm: Synergistic multimodal comprehension and creation
Dong, R., Han, C., Peng, Y., Qi, Z., Ge, Z., Yang, J., Zhao, L., Sun, J., Zhou, H., Wei, H., Kong, X., Zhang, X., Ma, K., and Yi, L · 2023
Cited alongside, same era.
Gpt4point: A unified framework for point-language understanding and generation, 2023
Qi, Z., Fang, Y., Sun, Z., Wu, X., Wu, T., Wang, J., Lin, D., and Zhao, H · 2023
Later among the works it cites.
Robovqa: Multimodal long-horizon reasoning for robotics
Sermanet, P., Ding, T., Zhao, J., Xia, F., Dwibedi, D., Gopalakrishnan, K., Chan, C., Dulac-Arnold, G., Maddineni, S., Joshi, N. J., Florence, P., Han, W., Baruch, R., Lu, Y., Mirchandani, S., Xu, P., Sanketi, P., Hausman, K., Shafran, I., Ichter, B., and Cao, Y · 2023
Later among the works it cites.
Shafiullah, N. M. M., Rai, A., Etukuru, H., Liu, Y., Misra, I., Chintala, S., and Pinto, L · 2023
Later among the works it cites.
MUTEX: Learning unified policies from multimodal task specifications
Shah, R., Martín-Martín, R., and Zhu, Y · 2023
Later among the works it cites.
Bridgedata v2: A dataset for robot learning at scale
Walke, H. R., Black, K., Zhao, T. Z., Vuong, Q., Zheng, C., Hansen-Estruch, P., He, A. W., Myers, V., Kim, M. J., Du, M., et al · 2023
Later among the works it cites.
Next-gpt: Any-to-any multimodal llm
Wu, S., Fei, H., Qu, L., Ji, W., and Chua, T.-S · 2023
Later among the works it cites.
Pointllm: Empowering large language models to understand point clouds, 2023
Xu, R., Wang, X., Wang, T., Chen, Y., Pang, J., and Lin, D · 2023
Later among the works it cites.
Uni3d: Exploring unified 3d representation at scale, 2023
Zhou, J., Wang, J., Ma, B., Liu, Y.-S., Huang, T., and Wang, X · 2023
Later among the works it cites.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
Zhu, D., Chen, J., Shen, X., Li, X., and Elhoseiny, M · 2023
Later among the works it cites.
Multiply: A multisensory object-centric embodied large language model in 3d world
Hong, Y., Zheng, Z., Chen, P., Wang, Y., Li, J., and Gan, C · 2024
Closest in time.
3dmit: 3d multi-modal instruction tuning for scene understanding, 2024
Li, Z., Zhang, C., Wang, X., Ren, R., Xu, Y., Ma, R., and Liu, X · 2024
Closest in time.
Grounded sam: Assembling open-world models for diverse visual tasks, 2024
Ren, T., Liu, S., Zeng, A., Lin, J., Li, K., Cao, H., Chen, J., Huang, X., Chen, Y., Yan, F., Zeng, Z., Zhang, H., Li, F., Yang, J., Li, H., Jiang, Q., and Zhang, L · 2024
Closest in time.
Playfusion: Skill acquisition via diffusion from language-annotated play
Chen, L., Bahl, S., and Pathak, D · 2029
Closest in time.