Fetching the paper…
Reading the bibliography…
Leveraging massive knowledge from large language models (LLMs), recent machine learning models show notable successes in general-purpose task solving in diverse domains such as computer vision and robotics.
An organizing principle for cerebral function: the unit module and the distributed system
Mountcastle, V. B · 1979
Earlier work this paper cites.
A bayesian skill rating system
Graepel, T., Minka, T., and Herbrich, R. T · 2007
Earlier work this paper cites.
Vqa: Visual question answering
Antol, S., Agrawal, A., Lu, J., Mitchell, M., Batra, D., Zitnick, C. L., and Parikh, D · 2015
Earlier work this paper cites.
Human-level concept learning through probabilistic program induction
Lake, B. M., Salakhutdinov, R., and Tenenbaum, J. B · 2015
Earlier work this paper cites.
Scannet: Richly-annotated 3d reconstructions of indoor scenes
Dai, A., Chang, A. X., Savva, M., Halber, M., Funkhouser, T., and Nießner, M · 2017
Earlier work this paper cites.
Building machines that learn and think like people
Lake, B. M., Ullman, T. D., Tenenbaum, J. B., and Gershman, S. J · 2017
Earlier work this paper cites.
Pointnet++: Deep hierarchical feature learning on point sets in a metric space
Qi, C. R., Yi, L., Su, H., and Guibas, L. J · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Kudo, T. and Richardson, J · 2018
Earlier work this paper cites.
Schmidhuber, J · 2018
Earlier work this paper cites.
Sentence-bert: Sentence embeddings using siamese bert-networks
Reimers, N. and Gurevych, I · 2019
Earlier work this paper cites.
Habitat: A platform for embodied ai research
Savva, M., Kadian, A., Maksymets, O., Zhao, Y., Wijmans, E., Jain, B., Straub, J., Liu, J., Koltun, V., Malik, J., et al · 2019
Earlier work this paper cites.
Rio: 3d object instance re-localization in changing indoor environments
Wald, J., Avetisyan, A., Navab, N., Tombari, F., and Nießner, M · 2019
Earlier work this paper cites.
Referit3d: Neural listeners for fine-grained 3d object identification in real-world scenes
Achlioptas, P., Abdelreheem, A., Xia, F., Elhoseiny, M., and Guibas, L · 2020
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Scanrefer: 3d object localization in rgb-d scans using natural language
Chen, D. Z., Chang, A. X., and Nießner, M · 2020
Earlier work this paper cites.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Earlier work this paper cites.
Dark, beyond deep: A paradigm shift to cognitive ai with humanlike common sense
Zhu, Y., Gao, T., Fan, L., Huang, S., Edmonds, M., Liu, H., Gao, F., Zhang, C., Qi, S., Wu, Y. N., et al · 2020
Earlier work this paper cites.
On the opportunities and risks of foundation models
Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M. S., Bohg, J., Bosselut, A., Brunskill, E., et al · 2021
Earlier work this paper cites.
Scan2cap: Context-aware dense captioning in rgb-d scans
Chen, Z., Gholami, A., Nießner, M., and Chang, A. X · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Earlier work this paper cites.
Habitat-matterport 3d dataset (hm3d): 1000 large-scale 3d environments for embodied ai
Ramakrishnan, S. K., Gokaslan, A., Wijmans, E., Maksymets, O., Clegg, A., Turner, J., Undersander, E., Galuba, W., Westbury, A., Chang, A. X., et al · 2021
Earlier work this paper cites.
Cliport: What and where pathways for robotic manipulation
Shridhar, M., Manuelli, L., and Fox, D · 2021
Earlier work this paper cites.
Embodied bert: A transformer model for embodied, language-guided visual task completion
Suglia, A., Gao, Q., Thomason, J., Thattai, G., and Sukhatme, G · 2021
Cited alongside, same era.
Multimodal few-shot learning with frozen language models
Tsimpoukelli, M., Menick, J. L., Cabi, S., Eslami, S., Vinyals, O., and Hill, F · 2021
Cited alongside, same era.
Scenegraphfusion: Incremental 3d scene graph prediction from rgb-d sequences
Wu, S.-C., Wald, J., Tateno, K., Navab, N., and Tombari, F · 2021
Cited alongside, same era.
3dvg-transformer: Relation modeling for visual grounding on point clouds
Zhao, L., Cai, D., Sheng, L., and Xu, D · 2021
Cited alongside, same era.
Do as i can, not as i say: Grounding language in robotic affordances
Ahn, M., Brohan, A., Brown, N., Chebotar, Y., Cortes, O., David, B., Finn, C., Fu, C., Gopalakrishnan, K., Hausman, K., et al · 2022
Cited alongside, same era.
Bang, Y., Cahyawijaya, S., Lee, N., Dai, W., Su, D., Wilie, B., Lovenia, H., Ji, Z., Yu, T., Chung, W., et al · 2023
Closest in time.
Rt-2: Vision-language-action models transfer web knowledge to robotic control
Brohan, A., Brown, N., Carbajal, J., Chebotar, Y., Chen, X., Choromanski, K., Ding, T., Driess, D., Dubey, A., Finn, C., et al · 2023
Closest in time.
End-to-end 3d dense captioning with vote2cap-detr
Chen, S., Zhu, H., Chen, X., Lei, Y., Yu, G., and Chen, T · 2023
Closest in time.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, 2023
Chiang, W.-L., Li, Z., Lin, Z., Sheng, Y., Wu, Z., Zhang, H., Zheng, L., Zhuang, S., Zhuang, Y., Gonzalez, J. E., Stoica, I., and Xing, E. P · 2023
Closest in time.
Instructblip: Towards general-purpose vision-language models with instruction tuning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Flamingo: a visual language model for few-shot learning
Alayrac, J.-B., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., et al · 2022
Cited alongside, same era.
Scanqa: 3d question answering for spatial scene understanding
Azuma, D., Miyanishi, T., Kurita, S., and Kawanabe, M · 2022
Cited alongside, same era.
Rt-1: Robotics transformer for real-world control at scale
Brohan, A., Brown, N., Carbajal, J., Chebotar, Y., Dabis, J., Finn, C., Gopalakrishnan, K., Hausman, K., Herzog, A., Hsu, J., et al · 2022
Cited alongside, same era.
3djcg: A unified framework for joint dense captioning and visual grounding on 3d point clouds
Cai, D., Zhao, L., Zhang, J., Sheng, L., and Xu, D · 2022
Cited alongside, same era.
Language conditioned spatial relation reasoning for 3d object grounding
Chen, S., Guhur, P.-L., Tapaswi, M., Schmid, C., and Laptev, I · 2022
Cited alongside, same era.
Scaling instruction-finetuned language models
Chung, H. W., Hou, L., Longpre, S., Zoph, B., Tay, Y., Fedus, W., Li, E., Wang, X., Dehghani, M., Brahma, S., et al · 2022
Cited alongside, same era.
Minedojo: Building open-ended embodied agents with internet-scale knowledge
Fan, L., Wang, G., Jiang, Y., Mandlekar, A., Yang, Y., Zhu, H., Tang, A., Huang, D.-A., Zhu, Y., and Anandkumar, A · 2022
Cited alongside, same era.
Dai, W., Li, J., Li, D., Tiong, A. M. H., Zhao, J., Wang, W., Li, B., Fung, P., and Hoi, S · 2023
Closest in time.
Objaverse: A universe of annotated 3d objects
Deitke, M., Schwenk, D., Salvador, J., Weihs, L., Michel, O., VanderBilt, E., Schmidt, L., Ehsani, K., Kembhavi, A., and Farhadi, A · 2023
Closest in time.
Palm-e: An embodied multimodal language model
Driess, D., Xia, F., Sajjadi, M. S., Lynch, C., Chowdhery, A., Ichter, B., Wahid, A., Tompson, J., Vuong, Q., Yu, T., et al · 2023
Closest in time.
Llama-adapter v2: Parameter-efficient visual instruction model
Gao, P., Han, J., Zhang, R., Lin, Z., Geng, S., Zhou, A., Zhang, W., Lu, P., He, C., Yue, X., et al · 2023
Closest in time.
From images to textual prompts: Zero-shot vqa with frozen large language models
Guo, J., Li, J., Li, D., Tiong, A. M. H., Li, B., Tao, D., and Hoi, S. C · 2023
Closest in time.
3d-llm: Injecting the 3d world into large language models
Hong, Y., Zhen, H., Chen, P., Zheng, S., Du, Y., Chen, Z., and Gan, C · 2023
Closest in time.
Vima: General robot manipulation with multimodal prompts
Jiang, Y., Gupta, A., Zhang, Z., Wang, G., Dou, Y., Chen, Y., Fei-Fei, L., Anandkumar, A., Zhu, Y., and Fan, L · 2023
Closest in time.
Lerf: Language embedded radiance fields
Kerr, J., Kim, C. M., Goldberg, K., Kanazawa, A., and Tancik, M · 2023
Closest in time.
Kirillov, A., Mintun, E., Ravi, N., Mao, H., Rolland, C., Gustafson, L., Xiao, T., Whitehead, S., Berg, A. C., Lo, W.-Y., et al · 2023
Closest in time.
Unified-io: A unified model for vision, language, and multi-modal tasks
Lu, J., Clark, C., Zellers, R., Mottaghi, R., and Kembhavi, A · 2023
Closest in time.
Scalable 3d captioning with pretrained models
Luo, T., Rockwell, C., Lee, H., and Johnson, J · 2023
Closest in time.
Sqa3d: Situated question answering in 3d scenes
Ma, X., Yong, S., Zheng, Z., Li, Q., Liang, Y., Zhu, S.-C., and Huang, S · 2023
Closest in time.
Embodiedgpt: Vision-language pre-training via embodied chain of thought
Mu, Y., Zhang, Q., Hu, M., Wang, W., Ding, M., Jin, J., Wang, B., Dai, J., Qiao, Y., and Luo, P · 2023
Closest in time.
OpenAI · 2023
Closest in time.
Stanford alpaca: An instruction-following llama model
Taori, R., Gulrajani, I., Zhang, T., Dubois, Y., Li, X., Guestrin, C., Liang, P., and Hashimoto, T. B · 2023
Closest in time.
Pointllm: Empowering large language models to understand point clouds
Xu, R., Wang, X., Wang, T., Chen, Y., Pang, J., and Lin, D · 2023
Closest in time.
mplug-owl: Modularization empowers large language models with multimodality
Ye, Q., Xu, H., Xu, G., Ye, J., Yan, M., Zhou, Y., Wang, J., Hu, A., Shi, P., Shi, Y., et al · 2023
Closest in time.
Lamm: Language-assisted multi-modal instruction-tuning dataset, framework, and benchmark
Yin, Z., Wang, J., Cao, J., Shi, Z., Liu, D., Li, M., Sheng, L., Bai, L., Huang, X., Wang, Z., et al · 2023
Closest in time.
Mmicl: Empowering vision-language model with multi-modal in-context learning
Zhao, H., Cai, Z., Si, S., Ma, X., An, K., Chen, L., Liu, Z., Wang, S., Han, W., and Chang, B · 2023
Closest in time.
Ll3da: Visual interactive instruction tuning for omni-3d understanding, reasoning, and planning
Chen, S., Chen, X., Zhang, C., Li, M., Yu, G., Fei, H., Zhu, H., Fan, J., and Chen, T · 2024
Closest in time.