Fetching the paper…
Reading the bibliography…
Pretrained vision language models (VLMs) present an opportunity to caption unlabeled 3D objects at scale.
Three-dimensional object recognition
Besl, P. J. and Jain, R. C · 1985
Earlier work this paper cites.
Wordnet: a lexical database for english
Miller, G. A · 1995
Earlier work this paper cites.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Wu, Y., Schuster, M., Chen, Z., Le, Q. V., Norouzi, M., Macherey, W., Krikun, M., Cao, Y., Gao, Q., Macherey, K., et al · 2016
Earlier work this paper cites.
Revisiting unreasonable effectiveness of data in deep learning era
Sun, C., Shrivastava, A., Singh, S., and Gupta, A · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Cer, D., Yang, Y., Kong, S.-y., Hua, N., Limtiaco, N., John, R. S., Constant, N., Guajardo-Cespedes, M., Yuan, S., Tar, C., et al · 2018
Earlier work this paper cites.
Blender - a 3D modelling and rendering package
Community, B. O · 2018
Earlier work this paper cites.
Compiling machine learning programs via high-level tracing
Frostig, R., Johnson, M. J., and Leary, C · 2018
Earlier work this paper cites.
Iqa: Visual question answering in interactive environments
Gordon, D., Kembhavi, A., Rastegari, M., Redmon, J., Fox, D., and Farhadi, A · 2018
Earlier work this paper cites.
Defining textual entailment
Korman, D. Z., Mack, E., Jett, J., and Renear, A. H · 2018
Earlier work this paper cites.
Kudo, T. and Richardson, J · 2018
Earlier work this paper cites.
3d scene graph: A structure for unified semantics, 3d space, and camera
Armeni, I., He, Z.-Y., Gwak, J., Zamir, A. R., Fischer, M., Malik, J., and Savarese, S · 2019
Earlier work this paper cites.
Lvis: A dataset for large vocabulary instance segmentation
Gupta, A., Dollar, P., and Girshick, R · 2019
Earlier work this paper cites.
Visualbert: A simple and performant baseline for vision and language
Li, L. H., Yatskar, M., Yin, D., Hsieh, C.-J., and Chang, K.-W · 2019
Earlier work this paper cites.
Auto-encoding scene graphs for image captioning
Yang, X., Tang, K., Zhang, H., and Cai, J · 2019
Earlier work this paper cites.
Objectnav revisited: On evaluation of embodied agents navigating to objects
Batra, D., Gokaslan, A., Kembhavi, A., Maksymets, O., Mottaghi, R., Savva, M., Toshev, A., and Wijmans, E · 2020
Earlier work this paper cites.
Say as you wish: Fine-grained control of image caption generation with abstract scene graphs
Chen, S., Jin, Q., Wang, P., and Wu, Q · 2020
Earlier work this paper cites.
Shapecaptioner: Generative caption network for 3d shapes by learning a mapping from parts detected in multiple views to sentences
Han, Z., Chen, C., Liu, Y.-S., and Zwicker, M · 2020
Earlier work this paper cites.
The origins and prevalence of texture bias in convolutional neural networks
Hermann, K., Chen, T., and Kornblith, S · 2020
Cited alongside, same era.
The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale
Kuznetsova, A., Rom, H., Alldrin, N., Uijlings, J., Krasin, I., Pont-Tuset, J., Kamali, S., Popov, S., Malloci, M., Kolesnikov, A., et al · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Cited alongside, same era.
Learning 3d semantic scene graphs from 3d indoor reconstructions
Wald, J., Dhamo, H., Navab, N., and Tombari, F · 2020
Cited alongside, same era.
Video object segmentation and tracking: A survey
Yao, R., Lin, G., Xia, S., Zhao, J., and Zhou, Y · 2020
Cited alongside, same era.
Scaling vision transformers
Zhai, X., Kolesnikov, A., Houlsby, N., and Beyer, L · 2022
Later among the works it cites.
Instructblip: Towards general-purpose vision-language models with instruction tuning, 2023
Dai, W., Li, J., Li, D., Tiong, A. M. H., Zhao, J., Wang, W., Li, B., Fung, P., and Hoi, S · 2023
Closest in time.
Scannerf: a scalable benchmark for neural radiance fields
De Luigi, L., Bolognini, D., Domeniconi, F., De Gregorio, D., Poggi, M., and Di Stefano, L · 2023
Closest in time.
Scaling vision transformers to 22 billion parameters
Dehghani, M., Djolonga, J., Mustafa, B., Padlewski, P., Heek, J., Gilmer, J., Steiner, A. P., Caron, M., Geirhos, R., Alabdulmohsin, I., et al · 2023
Closest in time.
Objaverse: A universe of annotated 3d objects
Deitke, M., Schwenk, D., Salvador, J., Weihs, L., Michel, O., VanderBilt, E., Schmidt, L., Ehsani, K., Kembhavi, A., and Farhadi, A · 2023
Closest in time.
Eva: Exploring the limits of masked visual representation learning at scale
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
He, Y., Yu, H., Liu, X., Yang, Z., Sun, W., Wang, Y., Fu, Q., Zou, Y., and Mian, A · 2021
Cited alongside, same era.
Scaling up visual and vision-language representation learning with noisy text supervision
Jia, C., Yang, Y., Xia, Y., Chen, Y.-T., Parekh, Z., Pham, H., Le, Q., Sung, Y.-H., Li, Z., and Duerig, T · 2021
Cited alongside, same era.
Intriguing properties of vision transformers
Naseer, M. M., Ranasinghe, K., Khan, S. H., Hayat, M., Shahbaz Khan, F., and Yang, M.-H · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Cited alongside, same era.
Open-vocabulary object detection using captions
Zareian, A., Rosa, K. D., Hu, D. H., and Chang, S.-F · 2021
Cited alongside, same era.
Flamingo: a visual language model for few-shot learning
Alayrac, J.-B., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., et al · 2022
Cited alongside, same era.
Scanqa: 3d question answering for spatial scene understanding. 2022 ieee
Azuma, D., Miyanishi, T., Kurita, S., and Kawanabe, M · 2022
Cited alongside, same era.
Fang, Y., Wang, W., Xie, B., Sun, Q., Wu, L., Wang, X., Huang, T., Wang, X., and Cao, Y · 2023
Closest in time.
Physically grounded vision-language models for robotic manipulation
Gao, J., Sarkar, B., Xia, F., Xiao, T., Wu, J., Ichter, B., Majumdar, A., and Sadigh, D · 2023
Closest in time.
Conceptgraphs: Open-vocabulary 3d scene graphs for perception and planning, 2023
Gu, Q., Kuwajerwala, A., Morin, S., Jatavallabhula, K. M., Sen, B., Agarwal, A., Rivera, C., Paul, W., Ellis, K., Chellappa, R., Gan, C., de Melo, C. M., Tenenbaum, J. B., Torralba, A., Shkurti, F., and Paull, L · 2023
Closest in time.
Visual programming: Compositional visual reasoning without training
Gupta, T. and Kembhavi, A · 2023
Closest in time.
3d-llm: Injecting the 3d world into large language models
Hong, Y., Zhen, H., Chen, P., Zheng, S., Du, Y., Chen, Z., and Gan, C · 2023
Closest in time.
Li, J., Li, D., Savarese, S., and Hoi, S · 2023
Closest in time.
Scalable 3d captioning with pretrained models
Luo, T., Rockwell, C., Lee, H., and Johnson, J · 2023
Closest in time.
Approaching human 3d shape perception with neurally mappable models
O’Connell, T. P., Bonnen, T., Friedman, Y., Tewari, A., Tenenbaum, J. B., Sitzmann, V., and Kanwisher, N · 2023
Closest in time.
Gpt-4 technical report, 2023
OpenAI · 2023
Closest in time.
Vipergpt: Visual inference via python execution for reasoning
Surís, D., Menon, S., and Vondrick, C · 2023
Closest in time.
Omniobject3d: Large-vocabulary 3d object dataset for realistic perception, reconstruction and generation
Wu, T., Zhang, J., Fu, X., Wang, Y., Ren, J., Pan, L., Wu, W., Yang, L., Wang, J., Qian, C., et al · 2023
Closest in time.
Chatgpt asks, blip-2 answers: Automatic questioning towards enriched visual descriptions
Zhu, D., Chen, J., Haydarov, K., Shen, X., Zhang, W., and Elhoseiny, M · 2023
Closest in time.