Fetching the paper…
Reading the bibliography…
Large language models (LLMs) and Vision-Language Models (VLMs) have been proven to excel at multiple tasks, such as commonsense reasoning.
Microsoft coco captions: Data collection and evaluation server
X. Chen, H. Fang, T.-Y. Lin, R. Vedantam, S. Gupta, P. Dollár, and C. L. Zitnick · 2015
Earlier work this paper cites.
Natural language object retrieval, 2016
R. Hu, H. Xu, M. Rohrbach, J. Feng, K. Saenko, and T. Darrell · 2016
Earlier work this paper cites.
Modeling context in referring expressions
L. Yu, P. Poirson, S. Yang, A. C. Berg, and T. L. Berg · 2016
Earlier work this paper cites.
Yin and Yang: Balancing and answering binary visual questions
P. Zhang, Y. Goyal, D. Summers-Stay, D. Batra, and D. Parikh · 2016
Earlier work this paper cites.
ScanNet: Richly-annotated 3D reconstructions of indoor scenes
A. Dai, A. X. Chang, M. Savva, M. Halber, T. A. Funkhouser, and M. Nießner · 2017
Earlier work this paper cites.
Making the V in VQA matter: Elevating the role of image understanding in Visual Question Answering
Y. Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
Earlier work this paper cites.
Deep hough voting for 3d object detection in point clouds
C. Qi, O. Litany, K. He, and L. J. Guibas · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever, et al · 2019
Earlier work this paper cites.
Habitat: A platform for embodied ai research
M. Savva, A. Kadian, O. Maksymets, Y. Zhao, E. Wijmans, B. Jain, J. Straub, J. Liu, V. Koltun, J. Malik, et al · 2019
Earlier work this paper cites.
Dd-ppo: Learning near-perfect pointgoal navigators from 2.5 billion frames
E. Wijmans, A. Kadian, A. Morcos, S. Lee, I. Essa, D. Parikh, M. Savva, and D. Batra · 2019
Earlier work this paper cites.
A fast and accurate one-stage approach to visual grounding, 2019
Z. Yang, B. Gong, L. Wang, W. Huang, D. Yu, and J. Luo · 2019
Earlier work this paper cites.
ReferIt3D: Neural listeners for fine-grained 3D object identification in real-world scenes
P. Achlioptas, A. Abdelreheem, F. Xia, M. Elhoseiny, and L. J. Guibas · 2020
Earlier work this paper cites.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Earlier work this paper cites.
ScanRefer: 3D object localization in RGB-D scans using natural language
D. Z. Chen, A. X. Chang, and M. Nießner · 2020
Earlier work this paper cites.
gradslam: Dense slam meets automatic differentiation
J. Krishna Murthy, S. Saryazdi, G. Iyer, and L. Paull · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu · 2020
Earlier work this paper cites.
Pix2seq: A language modeling framework for object detection
T. Chen, S. Saxena, L. Li, D. J. Fleet, and G. E. Hinton · 2021
Earlier work this paper cites.
Scan2cap: Context-aware dense captioning in rgb-d scans
Z. Chen, A. Gholami, M. Nießner, and A. X. Chang · 2021
Cited alongside, same era.
Free-form description guided 3d visual graph network for object grounding in point cloud
M. Feng, Z. Li, Q. Li, L. Zhang, X. Zhang, G. Zhu, H. Zhang, Y. Wang, and A. S. Mian · 2021
Cited alongside, same era.
Text-guided graph neural networks for referring 3D instance segmentation
P.-H. Huang, H.-H. Lee, H.-T. Chen, and T.-L. Liu · 2021
Cited alongside, same era.
Perceiver: General perception with iterative attention
A. Jaegle, F. Gimeno, A. Brock, A. Zisserman, O. Vinyals, and J. Carreira · 2021
Cited alongside, same era.
Scaling up visual and vision-language representation learning with noisy text supervision
C. Jia, Y. Yang, Y. Xia, Y.-T. Chen, Z. Parekh, H. Pham, Q. Le, Y.-H. Sung, Z. Li, and T. Duerig · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Direct voxel grid optimization: Super-fast convergence for radiance fields reconstruction
C. Sun, M. Sun, and H.-T. Chen · 2022
Later among the works it cites.
Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework
P. Wang, A. Yang, R. Men, J. Lin, S. Bai, Z. Li, J. Ma, C. Zhou, J. Zhou, and H. Yang · 2022
Later among the works it cites.
Habitat-matterport 3D semantics dataset
K. Yadav, R. Ramrakhya, S. K. Ramakrishnan, T. Gervet, J. Turner, A. Gokaslan, N. Maestre, A. X. Chang, D. Batra, M. Savva, et al · 2022
Later among the works it cites.
Openflamingo, Mar. 2023
A. Awadalla, I. Gao, J. Gardner, J. Hessel, Y. Hanafy, W. Zhu, K. Marathe, Y. Bitton, S. Gadre, J. Jitsev, S. Kornblith, P. W. Koh, G. Ilharco, M. Wortsman, and L. Schmidt · 2023
Closest in time.
Free dolly: Introducing the world’s first truly open instruction-tuned llm, 2023
Databricks · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Cited alongside, same era.
Habitat-matterport 3d dataset (hm3d): 1000 large-scale 3d environments for embodied ai
S. K. Ramakrishnan, A. Gokaslan, E. Wijmans, O. Maksymets, A. Clegg, J. Turner, E. Undersander, W. Galuba, A. Westbury, A. X. Chang, et al · 2021
Cited alongside, same era.
Habitat-matterport 3D dataset (HM3D): 1000 large-scale 3D environments for embodied AI
S. K. Ramakrishnan, A. Gokaslan, E. Wijmans, O. Maksymets, A. Clegg, J. M. Turner, E. Undersander, W. Galuba, A. Westbury, A. X. Chang, M. Savva, Y. Zhao, and D. Batra · 2021
Cited alongside, same era.
3D question answering
S. Ye, D. Chen, S. Han, and J. Liao · 2021
Cited alongside, same era.
ScanQA: 3D question answering for spatial scene understanding
D. Azuma, T. Miyanishi, S. Kurita, and M. Kawanabe · 2022
Cited alongside, same era.
Palm: Scaling language modeling with pathways
A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann, et al · 2022
Cited alongside, same era.
Scaling instruction-finetuned language models
H. W. Chung, L. Hou, S. Longpre, B. Zoph, Y. Tay, W. Fedus, E. Li, X. Wang, M. Dehghani, S. Brahma, et al · 2022
Cited alongside, same era.
T. Gong, C. Lyu, S. Zhang, Y. Wang, M. Zheng, Q. Zhao, K. Liu, W. Zhang, P. Luo, and K. Chen · 2023
Closest in time.
3D concept learning and reasoning from multi-view images, 2023
Y. Hong, C. Lin, Y. Du, Z. Chen, J. B. Tenenbaum, and C. Gan · 2023
Closest in time.
3d concept learning and reasoning from multi-view images
Y. Hong, C. Lin, Y. Du, Z. Chen, J. B. Tenenbaum, and C. Gan · 2023
Closest in time.
Visual language maps for robot navigation, 2023
C. Huang, O. Mees, A. Zeng, and W. Burgard · 2023
Closest in time.
Conceptfusion: Open-set multimodal 3D mapping, 2023
K. M. Jatavallabhula, A. Kuwajerwala, Q. Gu, M. Omama, T. Chen, S. Li, G. Iyer, S. Saryazdi, N. Keetha, A. Tewari, J. B. Tenenbaum, C. M. de Melo, M. Krishna, L. Paull, F. Shkurti, and A. Torralba · 2023
Closest in time.
BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models, 2023
J. Li, D. Li, S. Savarese, and S. Hoi · 2023
Closest in time.
J. Li, D. Li, S. Savarese, and S. Hoi · 2023
Closest in time.
H. Liu, C. Li, Q. Wu, and Y. J. Lee · 2023
Closest in time.
GPT-4 technical report, 2023
OpenAI · 2023
Closest in time.
OpenAI · 2023
Closest in time.
Principle-driven self-alignment of language models from scratch with minimal human supervision
Z. Sun, Y. Shen, Q. Zhou, H. Zhang, Z. Chen, D. Cox, Y. Yang, and C. Gan · 2023
Closest in time.
Chatgpt asks, blip-2 answers: Automatic questioning towards enriched visual descriptions, 2023
D. Zhu, J. Chen, K. Haydarov, X. Shen, W. Zhang, and M. Elhoseiny · 2023
Closest in time.