Fetching the paper…
Reading the bibliography…
In this work, we investigate extending the comprehension of Multi-modal Large Language Models (MLLMs) to regional objects.
Unifying Vision-and-Language Tasks via Text Generation
Cho, J.; Lei, J.; Tan, H.; and Bansal, M. 2021 · 1942
Earlier work this paper cites.
Microsoft COCO: Common Objects in Context
Lin, T.; Maire, M.; Belongie, S. J.; Hays, J.; Perona, P.; Ramanan, D.; Dollár, P.; and Zitnick, C. L. 2014 · 2014
Earlier work this paper cites.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
Plummer, B. A.; Wang, L.; Cervantes, C. M.; Caicedo, J. C.; Hockenmaier, J.; and Lazebnik, S. 2015 · 2015
Earlier work this paper cites.
Generation and comprehension of unambiguous object descriptions
Mao, J.; Huang, J.; Toshev, A.; Camburu, O.; Yuille, A. L.; and Murphy, K. 2016 · 2016
Earlier work this paper cites.
Modeling context in referring expressions
Yu, L.; Poirson, P.; Yang, S.; Berg, A. C.; and Berg, T. L. 2016 · 2016
Earlier work this paper cites.
ScanNet: Richly-Annotated 3D Reconstructions of Indoor Scenes
Dai, A.; Chang, A. X.; Savva, M.; Halber, M.; Funkhouser, T. A.; and Nießner, M. 2017 · 2017
Earlier work this paper cites.
Mask R-CNN
He, K.; Gkioxari, G.; Dollár, P.; and Girshick, R. B. 2017 · 2017
Earlier work this paper cites.
Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations
Krishna, R.; Zhu, Y.; Groth, O.; Johnson, J.; Hata, K.; Kravitz, J.; Chen, S.; Kalantidis, Y.; Li, L.-J.; Shamma, D. A.; Bernstein, M. S.; and Fei-Fei, L. 2017 · 2017
Earlier work this paper cites.
Conceptual Captions: A Cleaned, Hypernymed, Image Alt-text Dataset For Automatic Image Captioning
Sharma, P.; Ding, N.; Goodman, S.; and Soricut, R. 2018 · 2018
Earlier work this paper cites.
Decoupled Weight Decay Regularization
Loshchilov, I.; and Hutter, F. 2019 · 2019
Earlier work this paper cites.
ScanRefer: 3D Object Localization in RGB-D Scans Using Natural Language
Chen, D. Z.; Chang, A. X.; and Nießner, M. 2020 · 2020
Earlier work this paper cites.
spaCy: Industrial-strength natural language processing in python
Honnibal, M.; Montani, I.; Van Landeghem, S.; Boyd, A.; et al. 2020 · 2020
Earlier work this paper cites.
Conceptual 12M: Pushing Web-Scale Image-Text Pre-Training To Recognize Long-Tail Visual Concepts
Changpinyo, S.; Sharma, P.; Ding, N.; and Soricut, R. 2021 · 2021
Earlier work this paper cites.
Learning Transferable Visual Models From Natural Language Supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; Krueger, G.; and Sutskever, I. 2021 · 2021
Cited alongside, same era.
Flamingo: a Visual Language Model for Few-Shot Learning
Alayrac, J.; Donahue, J.; Luc, P.; Miech, A.; Barr, I.; Hasson, Y.; Lenc, K.; Mensch, A.; Millican, K.; Reynolds, M.; Ring, R.; Rutherford, E.; Cabi, S.; Han, T.; Gong, Z.; Samangooei, S.; Monteiro, M.; Menick, J. L.; Borgeaud, S.; Brock, A.; Nematzadeh, A.; Sharifzadeh, S.; Binkowski, M.; Barreira, R.; Vinyals, O.; Zisserman, A.; and Simonyan, K. 2022 · 2022
Cited alongside, same era.
Coyo-700m: Image-text pair dataset
Byeon, M.; Park, B.; Kim, H.; Lee, S.; Baek, W.; and Kim, S. 2022 · 2022
Cited alongside, same era.
Scaling Instruction-Finetuned Language Models
Chung, H. W.; Hou, L.; Longpre, S.; Zoph, B.; Tay, Y.; Fedus, W.; Li, E.; Wang, X.; Dehghani, M.; Brahma, S.; Webson, A.; Gu, S. S.; Dai, Z.; Suzgun, M.; Chen, X.; Chowdhery, A.; Narang, S.; Mishra, G.; Yu, A.; Zhao, V. Y.; Huang, Y.; Dai, A. M.; Yu, H.; Petrov, S.; Chi, E. H.; Dean, J.; Devlin, J.; Roberts, A.; Zhou, D.; Le, Q. V.; and Wei, J. 2022 · 2022
Cited alongside, same era.
InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
Dai, W.; Li, J.; Li, D.; Tiong, A. M. H.; Zhao, J.; Wang, W.; Li, B.; Fung, P.; and Hoi, S. 2023 · 2023
Closest in time.
MultiModal-GPT: A Vision and Language Model for Dialogue with Humans
Gong, T.; Lyu, C.; Zhang, S.; Wang, Y.; Zheng, M.; Zhao, Q.; Liu, K.; Zhang, W.; Luo, P.; and Chen, K. 2023 · 2023
Closest in time.
Kirillov, A.; Mintun, E.; Ravi, N.; Mao, H.; Rolland, C.; Gustafson, L.; Xiao, T.; Whitehead, S.; Berg, A. C.; Lo, W.-Y.; et al. 2023 · 2023
Closest in time.
Scalable 3D Captioning with Pretrained Models
Luo, T.; Rockwell, C.; Lee, H.; and Johnson, J. 2023 · 2023
Closest in time.
Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deitke, M.; Schwenk, D.; Salvador, J.; Weihs, L.; Michel, O.; VanderBilt, E.; Schmidt, L.; Ehsani, K.; Kembhavi, A.; and Farhadi, A. 2022 · 2022
Cited alongside, same era.
GLM: General Language Model Pretraining with Autoregressive Blank Infilling
Du, Z.; Qian, Y.; Liu, X.; Ding, M.; Qiu, J.; Yang, Z.; and Tang, J. 2022 · 2022
Cited alongside, same era.
EVA: Exploring the Limits of Masked Visual Representation Learning at Scale
Fang, Y.; Wang, W.; Xie, B.; Sun, Q.; Wu, L.; Wang, X.; Huang, T.; Wang, X.; and Cao, Y. 2022 · 2022
Cited alongside, same era.
LoRA: Low-Rank Adaptation of Large Language Models
Hu, E. J.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; and Chen, W. 2022 · 2022
Cited alongside, same era.
OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework
Wang, P.; Yang, A.; Men, R.; Lin, J.; Bai, S.; Li, Z.; Ma, J.; Zhou, C.; Zhou, J.; and Yang, H. 2022 · 2022
Cited alongside, same era.
PEVL: Position-enhanced Pre-training and Prompt Tuning for Vision-language Models
Yao, Y.; Chen, Q.; Zhang, A.; Ji, W.; Liu, Z.; Chua, T.; and Sun, M. 2022 · 2022
Cited alongside, same era.
Point-BERT: Pre-training 3D Point Cloud Transformers with Masked Point Modeling
Yu, X.; Tang, L.; Rao, Y.; Huang, T.; Zhou, J.; and Lu, J. 2022 · 2022
Cited alongside, same era.
OPT: Open Pre-trained Transformer Language Models
Zhang, S.; Roller, S.; Goyal, N.; Artetxe, M.; Chen, M.; Chen, S.; Dewan, C.; Diab, M. T.; Li, X.; Lin, X. V.; Mihaylov, T.; Ott, M.; Shleifer, S.; Shuster, K.; Simig, D.; Koura, P. S.; Sridhar, A.; Wang, T.; and Zettlemoyer, L. 2022 · 2022
Cited alongside, same era.
Lyu, C.; Wu, M.; Wang, L.; Huang, X.; Liu, B.; Du, Z.; Shi, S.; and Tu, Z. 2023 · 2023
Closest in time.
Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
Maaz, M.; Rasheed, H.; Khan, S.; and Khan, F. S. 2023 · 2023
Closest in time.
Kosmos-2: Grounding Multimodal Large Language Models to the World
Peng, Z.; Wang, W.; Dong, L.; Hao, Y.; Huang, S.; Ma, S.; and Wei, F. 2023 · 2023
Closest in time.
Caption anything: Interactive image description with diverse multimodal controls
Wang, T.; Zhang, J.; Fei, J.; Ge, Y.; Zheng, H.; Tang, Y.; Li, Z.; Gao, M.; Zhao, S.; Shan, Y.; et al. 2023 · 2023
Closest in time.
mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Ye, Q.; Xu, H.; Xu, G.; Ye, J.; Yan, M.; Zhou, Y.; Wang, J.; Hu, A.; Shi, P.; Shi, Y.; Li, C.; Xu, Y.; Chen, H.; Tian, J.; Qi, Q.; Zhang, J.; and Huang, F. 2023 · 2023
Closest in time.
LAMM: Language-Assisted Multi-Modal Instruction-Tuning Dataset, Framework, and Benchmark
Yin, Z.; Wang, J.; Cao, J.; Shi, Z.; Liu, D.; Li, M.; Sheng, L.; Bai, L.; Huang, X.; Wang, Z.; Shao, J.; and Ouyang, W. 2023 · 2023
Closest in time.
Transfer Visual Prompt Generator across LLMs
Zhang, A.; Fei, H.; Yao, Y.; Ji, W.; Li, L.; Liu, Z.; and Chua, T.-S. 2023 · 2023
Closest in time.
Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
Zhang, H.; Li, X.; and Bing, L. 2023 · 2023
Closest in time.
MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Zhu, D.; Chen, J.; Shen, X.; Li, X.; and Elhoseiny, M. 2023 · 2023
Closest in time.