Fetching the paper…
Reading the bibliography…
Large Multimodal Models (LMMs) have achieved significant progress by extending large language models.
Detect what you can: Detecting and representing objects using holistic models and body parts
Chen, X.; Mottaghi, R.; Liu, X.; Fidler, S.; Urtasun, R.; and Yuille, A. 2014 · 1978
Earlier work this paper cites.
Transfer of Rule-Based Expertise through a Tutorial Dialogue
Clancey, W. J. 1979 · 1979
Earlier work this paper cites.
Communication, Simulation, and Intelligent Agents: Implications of Personal Intelligent Machines for Medical Education
Clancey, W. J. 1983 · 1983
Earlier work this paper cites.
Strategic Explanations in Consultation—Duplicate
Hasling, D. W.; Clancey, W. J.; Rennels, G. R.; and Test, T. 1983 · 1983
Earlier work this paper cites.
Classification Problem Solving
Clancey, W. J. 1984 · 1984
Earlier work this paper cites.
Strategic explanations for a diagnostic consultation system
Hasling, D. W.; Clancey, W. J.; and Rennels, G. 1984 · 1984
Earlier work this paper cites.
Blackboard Systems
Engelmore, R.; and Morgan, A., eds. 1986 · 1986
Earlier work this paper cites.
Poligon: A System for Parallel Problem Solving
Rice, J. 1986 · 1986
Earlier work this paper cites.
VQA: Visual Question Answering
Antol, S.; Agrawal, A.; Lu, J.; Mitchell, M.; Batra, D.; Zitnick, C. L.; and Parikh, D. 2015 · 2015
Earlier work this paper cites.
Microsoft coco captions: Data collection and evaluation server
Chen, X.; Fang, H.; Lin, T.-Y.; Vedantam, R.; Gupta, S.; Dollár, P.; and Zitnick, C. L. 2015 · 2015
Earlier work this paper cites.
Deep Visual-Semantic Alignments for Generating Image Descriptions
Karpathy, A.; and Fei-Fei, L. 2015 · 2015
Earlier work this paper cites.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
Plummer, B. A.; Wang, L.; Cervantes, C. M.; Caicedo, J. C.; Hockenmaier, J.; and Lazebnik, S. 2015 · 2015
Earlier work this paper cites.
Modeling context in referring expressions
Yu, L.; Poirson, P.; Yang, S.; Berg, A. C.; and Berg, T. L. 2016 · 2016
Earlier work this paper cites.
The mapillary vistas dataset for semantic understanding of street scenes
Neuhold, G.; Ollmann, T.; Rota Bulo, S.; and Kontschieder, P. 2017 · 2017
Earlier work this paper cites.
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, L.; and Polosukhin, I. 2017 · 2017
Earlier work this paper cites.
Coco-stuff: Thing and stuff classes in context
Caesar, H.; Uijlings, J.; and Ferrari, V. 2018 · 2018
Earlier work this paper cites.
Pluto: The ’Other’ Red Planet
NASA. 2015 · 2018
Earlier work this paper cites.
Mattnet: Modular attention network for referring expression comprehension
Yu, L.; Lin, Z.; Shen, X.; Yang, J.; Lu, X.; Bansal, M.; and Berg, T. L. 2018 · 2018
Earlier work this paper cites.
Nocaps: Novel object captioning at scale
Agrawal, H.; Desai, K.; Wang, Y.; Chen, X.; Jain, R.; Johnson, M.; Batra, D.; Parikh, D.; Lee, S.; and Anderson, P. 2019 · 2019
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2019 · 2019
Cited alongside, same era.
Semantic understanding of scenes through the ade20k dataset
Zhou, B.; Zhao, H.; Puig, X.; Xiao, T.; Fidler, S.; Barriuso, A.; and Torralba, A. 2019 · 2019
Cited alongside, same era.
Multi-task collaborative network for joint referring expression comprehension and segmentation
Luo, G.; Zhou, Y.; Sun, X.; Cao, L.; Wu, C.; Deng, C.; and Ji, R. 2020 · 2020
Cited alongside, same era.
The Engineering of Qualitative Models
Clancey, W. J. 2021 · 2021
Cited alongside, same era.
Vision-language transformer and query generation for referring segmentation
Ding, H.; Liu, C.; Wang, S.; and Jiang, X. 2021 · 2021
Cited alongside, same era.
Shikra: Unleashing multimodal llm’s referential dialogue magic
Chen, K.; Zhang, Z.; Zeng, W.; Zhang, R.; Zhu, F.; and Zhao, R. 2023 · 2023
Later among the works it cites.
Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90%* ChatGPT Quality
Chiang, W.-L.; Li, Z.; Lin, Z.; Sheng, Y.; Wu, Z.; Zhang, H.; Zheng, L.; Zhuang, S.; Zhuang, Y.; Gonzalez, J. E.; Stoica, I.; and Xing, E. P. 2023 · 2023
Later among the works it cites.
InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
Dai, W.; Li, J.; Li, D.; Tiong, A. M. H.; Zhao, J.; Wang, W.; Li, B.; Fung, P.; and Hoi, S. 2023 · 2023
Later among the works it cites.
Segment anything
Kirillov, A.; Mintun, E.; Ravi, N.; Mao, H.; Rolland, C.; Gustafson, L.; Xiao, T.; Whitehead, S.; Berg, A. C.; Lo, W.-Y.; et al. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hu, E. J.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; and Chen, W. 2021 · 2021
Cited alongside, same era.
Locate then segment: A strong pipeline for referring image segmentation
Jing, Y.; Kong, T.; Wang, W.; Wang, L.; Li, L.; and Tan, T. 2021 · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021 · 2021
Cited alongside, same era.
Vennerød, C. B.; Kjærran, A.; and Bugge, E. S. 2021 · 2021
Cited alongside, same era.
Simvlm: Simple visual language model pretraining with weak supervision
Wang, Z.; Yu, J.; Yu, A. W.; Dai, Z.; Tsvetkov, Y.; and Cao, Y. 2021 · 2021
Cited alongside, same era.
Vinvl: Revisiting visual representations in vision-language models
Zhang, P.; Li, X.; Hu, X.; Yang, J.; Zhang, L.; Wang, L.; Choi, Y.; and Gao, J. 2021 · 2021
Cited alongside, same era.
Flamingo: a visual language model for few-shot learning
Alayrac, J.-B.; Donahue, J.; Luc, P.; Miech, A.; Barr, I.; Hasson, Y.; Lenc, K.; Mensch, A.; Millican, K.; Reynolds, M.; et al. 2022 · 2022
Cited alongside, same era.
Li, J.; Li, D.; Savarese, S.; and Hoi, S. 2023 · 2023
Later among the works it cites.
Kosmos-2: Grounding multimodal large language models to the world
Peng, Z.; Wang, W.; Dong, L.; Hao, Y.; Huang, S.; Ma, S.; and Wei, F. 2023 · 2023
Later among the works it cites.
Paco: Parts and attributes of common objects
Ramanathan, V.; Kalia, A.; Petrovic, V.; Wen, Y.; Zheng, B.; Guo, B.; Wang, R.; Marquez, A.; Kovvuri, R.; Kadian, A.; et al. 2023 · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Touvron, H.; Lavril, T.; Izacard, G.; Martinet, X.; Lachaux, M.-A.; Lacroix, T.; Rozière, B.; Goyal, N.; Hambro, E.; Azhar, F.; et al. 2023 · 2023
Later among the works it cites.
MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Zhu, D.; Chen, J.; Shen, X.; Li, X.; and Elhoseiny, M. 2023 · 2023
Later among the works it cites.
Language is not all you need: Aligning perception with language models
Huang, S.; Dong, L.; Wang, W.; Hao, Y.; Singhal, S.; Ma, S.; Lv, T.; Cui, L.; Mohammed, O. K.; Patra, B.; et al. 2024 · 2024
Closest in time.
Lisa: Reasoning segmentation via large language model
Lai, X.; Tian, Z.; Chen, Y.; Li, Y.; Yuan, Y.; Liu, S.; and Jia, J. 2024 · 2024
Closest in time.
EVCap: Retrieval-Augmented Image Captioning with External Visual-Name Memory for Open-World Comprehension
Li, J.; Vo, D. M.; Sugimoto, A.; and Nakayama, H. 2024 · 2024
Closest in time.
Glamm: Pixel grounding large multimodal model
Rasheed, H.; Maaz, M.; Shaji, S.; Shaker, A.; Khan, S.; Cholakkal, H.; Anwer, R. M.; Xing, E.; Yang, M.-H.; and Khan, F. S. 2024 · 2024
Closest in time.
Pixellm: Pixel reasoning with large multimodal model
Ren, Z.; Huang, Z.; Wei, Y.; Zhao, Y.; Fu, D.; Feng, J.; and Jin, X. 2024 · 2024
Closest in time.
LaSagnA: Language-based Segmentation Assistant for Complex Queries
Wei, C.; Tan, H.; Zhong, Y.; Yang, Y.; and Ma, L. 2024 · 2024
Closest in time.
F-LMM: Grounding Frozen Large Multimodal Models
Wu, S.; Jin, S.; Zhang, W.; Xu, L.; Liu, W.; Li, W.; and Loy, C. C. 2024 · 2024
Closest in time.
Gsva: Generalized segmentation via multimodal large language models
Xia, Z.; Han, D.; Han, Y.; Pan, X.; Song, S.; and Huang, G. 2024 · 2024
Closest in time.
Osprey: Pixel understanding with visual instruction tuning
Yuan, Y.; Li, W.; Liu, J.; Tang, D.; Luo, X.; Qin, C.; Zhang, L.; and Zhu, J. 2024 · 2024
Closest in time.