Fetching the paper…
Reading the bibliography…
Multimodal language models (MLMs) still face challenges in fundamental visual perception tasks where specialized models excel.
MACHINE PERCEPTION OF THREE-DIMENSIONAL, SO LIDS
PHILOSO EPHY DO CT OR OF · 1961
Earlier work this paper cites.
An introduction to computational geometry
Marvin Minsky and Seymour Papert · 1969
Earlier work this paper cites.
A combined corner and edge detector
Chris Harris, Mike Stephens, et al · 1988
Earlier work this paper cites.
Depth estimation from image structure
Antonio Torralba and Aude Oliva · 2002
Earlier work this paper cites.
Vision: A computational investigation into the human representation and processing of visual information
David Marr · 2010
Earlier work this paper cites.
Single-image depth perception in the wild
Weifeng Chen, Zhao Fu, Dawei Yang, and Jia Deng · 2016
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, et al · 2017
Earlier work this paper cites.
Neural discrete representation learning
Aaron Van Den Oord, Oriol Vinyals, et al · 2017
Earlier work this paper cites.
Scene parsing through ade20k dataset
Bolei Zhou, Hang Zhao, Xavier Puig, Sanja Fidler, Adela Barriuso, and Antonio Torralba · 2017
Earlier work this paper cites.
LVIS: A dataset for large vocabulary instance segmentation
Agrim Gupta, Piotr Dollar, and Ross Girshick · 2019
Earlier work this paper cites.
Semantic understanding of scenes through the ade20k dataset
Bolei Zhou, Hang Zhao, Xavier Puig, Tete Xiao, Sanja Fidler, Adela Barriuso, and Antonio Torralba · 2019
Earlier work this paper cites.
Pix2seq: A language modeling framework for object detection
Ting Chen, Saurabh Saxena, Lala Li, David J Fleet, and Geoffrey Hinton · 2021
Earlier work this paper cites.
Taming transformers for high-resolution image synthesis
Patrick Esser, Robin Rombach, and Bjorn Ommer · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen · 2021
Earlier work this paper cites.
Flamingo: a visual language model for few-shot learning, 2022
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katie Millican, Malcolm Reynolds, Roman Ring, Eliza Rutherford, Serkan Cabi, Tengda Han, Zhitao Gong, Sina Samangooei, Marianne Monteiro, Jacob Menick, Sebastian Borgeaud, Andrew Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikolaj Binkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, and Karen Simonyan · 2022
Cited alongside, same era.
Unified-io: A unified model for vision, language, and multi-modal tasks, 2022
Jiasen Lu, Christopher Clark, Rowan Zellers, Roozbeh Mottaghi, and Aniruddha Kembhavi · 2022
Cited alongside, same era.
Ofa: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework, 2022
Peng Wang, An Yang, Rui Men, Junyang Lin, Shuai Bai, Zhikang Li, Jianxin Ma, Chang Zhou, Jingren Zhou, and Hongxia Yang · 2022
Cited alongside, same era.
Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond, 2023
Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou · 2023
Cited alongside, same era.
Large language models are not robust multiple choice selectors
Chujie Zheng, Hao Zhou, Fandong Meng, Jie Zhou, and Minlie Huang · 2023
Later among the works it cites.
Guiding llms the right way: Fast, non-invasive constrained generation
Luca Beurer-Kellner, Marc Fischer, and Martin Vechev · 2024
Closest in time.
Spatialvlm: Endowing vision-language models with spatial reasoning capabilities
Boyuan Chen, Zhuo Xu, Sean Kirmani, Brain Ichter, Dorsa Sadigh, Leonidas Guibas, and Fei Xia · 2024
Closest in time.
A guide to structured outputs using constrained decoding, 2024
Aidan Cooper · 2024
Closest in time.
Blink: Multimodal large language models can see but not perceive
Xingyu Fu, Yushi Hu, Bangzheng Li, Yu Feng, Haoyu Wang, Xudong Lin, Dan Roth, Noah A Smith, Wei-Chiu Ma, and Ranjay Krishna · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale Fung, and Steven Hoi · 2023
Cited alongside, same era.
Grammar-constrained decoding for structured nlp tasks without finetuning
Saibo Geng, Martin Josifoski, Maxime Peyrard, and Robert West · 2023
Cited alongside, same era.
Cheng-Yu Hsieh, Chun-Liang Li, Chih-Kuan Yeh, Hootan Nakhost, Yasuhisa Fujii, Alexander Ratner, Ranjay Krishna, Chen-Yu Lee, and Tomas Pfister · 2023
Cited alongside, same era.
Generating images with multimodal language models, 2023
Jing Yu Koh, Daniel Fried, and Ruslan Salakhutdinov · 2023
Cited alongside, same era.
All in tokens: Unifying output space of visual tasks via soft token
Jia Ning, Chen Li, Zheng Zhang, Chunyu Wang, Zigang Geng, Qi Dai, Kun He, and Han Hu · 2023
Cited alongside, same era.
OpenAI · 2023
Cited alongside, same era.
Large language models sensitivity to the order of options in multiple-choice questions
Pouya Pezeshkpour and Estevam Hruschka · 2023
Cited alongside, same era.
Visual chatgpt: Talking, drawing and editing with visual foundation models, 2023
Chenfei Wu, Shengming Yin, Weizhen Qi, Xiaodong Wang, Zecheng Tang, and Nan Duan · 2023
Cited alongside, same era.
Unified language-vision pretraining in llm with dynamic discrete visual tokenization, 2024
Yang Jin, Kun Xu, Kun Xu, Liwei Chen, Chao Liao, Jianchao Tan, Quzhe Huang, Bin Chen, Chenyi Lei, An Liu, Chengru Song, Xiaoqiang Lei, Di Zhang, Wenwu Ou, Kun Gai, and Yadong Mu · 2024
Closest in time.
Lisa: Reasoning segmentation via large language model, 2024
Xin Lai, Zhuotao Tian, Yukang Chen, Yanwei Li, Yuhui Yuan, Shu Liu, and Jiaya Jia · 2024
Closest in time.
Llava-onevision: Easy visual task transfer
Bo Li, Yuanhan Zhang, Dong Guo, Renrui Zhang, Feng Li, Hao Zhang, Kaichen Zhang, Yanwei Li, Ziwei Liu, and Chunyuan Li · 2024
Closest in time.
Kosmos-g: Generating images in context with multimodal large language models, 2024
Xichen Pan, Li Dong, Shaohan Huang, Zhiliang Peng, Wenhu Chen, and Furu Wei · 2024
Closest in time.
Fast, high-fidelity llm decoding with regex constraints, 2024
Vivien Tran-Thien · 2024
Closest in time.
Next-gpt: Any-to-any multimodal llm, 2024
Shengqiong Wu, Hao Fei, Leigang Qu, Wei Ji, and Tat-Seng Chua · 2024
Closest in time.
Gsva: Generalized segmentation via multimodal large language models, 2024
Zhuofan Xia, Dongchen Han, Yizeng Han, Xuran Pan, Shiji Song, and Gao Huang · 2024
Closest in time.
mplug-owl: Modularization empowers large language models with multimodality, 2024
Qinghao Ye, Haiyang Xu, Guohai Xu, Jiabo Ye, Ming Yan, Yiyang Zhou, Junyang Wang, Anwen Hu, Pengcheng Shi, Yaya Shi, Chenliang Li, Yuanhong Xu, Hehong Chen, Junfeng Tian, Qi Qian, Ji Zhang, Fei Huang, and Jingren Zhou · 2024
Closest in time.