Fetching the paper…
Reading the bibliography…
Current multimodal language model (MLM) training approaches overlook the influence of instruction templates.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2019
Earlier work this paper cites.
Sentence-bert: Sentence embeddings using siamese bert-networks
N Reimers · 2019
Earlier work this paper cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen · 2021
Earlier work this paper cites.
Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity
Yao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel, and Pontus Stenetorp · 2021
Earlier work this paper cites.
Multitask prompted training enables zero-shot task generalization
Victor Sanh, Albert Webson, Colin Raffel, Stephen H Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Teven Le Scao, Arun Raja, et al · 2021
Earlier work this paper cites.
Demystifying prompts in language models via perplexity estimation
Hila Gonen, Srini Iyer, Terra Blevins, Noah A Smith, and Luke Zettlemoyer · 2022
Earlier work this paper cites.
Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi · 2022
Earlier work this paper cites.
Holistic evaluation of language models
Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, et al · 2022
Earlier work this paper cites.
Openflamingo: An open-source framework for training large autoregressive vision-language models, 2023
Anas Awadalla, Irena Gao, Josh Gardner, Jack Hessel, Yusuf Hanafy, Wanrong Zhu, Kalyani Marathe, Yonatan Bitton, Samir Gadre, Shiori Sagawa, Jenia Jitsev, Simon Kornblith, Pang Wei Koh, Gabriel Ilharco, Mitchell Wortsman, and Ludwig Schmidt · 2023
Earlier work this paper cites.
Qwen-vl: A frontier large vision-language model with versatile abilities
Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou · 2023
Earlier work this paper cites.
Misar: A multimodal instructional system with augmented reality, 2023
Jing Bi, Nguyen Manh Nguyen, Ali Vosoughi, and Chenliang Xu · 2023
Earlier work this paper cites.
Chengguang Gan and Tatsunori Mori · 2023
Earlier work this paper cites.
The language of prompting: What linguistic properties make a prompt successful?
Alina Leidinger, Robert Van Rooij, and Ekaterina Shutova · 2023
Earlier work this paper cites.
Mm-vid: Advancing video understanding with gpt-4v(ision), 2023
Kevin Lin, Faisal Ahmed, Linjie Li, Chung-Ching Lin, Ehsan Azarnasab, Zhengyuan Yang, Jianfeng Wang, Lin Liang, Zicheng Liu, Yumao Lu, Ce Liu, and Lijuan Wang · 2023
Earlier work this paper cites.
Improved baselines with visual instruction tuning, 2023
Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee · 2023
Earlier work this paper cites.
Chameleon: Plug-and-play compositional reasoning with large language models, 2023
Pan Lu, Baolin Peng, Hao Cheng, Michel Galley, Kai-Wei Chang, Ying Nian Wu, Song-Chun Zhu, and Jianfeng Gao · 2023
Cited alongside, same era.
Macaw-llm: Multi-modal language modeling with image, audio, video, and text integration, 2023
Chenyang Lyu, Minghao Wu, Longyue Wang, Xinting Huang, Bingshuai Liu, Zefeng Du, Shuming Shi, and Zhaopeng Tu · 2023
Cited alongside, same era.
What makes chain-of-thought prompting effective? a counterfactual study
Aman Madaan, Katherine Hermann, and Amir Yazdanbakhsh · 2023
Cited alongside, same era.
Artemis Panagopoulou, Le Xue, Ning Yu, Junnan Li, Dongxu Li, Shafiq R. Joty, Ran Xu, Silvio Savarese, Caiming Xiong, and Juan Carlos Niebles · 2023
Cited alongside, same era.
Kosmos-2: Grounding multimodal large language models to the world, 2023
Zhiliang Peng, Wenhui Wang, Li Dong, Yaru Hao, Shaohan Huang, Shuming Ma, and Furu Wei · 2023
Instructblip: Towards general-purpose vision-language models with instruction tuning
Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale N Fung, and Steven Hoi · 2024
Closest in time.
Scaling laws of synthetic images for model training… for now
Lijie Fan, Kaifeng Chen, Dilip Krishnan, Dina Katabi, Phillip Isola, and Yonglong Tian · 2024
Closest in time.
Blink: Multimodal large language models can see but not perceive
Xingyu Fu, Yushi Hu, Bangzheng Li, Yu Feng, Haoyu Wang, Xudong Lin, Dan Roth, Noah A Smith, Wei-Chiu Ma, and Ranjay Krishna · 2024
Closest in time.
What matters when building vision-language models?
Hugo Laurençon, Léo Tronchon, Matthieu Cord, and Victor Sanh · 2024
Closest in time.
Llava-onevision: Easy visual task transfer
Bo Li, Yuanhan Zhang, Dong Guo, Renrui Zhang, Feng Li, Hao Zhang, Kaichen Zhang, Yanwei Li, Ziwei Liu, and Chunyuan Li · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Quantifying language models’ sensitivity to spurious features in prompt design or: How i learned to start worrying about prompt formatting
Melanie Sclar, Yejin Choi, Yulia Tsvetkov, and Alane Suhr · 2023
Cited alongside, same era.
Unival: Unified model for image, video, audio and language tasks, 2023
Mustafa Shukor, Corentin Dancette, Alexandre Rame, and Matthieu Cord · 2023
Cited alongside, same era.
Fine-grained audio-visual joint representations for multimodal large language models, 2023
Guangzhi Sun, Wenyi Yu, Changli Tang, Xianzhao Chen, Tian Tan, Wei Li, Lu Lu, Zejun Ma, and Chao Zhang · 2023
Cited alongside, same era.
Llmva-gebc: Large language model with video adapter for generic event boundary captioning, 2023
Yunlong Tang, Jinrui Zhang, Xiangchen Wang, Teng Wang, and Feng Zheng · 2023
Cited alongside, same era.
Chatvideo: A tracklet-centric multimodal and versatile video understanding system, 2023
Junke Wang, Dongdong Chen, Chong Luo, Xiyang Dai, Lu Yuan, Zuxuan Wu, and Yu-Gang Jiang · 2023
Cited alongside, same era.
Split & merge: Unlocking the potential of visual adapters via sparse training
Qizhe Zhang, Bocheng Zou, Ruichuan An, Jiaming Liu, and Shanghang Zhang · 2023
Cited alongside, same era.
On large language models’ selection bias in multi-choice questions
Chujie Zheng, Hao Zhou, Fandong Meng, Jie Zhou, and Minlie Huang · 2023
Cited alongside, same era.
Closest in time.
Draw-and-understand: Leveraging visual prompts to enable mllms to comprehend what you want
Weifeng Lin, Xinyu Wei, Ruichuan An, Peng Gao, Bocheng Zou, Yulin Luo, Siyuan Huang, Shanghang Zhang, and Hongsheng Li · 2024
Closest in time.
Llava-next: Improved reasoning, ocr, and world knowledge, January 2024b
Haotian Liu, Chunyuan Li, Yuheng Li, Bo Li, Yuanhan Zhang, Sheng Shen, and Yong Jae Lee · 2024
Closest in time.
Llm as dataset analyst: Subpopulation structure discovery with large language model
Yulin Luo, Ruichuan An, Bocheng Zou, Yiming Tang, Jiaming Liu, and Shanghang Zhang · 2024
Closest in time.
m&m’s: A benchmark to evaluate tool-use for multi-step multi-modal tasks
Zixian Ma, Weikai Huang, Jieyu Zhang, Tanmay Gupta, and Ranjay Krishna · 2024
Closest in time.
State of what art? a call for multi-prompt llm evaluation
Moran Mizrahi, Guy Kaplan, Dan Malkin, Rotem Dror, Dafna Shahaf, and Gabriel Stanovsky · 2024
Closest in time.
Mind your format: Towards consistent evaluation of in-context learning improvements
Anton Voronov, Lena Wolf, and Max Ryabinin · 2024
Closest in time.
Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi
Xiang Yue, Yuansheng Ni, Kai Zhang, Tianyu Zheng, Ruoqi Liu, Ge Zhang, Samuel Stevens, Dongfu Jiang, Weiming Ren, Yuxuan Sun, et al · 2024
Closest in time.
Prosa: Assessing and understanding the prompt sensitivity of llms
Jingming Zhuo, Songyang Zhang, Xinyu Fang, Haodong Duan, Dahua Lin, and Kai Chen · 2024
Closest in time.
Mmbench: Is your multi-modal model an all-around player?
Yuan Liu, Haodong Duan, Yuanhan Zhang, Bo Li, Songyang Zhang, Wangbo Zhao, Yike Yuan, Jiaqi Wang, Conghui He, Ziwei Liu, et al · 2025
Closest in time.
Benchmarking prompt sensitivity in large language models
Amirhossein Razavi, Mina Soltangheis, Negar Arabzadeh, Sara Salamat, Morteza Zihayat, and Ebrahim Bagheri · 2025
Closest in time.