Fetching the paper…
Reading the bibliography…
Vision-language models (VLMs) show remarkable performance in multimodal tasks.
“Unified matrix treatment of the fast walsh-hadamard transform,”
Fino and Algazi, · 1976
Earlier work this paper cites.
“Microsoft coco captions: Data collection and evaluation server,”
Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Dollár, and C Lawrence Zitnick, · 2015
Earlier work this paper cites.
“Pointer sentinel mixture models,”
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher, · 2016
Earlier work this paper cites.
“Attention is all you need,”
A Vaswani, · 2017
Earlier work this paper cites.
“Multimodal intelligence: Representation learning, information fusion, and applications,”
Chao Zhang, Zichao Yang, Xiaodong He, and Li Deng, · 2020
Earlier work this paper cites.
“An image is worth 16x16 words: Transformers for image recognition at scale,”
Dosovitskiy Alexey, · 2020
Earlier work this paper cites.
“A survey on deep multimodal learning for computer vision: advances, trends, applications, and datasets,”
Khaled Bayoudh, Raja Knani, Fayçal Hamdaoui, and Abdellatif Mtibaa, · 2022
Earlier work this paper cites.
“Flashattention: Fast and memory-efficient exact attention with io-awareness,”
Tri Dao, Dan Fu, Stefano Ermon, Atri Rudra, and Christopher Ré, · 2022
Earlier work this paper cites.
“A survey on multimodal large language models,”
Shukang Yin, Chaoyou Fu, Sirui Zhao, Ke Li, Xing Sun, Tong Xu, and Enhong Chen, · 2023
Earlier work this paper cites.
“Rt-2: Vision-language-action models transfer web knowledge to robotic control,”
Anthony Brohan, Noah Brown, Justice Carbajal, Yevgen Chebotar, Xi Chen, Krzysztof Choromanski, Tianli Ding, Danny Driess, Avinava Dubey, Chelsea Finn, et al., · 2023
Earlier work this paper cites.
“Llama 2: Open foundation and fine-tuned chat models,”
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al., · 2023
Earlier work this paper cites.
“Efficient streaming language models with attention sinks,”
Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, and Mike Lewis, · 2023
Cited alongside, same era.
“Improved baselines with visual instruction tuning,” 2023
Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee, · 2023
Cited alongside, same era.
“SmoothQuant: Accurate and efficient post-training quantization for large language models,”
Guangxuan Xiao, Ji Lin, Mickael Seznec, Hao Wu, Julien Demouth, and Song Han, · 2023
Cited alongside, same era.
“Vision-language models for vision tasks: A survey,”
Jingyi Zhang, Jiaxing Huang, Sheng Jin, and Shijian Lu, · 2024
Cited alongside, same era.
“Kivi: A tuning-free asymmetric 2bit quantization for kv cache,”
Zirui Liu, Jiayi Yuan, Hongye Jin, Shaochen Zhong, Zhaozhuo Xu, Vladimir Braverman, Beidi Chen, and Xia Hu, · 2024
Cited alongside, same era.
“Quip#: Even better llm quantization with hadamard incoherence and lattice codebooks,”
Albert Tseng, Jerry Chee, Qingyao Sun, Volodymyr Kuleshov, and Christopher De Sa, · 2024
Later among the works it cites.
“Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution,”
Peng et al. Wang, · 2024
Later among the works it cites.
“Look-m: Look-once optimization in kv cache for efficient multimodal long-context inference,”
Zhongwei Wan, Ziang Wu, Che Liu, Jinfa Huang, Zhihong Zhu, Peng Jin, Longyue Wang, and Li Yuan, · 2024
Later among the works it cites.
June Yong et al. Yang, · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Zipcache: Accurate and efficient kv cache quantization with salient token identification,”
Yefei He, Luoming Zhang, Weijia Wu, Jing Liu, Hong Zhou, and Bohan Zhuang, · 2024
Cited alongside, same era.
“Skvq: Sliding-window key and value cache quantization for large language models,”
Haojie Duanmu, Zhihang Yuan, Xiuhong Li, Jiangfei Duan, Xingcheng Zhang, and Dahua Lin, · 2024
Cited alongside, same era.
“Visual instruction tuning,”
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee, · 2024
Cited alongside, same era.
“Intactkv: Improving large language model quantization by keeping pivot tokens intact,”
Ruikang Liu, Haoli Bai, Haokun Lin, Yuening Li, Han Gao, Zhengzhuo Xu, Lu Hou, Jun Yao, and Chun Yuan, · 2024
Cited alongside, same era.
“Kvquant: Towards 10 million context length llm inference with kv cache quantization,”
Coleman Hooper, Sehoon Kim, Hiva Mohammadzadeh, Michael W Mahoney, Yakun Sophia Shao, Kurt Keutzer, and Amir Gholami, · 2024
Cited alongside, same era.
“Quarot: Outlier-free 4-bit inference in rotated llms,”
Saleh Ashkboos, Amirkeivan Mohtashami, Maximilian L Croci, Bo Li, Pashmina Cameron, Martin Jaggi, Dan Alistarh, Torsten Hoefler, and James Hensman, · 2024
Cited alongside, same era.
Mingjie Sun, Xinlei Chen, J Zico Kolter, and Zhuang Liu, · 2024
Later among the works it cites.
Yefei He, Feng Chen, Jing Liu, Wenqi Shao, Hong Zhou, Kaipeng Zhang, and Bohan Zhuang, · 2024
Later among the works it cites.
“Spinquant–llm quantization with learned rotations,”
Zechun et al. Liu, · 2024
Later among the works it cites.
“Roformer: Enhanced transformer with rotary position embedding,”
Jianlin Su, Murtadha Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu, · 2024
Later among the works it cites.
“Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks,”
Zhe Chen, Jiannan Wu, Wenhai Wang, Weijie Su, Guo Chen, Sen Xing, Muyan Zhong, Qinglong Zhang, Xizhou Zhu, Lewei Lu, et al., · 2024
Later among the works it cites.
“Milebench: Benchmarking mllms in long context,”
Dingjie Song, Shunian Chen, Guiming Hardy Chen, Fei Yu, Xiang Wan, and Benyou Wang, · 2024
Later among the works it cites.