Fetching the paper…
Reading the bibliography…
Despite the exceptional performance of multi-modal large language models (MLLMs), their deployment requires substantial computational resources.
C. E. Shannon, “A mathematical theory of communication,” The Bell system technical journal , vol. 27, no. 3, pp. 379–423, 1948
1948
Earlier work this paper cites.
S. Kullback and R. A. Leibler, “On information and sufficiency,” The annals of mathematical statistics , vol. 22, no. 1, pp. 79–86, 1951
1951
Earlier work this paper cites.
Y. Bengio, R. Ducharme, and P. Vincent, “A neural probabilistic language model,” in NeurIPS , 2000
2000
Earlier work this paper cites.
M. Fazel, “Matrix rank minimization with applications,” Ph.D. dissertation, Stanford University, 2002
2002
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in CVPR , 2009
2009
Earlier work this paper cites.
K. Pelechrinis, M. Iliofotou, and S. V. Krishnamurthy, “Denial of service attacks in wireless networks: The case of jammers,” IEEE Communications surveys & tutorials , vol. 13, no. 2, pp. 245–257, 2010
2010
Earlier work this paper cites.
D. Chen and W. B. Dolan, “Collecting highly parallel data for paraphrase evaluation,” in ACL , 2011
2011
Earlier work this paper cites.
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick, “Microsoft coco: Common objects in context,” in ECCV , 2014
2014
Earlier work this paper cites.
I. J. Goodfellow, J. Shlens, and C. Szegedy, “Explaining and harnessing adversarial examples,” in ICLR , 2015
2015
Earlier work this paper cites.
C. Silberer, V. Ferrari, and M. Lapata, “Visually grounded meaning representations,” IEEE transactions on pattern analysis and machine intelligence , vol. 39, no. 11, pp. 2284–2297, 2016
2016
Earlier work this paper cites.
Y. Jang, Y. Song, Y. Yu, Y. Kim, and G. Kim, “Tgif-qa: Toward spatio-temporal reasoning in visual question answering,” in CVPR , 2017
2017
Earlier work this paper cites.
R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh, and D. Batra, “Grad-cam: Visual explanations from deep networks via gradient-based localization,” in ICCV , 2017
2017
Earlier work this paper cites.
Y. Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh, “Making the v in vqa matter: Elevating the role of image understanding in visual question answering,” in CVPR , 2017
2017
Earlier work this paper cites.
P. Anderson, X. He, C. Buehler, D. Teney, M. Johnson, S. Gould, and L. Zhang, “Bottom-up and top-down attention for image captioning and visual question answering,” in CVPR , 2018
2018
Earlier work this paper cites.
A. Madry, A. Makelov, L. Schmidt, D. Tsipras, and A. Vladu, “Towards deep learning models resistant to adversarial attacks,” in ICLR , 2018
2018
Earlier work this paper cites.
A. Rohrbach, L. A. Hendricks, K. Burns, T. Darrell, and K. Saenko, “Object hallucination in image captioning,” in EMNLP , 2018
2018
Earlier work this paper cites.
R. Zhang, P. Isola, A. A. Efros, E. Shechtman, and O. Wang, “The unreasonable effectiveness of deep features as a perceptual metric,” in CVPR , 2018
2018
Earlier work this paper cites.
A. Ilyas, L. Engstrom, A. Athalye, and J. Lin, “Black-box adversarial attacks with limited queries and information,” in ICML , 2018
2018
Earlier work this paper cites.
Y. Dong, F. Liao, T. Pang, H. Su, J. Zhu, X. Hu, and J. Li, “Boosting adversarial attacks with momentum,” in CVPR , 2018
2018
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
L. Peng, Y. Yang, Z. Wang, Z. Huang, and H. T. Shen, “Mra-net: Improving vqa via multi-modal relation attention network,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 44, no. 1, pp. 318–329, 2020
2020
Earlier work this paper cites.
A. Holtzman, J. Buys, L. Du, M. Forbes, and Y. Choi, “The curious case of neural text degeneration,” in ICLR , 2020
2020
Cited alongside, same era.
Y. Bai, Y. Zeng, Y. Jiang, Y. Wang, S.-T. Xia, and W. Guo, “Improving query efficiency of black-box adversarial attack,” in ECCV , 2020
2020
Cited alongside, same era.
J. Li, R. Selvaraju, A. Gotmare, S. Joty, C. Xiong, and S. C. H. Hoi, “Align before fuse: Vision and language representation learning with momentum distillation,” in NeurIPS , 2021
2021
Cited alongside, same era.
D. Patterson, J. Gonzalez, Q. Le, C. Liang, L.-M. Munguia, D. Rothchild, D. So, M. Texier, and J. Dean, “Carbon emissions and large neural network training,” 2021
2021
Cited alongside, same era.
I. Shumailov, Y. Zhao, D. Bates, N. Papernot, R. Mullins, and R. Anderson, “Sponge examples: Energy-latency attacks on neural networks,” in IEEE EuroS&P , 2021
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2021
Cited alongside, same era.
S. Hong, Y. Kaya, I.-V. Modoranu, and T. Dumitraş, “A panda? no, it’s a sloth: Slowdown attacks on adaptive multi-exit neural network inference,” in ICLR , 2021
2021
Cited alongside, same era.
J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds et al. , “Flamingo: a visual language model for few-shot learning,” in NeurIPS , 2022
2022
Cited alongside, same era.
J. Chen, H. Guo, K. Yi, B. Li, and M. Elhoseiny, “Visualgpt: Data-efficient adaptation of pretrained language models for image captioning,” in CVPR , 2022
2022
Cited alongside, same era.
J. Li, D. Li, C. Xiong, and S. Hoi, “Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,” in ICML , 2022
2022
Cited alongside, same era.
S. Chen, Z. Song, M. Haque, C. Liu, and W. Yang, “Nicgslowdown: Evaluating the efficiency robustness of neural image caption generation models,” in CVPR , 2022
2022
Cited alongside, same era.
2022
Cited alongside, same era.
S. Chen, C. Liu, M. Haque, Z. Song, and W. Yang, “Nmtsloth: understanding and testing efficiency degradation of neural machine translation systems,” in Proceedings of the 30th ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering , 2022, pp. 1148–1160
2022
Cited alongside, same era.
Z. Zhang, K. Chen, R. Wang, M. Utiyama, E. Sumita, Z. Li, and H. Zhao, “Universal multimodal representation for language understanding,” IEEE Transactions on Pattern Analysis and Machine Intelligence , 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
Y. Dong, H. Chen, J. Chen, Z. Fang, X. Yang, Y. Zhang, Y. Tian, H. Su, and J. Zhu, “How robust is google’s bard to adversarial image attacks?” in NeurIPS Workshop , 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
S. Chen, H. Chen, M. Haque, C. Liu, and W. Yang, “The dark side of dynamic routing neural networks: Towards efficiency backdoor injection,” in CVPR , 2023
2023
Later among the works it cites.
H. Liu, Y. Wu, Z. Yu, Y. Vorobeychik, and N. Zhang, “Slowlidar: Increasing the latency of lidar-based detection using adversarial examples,” in CVPR , 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
K. Li, Y. Wang, Y. He, Y. Li, Y. Wang, Y. Liu, Z. Wang, J. Xu, G. Chen, P. Luo et al. , “Mvbench: A comprehensive multi-modal video understanding benchmark,” in CVPR , 2024
2024
Closest in time.
2024
Closest in time.
K. Gao, Y. Bai, J. Gu, S.-T. Xia, P. Torr, Z. Li, and W. Liu, “Inducing high energy-latency of large vision-language models with verbose images,” in ICLR , 2024
2024
Closest in time.
K. Gao, Y. Bai, J. Bai, Y. Yang, and S.-T. Xia, “Adversarial robustness for visual grounding of multimodal large language models,” in ICLR Workshop , 2024
2024
Closest in time.
J. Bai, K. Gao, S. Min, S.-T. Xia, Z. Li, and W. Liu, “Badclip: Trigger-aware prompt learning for backdoor attacks on clip,” in CVPR , 2024
2024
Closest in time.
H. Luo, J. Gu, F. Liu, and P. Torr, “An image is worth 1000 lies: Adversarial transferability across prompts on vision-language models,” in ICLR , 2024
2024
Closest in time.
H. Wang, K. Dong, Z. Zhu, H. Qin, A. Liu, X. Fang, J. Wang, and X. Liu, “Transferable multimodal attack on vision-language pre-training models,” in 2024 IEEE Symposium on Security and Privacy (SP) . IEEE Computer Society, 2024, pp. 102–102
2024
Closest in time.
X. Qi, K. Huang, A. Panda, M. Wang, and P. Mittal, “Visual adversarial examples jailbreak large language models,” in AAAI , 2024
2024
Closest in time.