Fetching the paper…
Reading the bibliography…
Vision-Language Models (VLMs) play a crucial role in the advancement of Artificial General Intelligence (AGI).
Intriguing properties of neural networks
Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus. 2013 · 2013
Earlier work this paper cites.
Artificial general intelligence: concept, state of the art, and future prospects
Ben Goertzel. 2014 · 2014
Earlier work this paper cites.
Explaining and harnessing adversarial examples
Ian J Goodfellow, Jonathon Shlens, and Christian Szegedy. 2014 · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context. In ECCV . 740–755
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. 2014 · 2014
Earlier work this paper cites.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models. In ICCV . 2641–2649
Bryan A Plummer, Liwei Wang, Chris M Cervantes, Juan C Caicedo, Julia Hockenmaier, and Svetlana Lazebnik. 2015 · 2015
Earlier work this paper cites.
Adversarial machine learning at scale
Alexey Kurakin, Ian Goodfellow, and Samy Bengio. 2016 · 2016
Earlier work this paper cites.
Accessorize to a crime: Real and stealthy attacks on state-of-the-art face recognition. In Proceedings of the 2016 acm sigsac conference on computer and communications security . 1528–1540
Mahmood Sharif, Sruti Bhagavatula, Lujo Bauer, and Michael K Reiter. 2016 · 2016
Earlier work this paper cites.
Msr-vtt: A large video description dataset for bridging video and language. In CVPR . 5288–5296
Jun Xu, Tao Mei, Ting Yao, and Yong Rui. 2016 · 2016
Earlier work this paper cites.
Localizing moments in video with natural language. In ICCV . 5803–5812
Lisa Anne Hendricks, Oliver Wang, Eli Shechtman, Josef Sivic, Trevor Darrell, and Bryan Russell. 2017 · 2017
Earlier work this paper cites.
Zoo: Zeroth order optimization based black-box attacks to deep neural networks without training substitute models. In Proceedings of the 10th ACM workshop on artificial intelligence and security . 15–26
Pin-Yu Chen, Huan Zhang, Yash Sharma, Jinfeng Yi, and Cho-Jui Hsieh. 2017 · 2017
Earlier work this paper cites.
Towards deep learning models resistant to adversarial attacks
Aleksander Madry. 2017 · 2017
Earlier work this paper cites.
On detecting adversarial perturbations
Jan Hendrik Metzen, Tim Genewein, Volker Fischer, and Bastian Bischoff. 2017 · 2017
Earlier work this paper cites.
Growing a brain: Fine-tuning by increasing model capacity. In CVPR . 2471–2480
Yu-Xiong Wang, Deva Ramanan, and Martial Hebert. 2017 · 2017
Earlier work this paper cites.
Adversarial detection with model interpretation. In KDD . 1803–1811
Ninghao Liu, Hongxia Yang, and Xia Hu. 2018 · 2018
Earlier work this paper cites.
Generalizing to unseen domains via adversarial data augmentation
Riccardo Volpi, Hongseok Namkoong, Ozan Sener, John C Duchi, Vittorio Murino, and Silvio Savarese. 2018 · 2018
Earlier work this paper cites.
Adversarial attacks on medical machine learning
Samuel G Finlayson, John D Bowers, Joichi Ito, Jonathan L Zittrain, Andrew L Beam, and Isaac S Kohane. 2019 · 2019
Earlier work this paper cites.
Towards artificial general intelligence with hybrid Tianjic chip architecture
Jing Pei, Lei Deng, Sen Song, Mingguo Zhao, Youhui Zhang, Shuang Wu, Guanrui Wang, Zhe Zou, Zhenzhi Wu, Wei He, et al · 2019
Earlier work this paper cites.
Sparse adversarial perturbations for videos. In AAAI . 8973–8980
Xingxing Wei, Jun Zhu, Sha Yuan, and Hang Su. 2019 · 2019
Earlier work this paper cites.
Square attack: a query-efficient black-box adversarial attack via random search. In ECCV . Springer, 484–501
Maksym Andriushchenko, Francesco Croce, Nicolas Flammarion, and Matthias Hein. 2020 · 2020
Earlier work this paper cites.
Large-scale adversarial training for vision-and-language representation learning
Zhe Gan, Yen-Chun Chen, Linjie Li, Chen Zhu, Yu Cheng, and Jingjing Liu. 2020 · 2020
Earlier work this paper cites.
Adversarial training for large neural language models
Xiaodong Liu, Hao Cheng, Pengcheng He, Weizhu Chen, Yu Wang, Hoifung Poon, and Jianfeng Gao. 2020 · 2020
Earlier work this paper cites.
TextAttack: A Framework for Adversarial Attacks, Data Augmentation, and Adversarial Training in NLP. In EMNLP . 119–126
John Morris, Eli Lifland, Jin Yong Yoo, Jake Grigsby, Di Jin, and Yanjun Qi. 2020 · 2020
Earlier work this paper cites.
Bag of tricks for adversarial training
Tianyu Pang, Xiao Yang, Yinpeng Dong, Hang Su, and Jun Zhu. 2020 · 2020
Cited alongside, same era.
Frozen in time: A joint video and image encoder for end-to-end retrieval. In ICCV . 1728–1738
Max Bain, Arsha Nagrani, Gül Varol, and Andrew Zisserman. 2021 · 2021
Cited alongside, same era.
Intelligent driving intelligence test for autonomous vehicles with naturalistic and adversarial environment
Shuo Feng, Xintao Yan, Haowei Sun, Yiheng Feng, and Henry X Liu. 2021 · 2021
Cited alongside, same era.
Lora: Low-rank adaptation of large language models
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2021 · 2021
Cited alongside, same era.
Understanding adversarial attacks on deep learning based medical image analysis systems
Xingjun Ma, Yuhao Niu, Lin Gu, Yisen Wang, Yitian Zhao, James Bailey, and Feng Lu. 2021 · 2021
One transformer fits all distributions in multi-modal diffusion at scale
Fan Bao, Shen Nie, Kaiwen Xue, Chongxuan Li, Shi Pu, Yaole Wang, Gang Yue, Yue Cao, Hang Su, and Jun Zhu. 2023 · 2023
Later among the works it cites.
Yichao Cai, Yuhang Liu, Zhen Zhang, and Javen Qinfeng Shi. 2023 · 2023
Later among the works it cites.
Qlora: Efficient finetuning of quantized llms
Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer. 2023 · 2023
Later among the works it cites.
Clip-adapter: Better vision-language models with feature adapters
Peng Gao, Shijie Geng, Renrui Zhang, Teli Ma, Rongyao Fang, Yongfeng Zhang, Hongsheng Li, and Yu Qiao. 2023 · 2023
Later among the works it cites.
GPT understands, too
Xiao Liu, Yanan Zheng, Zhengxiao Du, Ming Ding, Yujie Qian, Zhilin Yang, and Jie Tang. 2023 · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Re-identification of individuals in genomic datasets using public face images
Rajagopal Venkatesaramani, Bradley A Malin, and Yevgeniy Vorobeychik. 2021 · 2021
Cited alongside, same era.
Feature purification: How adversarial training performs robust deep learning. In FOCS . IEEE, 977–988
Zeyuan Allen-Zhu and Yuanzhi Li. 2022 · 2022
Cited alongside, same era.
Exploring visual prompts for adapting large-scale models
Hyojin Bahng, Ali Jahanian, Swami Sankaranarayanan, and Phillip Isola. 2022 · 2022
Cited alongside, same era.
Towards artificial general intelligence via a multimodal foundation model
Nanyi Fei, Zhiwu Lu, Yizhao Gao, Guoxing Yang, Yuqi Huo, Jingyuan Wen, Haoyu Lu, Ruihua Song, Xin Gao, Tao Xiang, et al · 2022
Cited alongside, same era.
Listen and look: Multi-modal aggregation and co-attention network for video-audio retrieval. In ICME . 1–6
Xiaoshuai Hao, Wanqian Zhang, Dayan Wu, Fei Zhu, and Bo Li. 2022 · 2022
Cited alongside, same era.
Visual prompt tuning. In ECCV . 709–727
Menglin Jia, Luming Tang, Bor-Chun Chen, Claire Cardie, Serge Belongie, Bharath Hariharan, and Ser-Nam Lim. 2022 · 2022
Cited alongside, same era.
Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In ICML . 12888–12900
Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. 2022 · 2022
Cited alongside, same era.
Later among the works it cites.
UniAdapter: Unified Parameter-Efficient Transfer Learning for Cross-modal Modeling
Haoyu Lu, Mingyu Ding, Yuqi Huo, Guoxing Yang, Zhiwu Lu, Masayoshi Tomizuka, and Wei Zhan. 2023 · 2023
Later among the works it cites.
Parameter-efficient Tuning of Large-scale Multimodal Foundation Model. In NeurIPS . 15752–15774
Haixin Wang, Xinlong Yang, Jianlong Chang, Dian Jin, Jinan Sun, Shikun Zhang, Xiao Luo, and Qi Tian. 2023 · 2023
Later among the works it cites.
Adversarial prompt tuning for vision-language models
Jiaming Zhang, Xingjun Ma, Xin Wang, Lingyu Qiu, Jiaqi Wang, Yu-Gang Jiang, and Jitao Sang. 2023b · 2023
Later among the works it cites.
Adaptive budget allocation for parameter-efficient fine-tuning
Qingru Zhang, Minshuo Chen, Alexander Bukharin, Pengcheng He, Yu Cheng, Weizhu Chen, and Tuo Zhao. 2023a · 2023
Later among the works it cites.
On evaluating adversarial robustness of large vision-language models
Yunqing Zhao, Tianyu Pang, Chao Du, Xiao Yang, Chongxuan Li, Ngai-Man Cheung, and Min Lin. 2023 · 2023
Later among the works it cites.
TASAR: Transfer-based Attack on Skeletal Action Recognition
Yunfeng Diao, Baiqi Wu, Ruixuan Zhang, Ajian Liu, Xingxing Wei, Meng Wang, and He Wang. 2024 · 2024
Closest in time.
Uncertainty-aware alignment network for cross-domain video-text retrieval
Xiaoshuai Hao and Wanqian Zhang. 2024 · 2024
Closest in time.
Flipattack: Jailbreak llms via flipping
Yue Liu, Xiaoxin He, Miao Xiong, Jinlan Fu, Shumin Deng, and Bryan Hooi. 2024a · 2024
Closest in time.
AFLoRA: Adaptive Freezing of Low Rank Adaptation in Parameter Efficient Fine-Tuning of Large Models
Zeyu Liu, Souvik Kundu, Anni Li, Junrui Wan, Lianghao Jiang, and Peter Anthony Beerel. 2024b · 2024
Closest in time.
LISA: Layerwise Importance Sampling for Memory-Efficient Large Language Model Fine-Tuning
Rui Pan, Xiang Liu, Shizhe Diao, Renjie Pi, Jipeng Zhang, Chi Han, and Tong Zhang. 2024 · 2024
Closest in time.
Rushi Qiang, Ruiyi Zhang, and Pengtao Xie. 2024 · 2024
Closest in time.
LoRA Meets Dropout under a Unified Framework
Sheng Wang, Liheng Chen, Jiyue Jiang, Boyang Xue, Lingpeng Kong, and Chuan Wu. 2024 · 2024
Closest in time.
Fulllora-at: Efficiently boosting the robustness of pretrained vision transformers
Zheng Yuan, Jie Zhang, and Shiguang Shan. 2024 · 2024
Closest in time.
Galore: Memory-efficient llm training by gradient low-rank projection
Jiawei Zhao, Zhenyu Zhang, Beidi Chen, Zhangyang Wang, Anima Anandkumar, and Yuandong Tian. 2024 · 2024
Closest in time.
Multi-LoRA Composition for Image Generation
Ming Zhong, Yelong Shen, Shuohang Wang, Yadong Lu, Yizhu Jiao, Siru Ouyang, Donghan Yu, Jiawei Han, and Weizhu Chen. 2024 · 2024
Closest in time.
GuardReasoner: Towards Reasoning-based LLM Safeguards
Yue Liu, Hongcheng Gao, Shengfang Zhai, Jun Xia, Tianyi Wu, Zhiwei Xue, Yulin Chen, Kenji Kawaguchi, Jiaheng Zhang, and Bryan Hooi. 2025 · 2025
Closest in time.