Fetching the paper…
Reading the bibliography…
The systems and software powered by Large Language Models (LLMs) and Multi-Modal LLMs (MLLMs) have played a critical role in numerous scenarios.
WordNet: a lexical database for English
George A Miller. 1995 · 1995
Earlier work this paper cites.
Similarity measures for text document clustering. In Proceedings of the sixth new zealand computer science research student conference (NZCSRSC2008), Christchurch, New Zealand , Vol. 4. 9–56
Anna Huang et al · 2008
Earlier work this paper cites.
Cohen’s kappa coefficient as a performance measure for feature selection. In FUZZ-IEEE 2010, IEEE International Conference on Fuzzy Systems, Barcelona, Spain, 18-23 July, 2010, Proceedings . IEEE, 1–8
Susana M. Vieira, Uzay Kaymak, and João M. C. Sousa. 2010 · 2010
Earlier work this paper cites.
Understanding bag-of-words model: a statistical framework
Yin Zhang, Rong Jin, and Zhi-Hua Zhou. 2010 · 2010
Earlier work this paper cites.
Interrater reliability: the kappa statistic
Mary L McHugh. 2012 · 2012
Earlier work this paper cites.
Explaining and harnessing adversarial examples
Ian J Goodfellow, Jonathon Shlens, and Christian Szegedy. 2014 · 2014
Earlier work this paper cites.
Hyperopt-Sklearn: Automatic Hyperparameter Configuration for Scikit-Learn.. In Scipy . 32–37
Brent Komer, James Bergstra, and Chris Eliasmith. 2014 · 2014
Earlier work this paper cites.
Word2Vec
Kenneth Ward Church. 2017 · 2017
Earlier work this paper cites.
Countering adversarial images using input transformations
Chuan Guo, Mayank Rana, Moustapha Cisse, and Laurens Van Der Maaten. 2017 · 2017
Earlier work this paper cites.
Towards deep learning models resistant to adversarial attacks
Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu. 2017 · 2017
Earlier work this paper cites.
On detecting adversarial perturbations
Jan Hendrik Metzen, Tim Genewein, Volker Fischer, and Bastian Bischoff. 2017 · 2017
Earlier work this paper cites.
Ontonotes: Large scale multi-layer, multi-lingual, distributed annotation
Sameer Pradhan and Lance Ramshaw. 2017 · 2017
Earlier work this paper cites.
Feature squeezing: Detecting adversarial examples in deep neural networks
Weilin Xu, David Evans, and Yanjun Qi. 2017 · 2017
Earlier work this paper cites.
Obfuscated gradients give a false sense of security: Circumventing defenses to adversarial examples. In International conference on machine learning . PMLR, 274–283
Anish Athalye, Nicholas Carlini, and David Wagner. 2018 · 2018
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Earlier work this paper cites.
Mma training: Direct input space margin maximization through adversarial training
Gavin Weiguang Ding, Yash Sharma, Kry Yik Chau Lui, and Ruitong Huang. 2018 · 2018
Earlier work this paper cites.
Unsupervised representation learning by predicting image rotations
Spyros Gidaris, Praveer Singh, and Nikos Komodakis. 2018 · 2018
Earlier work this paper cites.
Adversarial examples in the physical world
Alexey Kurakin, Ian J Goodfellow, and Samy Bengio. 2018 · 2018
Earlier work this paper cites.
On the dimensionality of word embedding
Zi Yin and Yuanyuan Shen. 2018 · 2018
Earlier work this paper cites.
Certified adversarial robustness via randomized smoothing. In international conference on machine learning . PMLR, 1310–1320
Jeremy Cohen, Elan Rosenfeld, and Zico Kolter. 2019 · 2019
Earlier work this paper cites.
Augmix: A simple data processing method to improve robustness and uncertainty
Dan Hendrycks, Norman Mu, Ekin D Cubuk, Barret Zoph, Justin Gilmer, and Balaji Lakshminarayanan. 2019 · 2019
Earlier work this paper cites.
Improving robustness without sacrificing accuracy with patch gaussian augmentation
Raphael Gontijo Lopes, Dong Yin, Ben Poole, Justin Gilmer, and Ekin D Cubuk. 2019 · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Earlier work this paper cites.
Adversarial defense by stratified convolutional sparse coding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 11447–11456
Bo Sun, Nian-hsuan Tsai, Fangchen Liu, Ronald Yu, and Hao Su. 2019 · 2019
Earlier work this paper cites.
Triple-to-text: Converting RDF triples into high-quality natural languages via optimizing an inverse KL divergence. In Proceedings of the 42nd International ACM SIGIR Conference on Research and Development in Information Retrieval . 455–464
Yaoming Zhu, Juncheng Wan, Zhiming Zhou, Liheng Chen, Lin Qiu, Weinan Zhang, Xin Jiang, and Yong Yu. 2019 · 2019
Earlier work this paper cites.
Square attack: a query-efficient black-box adversarial attack via random search. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XXIII . Springer, 484–501
Maksym Andriushchenko, Francesco Croce, Nicolas Flammarion, and Matthias Hein. 2020 · 2020
Earlier work this paper cites.
Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacks. In International conference on machine learning . PMLR, 2206–2216
Francesco Croce and Matthias Hein. 2020 · 2020
Earlier work this paper cites.
Randaugment: Practical automated data augmentation with a reduced search space. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition workshops . 702–703
Ekin D Cubuk, Barret Zoph, Jonathon Shlens, and Quoc V Le. 2020 · 2020
Earlier work this paper cites.
Certified defense to image transformations via randomized smoothing
Marc Fischer, Maximilian Baader, and Martin Vechev. 2020 · 2020
Earlier work this paper cites.
Watch out! motion is blurring the vision of your deep neural networks
Qing Guo, Felix Juefei-Xu, Xiaofei Xie, Lei Ma, Jian Wang, Bing Yu, Wei Feng, and Yang Liu. 2020 · 2020
Earlier work this paper cites.
Adam Hare, Yu Chen, Yinan Liu, Zhenming Liu, and Christopher G Brinton. 2020 · 2020
Earlier work this paper cites.
Momentum contrast for unsupervised visual representation learning. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 9729–9738
Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick. 2020 · 2020
Earlier work this paper cites.
Adv-watermark: A novel watermark perturbation for adversarial examples. In Proceedings of the 28th ACM International Conference on Multimedia . 1579–1587
Xiaojun Jia, Xingxing Wei, Xiaochun Cao, and Xiaoguang Han. 2020 · 2020
Earlier work this paper cites.
Systematic literature reviews in software engineering - enhancement of the study selection process using Cohen’s Kappa statistic
Jorge E. Pérez, Jessica Díaz, Javier García Martín, and Bernardo Tabuenca. 2020 · 2020
Earlier work this paper cites.
On finding similar verses from the Holy Quran using word embeddings. In 2020 International Conference on Emerging Trends in Smart Technologies (ICETST) . IEEE, 1–6
Sumaira Saeed, Sajjad Haider, and Quratulain Rajput. 2020 · 2020
Earlier work this paper cites.
SAFER: A Structure-free Approach for Certified Robustness to Adversarial Word Substitutions. In Annual Meeting of the Association for Computational Linguistics (ACL)
Mao Ye, Chengyue Gong, and Qiang Liu. 2020 · 2020
Earlier work this paper cites.
Demystifying limited adversarial transferability in automatic speech recognition systems. In International Conference on Learning Representations (ICLR)
Hadi Abdullah, Aditya Karlekar, Vincent Bindschaedler, and Patrick Traynor. 2021 · 2021
Earlier work this paper cites.
Evaluating large language models trained on code
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al · 2021
Earlier work this paper cites.
Achieving rotational invariance with bessel-convolutional neural networks
Valentin Delchevalerie, Adrien Bibal, Benoît Frénay, and Alexandre Mayer. 2021 · 2021
Earlier work this paper cites.
Black-box detection of backdoor attacks with limited information and data. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 16482–16491
Yinpeng Dong, Xiao Yang, Zhijie Deng, Tianyu Pang, Zihao Xiao, Hang Su, and Jun Zhu. 2021 · 2021
Earlier work this paper cites.
Yunpeng Gong, Liqing Huang, and Lifei Chen. 2021 · 2021
Earlier work this paper cites.
AEDA: An Easier Data Augmentation Technique for Text Classification. In Findings of the Association for Computational Linguistics: EMNLP 2021 . 2748–2754
Akbar Karimi, Leonardo Rossi, and Andrea Prati. 2021 · 2021
Earlier work this paper cites.
Reading Isn’t Believing: Adversarial Attacks On Multi-Modal Neurons
David A Noever and Samantha E Miller Noever. 2021 · 2021
Earlier work this paper cites.
Helper-based adversarial training: Reducing excessive margin to achieve a better accuracy vs. robustness trade-off. In ICML 2021 Workshop on Adversarial Machine Learning
Rahul Rade and Seyed-Mohsen Moosavi-Dezfooli. 2021 · 2021
Earlier work this paper cites.
Robust learning meets generative models: Can proxy distributions improve adversarial robustness?
Vikash Sehwag, Saeed Mahloujifar, Tinashe Handina, Sihui Dai, Chong Xiang, Mung Chiang, and Prateek Mittal. 2021 · 2021
Earlier work this paper cites.
AVA: Adversarial vignetting attack against visual recognition
Binyu Tian, Felix Juefei-Xu, Qing Guo, Xiaofei Xie, Xiaohong Li, and Yang Liu. 2021 · 2021
Cited alongside, same era.
Directional self-supervised learning for heavy image augmentations. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 16692–16701
Yalong Bai, Yifan Yang, Wei Zhang, and Tao Mei. 2022 · 2022
Cited alongside, same era.
A survey on data augmentation for text classification
Markus Bayer, Marc-André Kaufhold, and Christian Reuter. 2022 · 2022
Cited alongside, same era.
Can you spot the chameleon? adversarially camouflaging images from co-salient object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 2150–2159
Ruijun Gao, Qing Guo, Felix Juefei-Xu, Hongkai Yu, Huazhu Fu, Wei Feng, Yang Liu, and Song Wang. 2022 · 2022
Cited alongside, same era.
Prompt injection attacks against GPT-3
Delimiters won’t save you from prompt injection
Simon Willison. 2023 · 2023
Closest in time.
Hacker Reveals Microsoft’s New AI-Powered Bing Chat Search Secrets
Davey Winder. 2023 · 2023
Closest in time.
{ \{ KENKU } \} : Towards Efficient and Stealthy Black-box Adversarial Attacks against { \{ ASR } \} Systems. In 32nd USENIX Security Symposium (USENIX Security 23) . 247–264
Xinghui Wu, Shiqing Ma, Chao Shen, Chenhao Lin, Qian Wang, Qi Li, and Yuan Rao. 2023 · 2023
Closest in time.
Defending chatgpt against jailbreak attack via self-reminders
Yueqi Xie, Jingwei Yi, Jiawei Shao, Justin Curl, Lingjuan Lyu, Qifeng Chen, Xing Xie, and Fangzhao Wu. 2023 · 2023
Closest in time.
Instructions as Backdoors: Backdoor Vulnerabilities of Instruction Tuning for Large Language Models
Jiashu Xu, Mingyu Derek Ma, Fei Wang, Chaowei Xiao, and Muhao Chen. 2023 · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Riley Goodside. 2022 · 2022
Cited alongside, same era.
Gsmooth: Certified robustness against semantic transformations via generalized randomized smoothing. In International Conference on Machine Learning . PMLR, 8465–8483
Zhongkai Hao, Chengyang Ying, Yinpeng Dong, Hang Su, Jian Song, and Jun Zhu. 2022 · 2022
Cited alongside, same era.
DISCO: Adversarial Defense with Local Implicit Functions
Chih-Hui Ho and Nuno Vasconcelos. 2022 · 2022
Cited alongside, same era.
Complex backdoor detection by symmetric feature differencing. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 15003–15013
Yingqi Liu, Guangyu Shen, Guanhong Tao, Zhenting Wang, Shiqing Ma, and Xiangyu Zhang. 2022 · 2022
Cited alongside, same era.
Software engineering for AI-based systems: a survey
Silverio Martínez-Fernández, Justus Bogner, Xavier Franch, Marc Oriol, Julien Siebert, Adam Trendowicz, Anna Maria Vollmer, and Stefan Wagner. 2022 · 2022
Cited alongside, same era.
Data augmentation: A comprehensive survey of modern approaches
Alhassan Mumuni and Fuseini Mumuni. 2022 · 2022
Cited alongside, same era.
Diffusion models for adversarial purification
Weili Nie, Brandon Guo, Yujia Huang, Chaowei Xiao, Arash Vahdat, and Anima Anandkumar. 2022 · 2022
Cited alongside, same era.
Ignore Previous Prompt: Attack Techniques For Language Models. In NeurIPS ML Safety Workshop
Fábio Perez and Ian Ribeiro. 2022 · 2022
Cited alongside, same era.
Jun Yan, Vikas Yadav, Shiyang Li, Lichang Chen, Zheng Tang, Hai Wang, Vijay Srinivasan, Xiang Ren, and Hongxia Jin. 2023 · 2023
Closest in time.
Benchmarking and defending against indirect prompt injection attacks on large language models
Jingwei Yi, Yueqi Xie, Bin Zhu, Keegan Hines, Emre Kiciman, Guangzhong Sun, Xing Xie, and Fangzhao Wu. 2023 · 2023
Closest in time.
Gptfuzzer: Red teaming large language models with auto-generated jailbreak prompts
Jiahao Yu, Xingwei Lin, and Xinyu Xing. 2023 · 2023
Closest in time.
Certified robustness to text adversarial attacks by randomized [mask]
Jiehang Zeng, Jianhan Xu, Xiaoqing Zheng, and Xuanjing Huang. 2023 · 2023
Closest in time.
Judging LLM-as-a-judge with MT-Bench and Chatbot Arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric. P Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2023 · 2023
Closest in time.
Instruction-Following Evaluation for Large Language Models
Jeffrey Zhou, Tianjian Lu, Swaroop Mishra, Siddhartha Brahma, Sujoy Basu, Yi Luan, Denny Zhou, and Le Hou. 2023 · 2023
Closest in time.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny. 2023a · 2023
Closest in time.
AutoDAN: Automatic and Interpretable Adversarial Attacks on Large Language Models
Sicheng Zhu, Ruiyi Zhang, Bang An, Gang Wu, Joe Barrow, Zichao Wang, Furong Huang, Ani Nenkova, and Tong Sun. 2023b · 2023
Closest in time.
Universal and transferable adversarial attacks on aligned language models
Andy Zou, Zifan Wang, J Zico Kolter, and Matt Fredrikson. 2023 · 2023
Closest in time.
Hands-on AI Demos for Human Resources
2024 · 2024
Closest in time.
spaCy: Industrial-strength Natural Language Processing in Python
2024 · 2024
Closest in time.
The Website of JailGuard
2024 · 2024
Closest in time.
Wizard-Vicuna-13B-Uncensored
2024 · 2024
Closest in time.
Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
Maksym Andriushchenko, Francesco Croce, and Nicolas Flammarion. 2024 · 2024
Closest in time.
StruQ: Defending Against Prompt Injection with Structured Queries
Sizhe Chen, Julien Piet, Chawin Sitawarin, and David Wagner. 2024b · 2024
Closest in time.
OR-Bench: An Over-Refusal Benchmark for Large Language Models
Justin Cui, Wei-Lin Chiang, Ion Stoica, and Cho-Jui Hsieh. 2024 · 2024
Closest in time.
Alpacafarm: A simulation framework for methods that learn from human feedback
Yann Dubois, Chen Xuechen Li, Rohan Taori, Tianyi Zhang, Ishaan Gulrajani, Jimmy Ba, Carlos Guestrin, Percy S Liang, and Tatsunori B Hashimoto. 2024 · 2024
Closest in time.
Coercing LLMs to do and reveal (almost) anything
Jonas Geiping, Alex Stein, Manli Shu, Khalid Saifullah, Yuxin Wen, and Tom Goldstein. 2024 · 2024
Closest in time.
Eyes closed, safety on: Protecting multimodal llms via image-to-text transformation
Yunhao Gou, Kai Chen, Zhili Liu, Lanqing Hong, Hang Xu, Zhenguo Li, Dit-Yan Yeung, James T Kwok, and Yu Zhang. 2024 · 2024
Closest in time.
Rethinking software engineering in the era of foundation models: A curated catalogue of challenges in the development of trustworthy fmware. In Companion Proceedings of the 32nd ACM International Conference on the Foundations of Software Engineering . 294–305
Ahmed E Hassan, Dayi Lin, Gopi Krishnan Rajbahadur, Keheliya Gallaba, Filipe Roseiro Cogo, Boyuan Chen, Haoxiang Zhang, Kishanthan Thangarajah, Gustavo Oliva, Jiahuei Lin, et al · 2024
Closest in time.
Comparative Performance of Advanced NLP Models and LLMs in Multilingual Geo-Entity Detection. In Proceedings of the Cognitive Models and Artificial Intelligence Conference . 106–110
Kalin Kopanov. 2024 · 2024
Closest in time.
Visual instruction tuning
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2024c · 2024
Closest in time.
Automatic and Universal Prompt Injection Attacks against Large Language Models
Xiaogeng Liu, Zhiyuan Yu, Yizhe Zhang, Ning Zhang, and Chaowei Xiao. 2024d · 2024
Closest in time.
A Hitchhiker’s Guide to Jailbreaking ChatGPT via Prompt Engineering. In Proceedings of the 4th International Workshop on Software Engineering and AI for Data Quality in Cyber-Physical Systems/Internet of Things, SEA4DQ 2024, Porto de Galinhas, Brazil, 15 July 2024 , Tim Menzies, Bowen Xu, Hong Jin Kang, and Jie M. Zhang (Eds.). ACM, 12–21
Yi Liu, Gelei Deng, Zhengzi Xu, Yuekang Li, Yaowen Zheng, Ying Zhang, Lida Zhao, Tianwei Zhang, and Kailong Wang. 2024a · 2024
Closest in time.
Meet Your New Assistant: Meta AI, Built With Llama 3
Meta. 2024 · 2024
Closest in time.
Microsoft Copilot
Microsoft. 2024 · 2024
Closest in time.
Jailbreaking attack against multimodal large language model
Zhenxing Niu, Haodong Ren, Xinbo Gao, Gang Hua, and Rong Jin. 2024 · 2024
Closest in time.
Hello GPT-4o
OpenAI. 2024 · 2024
Closest in time.
MLLM-Protector: Ensuring MLLM’s Safety without Hurting Performance
Renjie Pi, Tianyang Han, Yueqi Xie, Rui Pan, Qing Lian, Hanze Dong, Jipeng Zhang, and Tong Zhang. 2024 · 2024
Closest in time.
Visual adversarial examples jailbreak aligned large language models. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 38. 21527–21536
Xiangyu Qi, Kaixuan Huang, Ashwinee Panda, Peter Henderson, Mengdi Wang, and Prateek Mittal. 2024 · 2024
Closest in time.
Prompt Stealing Attacks Against Large Language Models
Zeyang Sha and Yang Zhang. 2024 · 2024
Closest in time.
Rapid Optimization for Jailbreaking LLMs via Subconscious Exploitation and Echopraxia
Guangyu Shen, Siyuan Cheng, Kaiyuan Zhang, Guanhong Tao, Shengwei An, Lu Yan, Zhuo Zhang, Shiqing Ma, and Xiangyu Zhang. 2024b · 2024
Closest in time.
Large Language Models as Software Components: A Taxonomy for LLM-Integrated Applications
Irene Weber. 2024 · 2024
Closest in time.
Jailbroken: How does llm safety training fail?
Alexander Wei, Nika Haghtalab, and Jacob Steinhardt. 2024 · 2024
Closest in time.
LLMs Can Defend Themselves Against Jailbreaking in a Practical Manner: A Vision Paper
Daoyuan Wu, Shuai Wang, Yang Liu, and Ning Liu. 2024 · 2024
Closest in time.
LLM Jailbreak Attack versus Defense Techniques–A Comprehensive Study
Zihao Xu, Yi Liu, Gelei Deng, Yuekang Li, and Stjepan Picek. 2024 · 2024
Closest in time.
Goal-guided Generative Prompt Injection Attack on Large Language Models
Chong Zhang, Mingyu Jin, Qinkai Yu, Chengzhi Liu, Haochen Xue, and Xiaobo Jin. 2024 · 2024
Closest in time.
Robust prompt optimization for defending language models against jailbreaking attacks
Andy Zhou, Bo Li, and Haohan Wang. 2024 · 2024
Closest in time.
Man who exploded Cybertruck in Las Vegas used ChatGPT in planning, police say
The Associated Press. 2025 · 2025
Closest in time.