Fetching the paper…
Reading the bibliography…
We introduce AutoAdvExBench, a benchmark to evaluate if large language models (LLMs) can autonomously exploit defenses to adversarial examples.
Using pre-training can improve model robustness and uncertainty, 2019
Dan Hendrycks, Kimin Lee, and Mantas Mazeika · 1901
Earlier work this paper cites.
Theoretically principled trade-off between robustness and accuracy, 2019
Hongyang Zhang, Yaodong Yu, Jiantao Jiao, Eric P. Xing, Laurent El Ghaoui, and Michael I. Jordan · 1901
Earlier work this paper cites.
Improving adversarial robustness via guided complement entropy, 2019
Hao-Yun Chen, Jhao-Hong Liang, Shih-Chieh Chang, Jia-Yu Pan, Yu-Ting Chen, Wei Wei, and Da-Cheng Juan · 1903
Earlier work this paper cites.
Gotta catch ’em all: Using honeypots to catch adversarial attacks on neural networks, 2019
Shawn Shan, Emily Wenger, Bolun Wang, Bo Li, Haitao Zheng, and Ben Y. Zhao · 1904
Earlier work this paper cites.
Rethinking softmax cross-entropy loss for adversarial robustness, 2019a
Tianyu Pang, Kun Xu, Yinpeng Dong, Chao Du, Ning Chen, and Jun Zhu · 1905
Earlier work this paper cites.
Defending against adversarial examples with k-nearest neighbor, 2019a
Chawin Sitawarin and David Wagner · 1906
Earlier work this paper cites.
Defending against adversarial examples with k-nearest neighbor
Chawin Sitawarin and David Wagner · 1906
Earlier work this paper cites.
Mixup inference: Better exploiting mixup to defend adversarial attacks
Tianyu Pang, Kun Xu, and Jun Zhu · 1909
Earlier work this paper cites.
Adversarial weight perturbation helps robust generalization, 2020
Dongxian Wu, Shu tao Xia, and Yisen Wang · 2004
Earlier work this paper cites.
Label smoothing and adversarial robustness, 2020a
Chaohao Fu, Hongbin Chen, Na Ruan, and Weijia Jia · 2009
Earlier work this paper cites.
Label smoothing and adversarial robustness
Chaohao Fu, Hongbin Chen, Na Ruan, and Weijia Jia · 2009
Earlier work this paper cites.
Metasploit: the penetration tester’s guide
David Kennedy, Jim O’gorman, Devon Kearns, and Mati Aharoni · 2011
Earlier work this paper cites.
Explaining and harnessing adversarial examples
Ian J Goodfellow, Jonathon Shlens, and Christian Szegedy · 2014
Earlier work this paper cites.
Distillation as a defense to adversarial perturbations against deep neural networks, 2015
Nicolas Papernot, Patrick McDaniel, Xi Wu, Somesh Jha, and Ananthram Swami · 2015
Earlier work this paper cites.
Towards evaluating the robustness of neural networks
Nicholas Carlini and David Wagner · 2017
Earlier work this paper cites.
Countering adversarial images using input transformations
Chuan Guo, Mayank Rana, Moustapha Cisse, and Laurens Van Der Maaten · 2017
Earlier work this paper cites.
Adversarial examples detection in deep networks with convolutional filter statistics
Xin Li and Fuxin Li · 2017
Earlier work this paper cites.
Towards deep learning models resistant to adversarial attacks, 2017
Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu · 2017
Earlier work this paper cites.
Magnet: a two-pronged defense against adversarial examples, 2017
Dongyu Meng and Hao Chen · 2017
Earlier work this paper cites.
Practical black-box attacks against machine learning
Nicolas Papernot, Patrick McDaniel, Ian Goodfellow, Somesh Jha, Z Berkay Celik, and Ananthram Swami · 2017
Earlier work this paper cites.
Feature squeezing: Detecting adversarial examples in deep neural networks, 2017
Weilin Xu, David Evans, and Yanjun Qi · 2017
Earlier work this paper cites.
Obfuscated gradients give a false sense of security: Circumventing defenses to adversarial examples
Anish Athalye, Nicholas Carlini, and David Wagner · 2018
Earlier work this paper cites.
Thermometer encoding: One hot way to resist adversarial examples
Jacob Buckman, Aurko Roy, Colin Raffel, and Ian Goodfellow · 2018
Earlier work this paper cites.
Stochastic activation pruning for robust adversarial defense, 2018
Guneet S. Dhillon, Kamyar Azizzadenesheli, Zachary C. Lipton, Jeremy Bernstein, Jean Kossaifi, Aran Khanna, and Anima Anandkumar · 2018
Earlier work this paper cites.
Adversarial logit pairing, 2018
Harini Kannan, Alexey Kurakin, and Ian Goodfellow · 2018
Earlier work this paper cites.
Characterizing adversarial subspaces using local intrinsic dimensionality, 2018
Xingjun Ma, Bo Li, Yisen Wang, Sarah M. Erfani, Sudanthi Wijewickrema, Grant Schoenebeck, Dawn Song, Michael E. Houle, and James Bailey · 2018
Earlier work this paper cites.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang · 2018
Earlier work this paper cites.
On evaluating adversarial robustness
Nicholas Carlini, Anish Athalye, Nicolas Papernot, Wieland Brendel, Jonas Rauber, Dimitris Tsipras, Ian Goodfellow, Aleksander Madry, and Alexey Kurakin · 2019
Earlier work this paper cites.
A new defense against adversarial images: Turning a weakness into a strength
Shengyuan Hu, Tao Yu, Chuan Guo, Wei-Lun Chao, and Kilian Q Weinberger · 2019
Earlier work this paper cites.
Feature distillation: Dnn-oriented jpeg compression against adversarial examples
Zihao Liu, Qi Liu, Tao Liu, Nuo Xu, Xue Lin, Yanzhi Wang, and Wujie Wen · 2019
Cited alongside, same era.
Barrage of random transforms for adversarially robust defense
Edward Raff, Jared Sylvester, Steven Forsyth, and Mark McLean · 2019
Cited alongside, same era.
Error correcting output codes improve probability estimation and adversarial robustness of deep neural networks
Gunjan Verma and Ananthram Swami · 2019
Cited alongside, same era.
Superglue: A stickier benchmark for general-purpose language understanding systems
Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman · 2019
Cited alongside, same era.
A partial break of the honeypots defense to catch adversarial attacks
Nicholas Carlini · 2020
Cited alongside, same era.
Stratified adversarial robustness with rejection, 2023
Jiefeng Chen, Jayaram Raghuram, Jihye Choi, Xi Wu, Yingyu Liang, and Somesh Jha · 2023
Later among the works it cites.
Decoupled kullback-leibler divergence loss, 2023
Jiequan Cui, Zhuotao Tian, Zhisheng Zhong, Xiaojuan Qi, Bei Yu, and Hanwang Zhang · 2023
Later among the works it cites.
Pentestgpt: An llm-empowered automatic penetration testing tool
Gelei Deng, Yi Liu, Víctor Mayoral-Vilches, Peng Liu, Yuekang Li, Yuan Xu, Tianwei Zhang, Yang Liu, Martin Pinzger, and Stefan Rass · 2023
Later among the works it cites.
The best defense is a good offense: Adversarial augmentation against adversarial attacks, 2023
Iuri Frosio and Jan Kautz · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Self-study course in evaluating adversarial robustness
Nicholas Carlini and Alex Kurakin · 2020
Cited alongside, same era.
Robustbench: a standardized adversarial robustness benchmark
Francesco Croce, Maksym Andriushchenko, Vikash Sehwag, Edoardo Debenedetti, Nicolas Flammarion, Mung Chiang, Prateek Mittal, and Matthias Hein · 2020
Cited alongside, same era.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu · 2020
Cited alongside, same era.
Empir: Ensembles of mixed precision deep networks for increased robustness against adversarial attacks
Sanchari Sen, Balaraman Ravindran, and Anand Raghunathan · 2020
Cited alongside, same era.
On adaptive attacks to adversarial example defenses
Florian Tramer, Nicholas Carlini, Wieland Brendel, and Aleksander Madry · 2020
Cited alongside, same era.
Improving adversarial robustness requires revisiting misclassified examples
Yisen Wang, Difan Zou, Jinfeng Yi, James Bailey, Xingjun Ma, and Quanquan Gu · 2020
Cited alongside, same era.
Andreas Happe and Jürgen Cito · 2023
Later among the works it cites.
An overview of catastrophic ai risks
Dan Hendrycks, Mantas Mazeika, and Thomas Woodside · 2023
Later among the works it cites.
Swe-bench: Can language models resolve real-world github issues?
Carlos E Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan · 2023
Later among the works it cites.
Improved adversarial training through adaptive instance-wise loss smoothing, 2023
Lin Li and Michael Spratling · 2023
Later among the works it cites.
Agentbench: Evaluating llms as agents
Xiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu, Xuanyu Lei, Hanyu Lai, Yu Gu, Hangliang Ding, Kaiwen Men, Kejuan Yang, et al · 2023
Later among the works it cites.
Gpqa: A graduate-level google-proof q&a benchmark
David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, and Samuel R Bowman · 2023
Later among the works it cites.
New adversarial image detection based on sentiment analysis, 2023
Yulong Wang, Tianxiang Li, Shenghong Li, Xin Yuan, and Wei Ni · 2023
Later among the works it cites.
Webarena: A realistic web environment for building autonomous agents
Shuyan Zhou, Frank F Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, et al · 2023
Later among the works it cites.
Swe-search: Enhancing software agents with monte carlo tree search and iterative refinement
Antonis Antoniades, Albert Örwall, Kexun Zhang, Yuxi Xie, Anirudh Goyal, and William Wang · 2024
Later among the works it cites.
Cyberseceval 2: A wide-ranging cybersecurity evaluation suite for large language models
Manish Bhatt, Sahana Chennabasappa, Yue Li, Cyrus Nikolaidis, Daniel Song, Shengye Wan, Faizan Ahmad, Cornelius Aschermann, Yaohui Chen, Dhaval Kapil, et al · 2024
Later among the works it cites.
Arc prize 2024: Technical report
Francois Chollet, Mike Knoop, Gregory Kamradt, and Bryan Landers · 2024
Later among the works it cites.
Agentdojo: A dynamic environment to evaluate attacks and defenses for llm agents
Edoardo Debenedetti, Jie Zhang, Mislav Balunović, Luca Beurer-Kellner, Marc Fischer, and Florian Tramèr · 2024
Later among the works it cites.
Sabre: Cutting through adversarial noise with adaptive spectral filtering and input reconstruction
Alec F Diallo and Paul Patras · 2024
Later among the works it cites.
Llm agents can autonomously hack websites
Richard Fang, Rohan Bindu, Akul Gupta, Qiusi Zhan, and Daniel Kang · 2024
Later among the works it cites.
Ensemble everything everywhere: Multi-scale aggregation for adversarial robustness
Stanislav Fort and Balaji Lakshminarayanan · 2024
Later among the works it cites.
Frontiermath: A benchmark for evaluating advanced mathematical reasoning in ai
Elliot Glazer, Ege Erdil, Tamay Besiroglu, Diego Chicharro, Evan Chen, Alex Gunning, Caroline Falkman Olsson, Jean-Stanislas Denain, Anson Ho, Emily de Oliveira Santos, et al · 2024
Later among the works it cites.
Can implicit bias imply adversarial robustness?
Hancheng Min and René Vidal · 2024
Later among the works it cites.
Learning to reason with llms
OpenAI · 2024
Later among the works it cites.
Openai o3 and o3-mini—12 days of openai: Day 12
OpenAI · 2024
Later among the works it cites.
Nyu ctf dataset: A scalable open-source benchmark dataset for evaluating llms in offensive security
Minghao Shao, Sofija Jancheska, Meet Udeshi, Brendan Dolan-Gavitt, Haoran Xi, Kimberly Milner, Boyuan Chen, Max Yin, Siddharth Garg, Prashanth Krishnamurthy, et al · 2024
Later among the works it cites.
Zachary S Siegel, Sayash Kapoor, Nitya Nagdir, Benedikt Stroebl, and Arvind Narayanan · 2024
Later among the works it cites.
Re-bench: Evaluating frontier ai r&d capabilities of language model agents against human experts
Hjalmar Wijk, Tao Lin, Joel Becker, Sami Jawhar, Neev Parikh, Thomas Broadley, Lawrence Chan, Michael Chen, Josh Clymer, Jai Dhyani, et al · 2024
Later among the works it cites.
Swe-agent: Agent-computer interfaces enable automated software engineering
John Yang, Carlos E Jimenez, Alexander Wettig, Kilian Lieret, Shunyu Yao, Karthik Narasimhan, and Ofir Press · 2024
Later among the works it cites.
Cybench: A framework for evaluating cybersecurity capabilities and risk of language models
Andy K Zhang, Neil Perry, Riya Dulepet, Eliot Jones, Justin W Lin, Joey Ji, Celeste Menders, Gashon Hussein, Samantha Liu, Donovan Jasper, et al · 2024
Later among the works it cites.