Intriguing properties of neural networks
Original
C. Szegedy, W. Zaremba, I. Sutskever, J. Bruna, D. Erhan, I. Goodfellow, and R. Fergus · 2013
Earlier work this paper cites.
Explaining and harnessing adversarial examples
Original
I. J. Goodfellow, J. Shlens, and C. Szegedy · 2014
Earlier work this paper cites.
Vqa: Visual question answering
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh · 2015
Earlier work this paper cites.
Research on the technology of ios jailbreak
F. Liu, K.-S. Liu, C. Chang, and Y. Wang · 2016
Earlier work this paper cites.
Adversarial examples are not easily detected: Bypassing ten detection methods
N. Carlini and D. Wagner · 2017
Earlier work this paper cites.
Hotflip: White-box adversarial examples for text classification
Original
J. Ebrahimi, A. Rao, D. Lowd, and D. Dou · 2017
Earlier work this paper cites.
Towards deep learning models resistant to adversarial attacks
Original
A. Madry, A. Makelov, L. Schmidt, D. Tsipras, and A. Vladu · 2017
Earlier work this paper cites.
Practical black-box attacks against machine learning
N. Papernot, P. McDaniel, I. Goodfellow, S. Jha, Z. B. Celik, and A. Swami · 2017
Earlier work this paper cites.
Generating natural adversarial examples
Original
Z. Zhao, D. Dua, and S. Singh · 2017
Earlier work this paper cites.
Generating natural language adversarial examples
Original
M. Alzantot, Y. Sharma, A. Elgohary, B.-J. Ho, M. Srivastava, and K.-W. Chang · 2018
Earlier work this paper cites.
Obfuscated gradients give a false sense of security: Circumventing defenses to adversarial examples
A. Athalye, N. Carlini, and D. Wagner · 2018
Earlier work this paper cites.
Audio adversarial examples: Targeted attacks on speech-to-text
N. Carlini and D. Wagner · 2018
Earlier work this paper cites.
Black-box adversarial attacks with limited queries and information
A. Ilyas, L. Engstrom, A. Athalye, and J. Lin · 2018
Earlier work this paper cites.
Certified adversarial robustness via randomized smoothing
J. Cohen, E. Rosenfeld, and Z. Kolter · 2019
Earlier work this paper cites.
The curious case of neural text degeneration
Original
A. Holtzman, J. Buys, L. Du, M. Forbes, and Y. Choi · 2019
Earlier work this paper cites.
Universal adversarial triggers for attacking and analyzing nlp
Original
E. Wallace, S. Feng, N. Kandpal, M. Gardner, and S. Singh · 2019
Earlier work this paper cites.
From recognition to cognition: Visual commonsense reasoning
R. Zellers, Y. Bisk, A. Farhadi, and Y. Choi · 2019
Earlier work this paper cites.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Earlier work this paper cites.
Robustbench: a standardized adversarial robustness benchmark
Original
F. Croce, M. Andriushchenko, V. Sehwag, E. Debenedetti, N. Flammarion, M. Chiang, P. Mittal, and M. Hein · 2020
Earlier work this paper cites.
Realtoxicityprompts: Evaluating neural toxic degeneration in language models
Original
S. Gehman, S. Gururangan, M. Sap, Y. Choi, and N. A. Smith · 2020
Earlier work this paper cites.
Detoxify
L. Hanu and Unitary team · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
J. Ho, A. Jain, and P. Abbeel · 2020
Earlier work this paper cites.
Autoprompt: Eliciting knowledge from language models with automatically generated prompts
Original
T. Shin, Y. Razeghi, R. L. Logan IV, E. Wallace, and S. Singh · 2020
Earlier work this paper cites.
On the opportunities and risks of foundation models
Original
R. Bommasani, D. A. Hudson, E. Adeli, R. Altman, S. Arora, S. von Arx, M. S. Bernstein, J. Bohg, A. Bosselut, E. Brunskill, et al · 2021
Earlier work this paper cites.
Invisible for both camera and lidar: Security of multi-sensor fusion based perception in autonomous driving under physical-world attacks
Y. Cao, N. Wang, C. Xiao, D. Yang, J. Fang, R. Yang, Q. A. Chen, M. Liu, and B. Li · 2021
Earlier work this paper cites.