Fetching the paper…
Reading the bibliography…
Misaligned research objectives have considerably hindered progress in adversarial robustness research over the past decade.
Learning disjunction of conjunctions
Valiant, L. G · 1985
Earlier work this paper cites.
Learning in the Presence of Malicious Errors
Kearns, M. and Li, M · 1993
Earlier work this paper cites.
Kerckhoffs’ principle for intrusion detection
Mrdovic, S. and Perunicic, B · 2008
Earlier work this paper cites.
The security of machine learning
Barreno, M., Nelson, B., Joseph, A. D., and Tygar, J. D · 2010
Earlier work this paper cites.
Ethical hackers: putting on the white hat
Caldwell, T · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E · 2012
Earlier work this paper cites.
Intriguing properties of neural networks
Szegedy, C., Zaremba, W., Sutskever, I., Bruna, J., Erhan, D., Goodfellow, I. J., and Fergus, R · 2014
Earlier work this paper cites.
Explaining and harnessing adversarial examples
Goodfellow, I. J., Shlens, J., and Szegedy, C · 2015
Earlier work this paper cites.
Adversarial examples are not easily detected: Bypassing ten detection methods
Carlini, N. and Wagner, D · 2017
Earlier work this paper cites.
Reluplex: An efficient smt solver for verifying deep neural networks
Katz, G., Barrett, C., Dill, D. L., Julian, K., and Kochenderfer, M. J · 2017
Earlier work this paper cites.
Practical black-box attacks against machine learning
Papernot, N., McDaniel, P. D., Goodfellow, I. J., Jha, S., Celik, Z. B., and Swami, A · 2017
Earlier work this paper cites.
Evaluating robustness of neural networks with mixed integer programming
Tjeng, V., Xiao, K., and Tedrake, R · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A · 2017
Earlier work this paper cites.
On the effectiveness of interval bound propagation for training verifiably robust models
Gowal, S., Dvijotham, K., Stanforth, R., Bunel, R., Qin, C., Uesato, J., Arandjelovic, R., Mann, T., and Kohli, P · 2018
Earlier work this paper cites.
Physical adversarial examples for object detectors
Song, D., Eykholt, K., Evtimov, I., Fernandes, E., Li, B., Rahmati, A., Tramer, F., Prakash, A., and Kohno, T · 2018
Earlier work this paper cites.
Adversarial risk and the dangers of evaluating against weak attacks
Uesato, J., O’donoghue, B., Kohli, P., and Oord, A · 2018
Earlier work this paper cites.
On evaluating adversarial robustness
Carlini, N., Athalye, A., Papernot, N., Brendel, W., Rauber, J., Tsipras, D., Goodfellow, I., Madry, A., and Kurakin, A · 2019
Earlier work this paper cites.
Certified adversarial robustness via randomized smoothing
Cohen, J., Rosenfeld, E., and Kolter, Z · 2019
Earlier work this paper cites.
Show your work: Improved reporting of experimental results
Dodge, J., Gururangan, S., Card, D., Schwartz, R., and Smith, N. A · 2019
Earlier work this paper cites.
Testing robustness against unforeseen adversaries
Kaufmann, M., Kang, D., Sun, Y., Basart, S., Yin, X., Mazeika, M., Arora, A., Dziedzic, A., Boenisch, F., Brown, T., et al · 2019
Earlier work this paper cites.
Functional adversarial attacks
Laidlaw, C. and Feizi, S · 2019
Earlier work this paper cites.
Adversarial example games
Bose, J., Gidel, G., Berard, H., Cianflone, A., Vincent, P., Lacoste-Julien, S., and Hamilton, W · 2020
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Cited alongside, same era.
With little power comes great responsibility
Card, D., Henderson, P., Khandelwal, U., Jia, R., Mahowald, K., and Jurafsky, D · 2020
Cited alongside, same era.
Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacks
Croce, F. and Hein, M · 2020
Cited alongside, same era.
Robustbench: a standardized adversarial robustness benchmark
Croce, F., Andriushchenko, M., Sehwag, V., Debenedetti, E., Flammarion, N., Chiang, M., Mittal, P., and Hein, M · 2020
Cited alongside, same era.
A metric learning reality check
Musgrave, K., Belongie, S., and Lim, S.-N · 2020
Cited alongside, same era.
Defending against unforeseen failure modes with latent adversarial training
Casper, S., Schulze, L., Patel, O., and Hadfield-Menell, D · 2024
Later among the works it cites.
Privacy side channels in machine learning systems
Debenedetti, E., Severi, G., Carlini, N., Choquette-Choo, C. A., Jagielski, M., Nasr, M., Wallace, E., and Tramèr, F · 2024
Later among the works it cites.
Adversarial perturbations cannot reliably protect artists from generative ai
Hönig, R., Rando, J., Carlini, N., and Tramèr, F · 2024
Later among the works it cites.
Hurst, A., Lerer, A., Goucher, A. P., Perelman, A., Ramesh, A., Clark, A., Ostrow, A., Welihinda, A., Hayes, A., Radford, A., et al · 2024
Later among the works it cites.
Llm defenses are not robust to multi-turn human jailbreaks yet
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
On Adaptive Attacks to Adversarial Example Defenses
Tramer, F., Carlini, N., Brendel, W., and Madry, A · 2020
Cited alongside, same era.
Are we learning yet? a meta review of evaluation failures across machine learning
Liao, T., Taori, R., Raji, I. D., and Schmidt, L · 2021
Cited alongside, same era.
On the adversarial robustness of vision transformers
Shao, R., Shi, Z., Yi, J., Chen, P.-Y., and Hsieh, C.-J · 2021
Cited alongside, same era.
X-risk analysis for ai research
Hendrycks, D. and Mazeika, M · 2022
Cited alongside, same era.
Leakage and the reproducibility crisis in ml-based science
Kapoor, S. and Narayanan, A · 2022
Cited alongside, same era.
Are aligned neural networks adversarially aligned?
Carlini, N., Nasr, M., Choquette-Choo, C. A., Jagielski, M., Gao, I., Awadalla, A., Koh, P. W., Ippolito, D., Lee, K., Tramer, F., et al · 2023
Cited alongside, same era.
Jailbreaking black box large language models in twenty queries
Chao, P., Robey, A., Dobriban, E., Hassani, H., Pappas, G. J., and Wong, E · 2023
Cited alongside, same era.
Li, N., Han, Z., Steneker, I., Primack, W., Goodside, R., Zhang, H., Wang, Z., Menghini, C., and Yue, S · 2024
Later among the works it cites.
Tofu: A task of fictitious unlearning for llms
Maini, P., Feng, Z., Schwarzschild, A., Lipton, Z. C., and Kolter, J. Z · 2024
Later among the works it cites.
Harmbench: A standardized evaluation framework for automated red teaming and robust refusal
Mazeika, M., Phan, L., Yin, X., Zou, A., Wang, Z., Mu, N., Sakhaee, E., Li, N., Basart, S., Li, B., et al · 2024
Later among the works it cites.
On evaluating the durability of safeguards for open-weight llms
Qi, X., Wei, B., Carlini, N., Huang, Y., Xie, T., He, L., Jagielski, M., Nasr, M., Mittal, P., and Henderson, P · 2024
Later among the works it cites.
Revisiting the robust alignment of circuit breakers
Schwinn, L. and Geisler, S · 2024
Later among the works it cites.
Soft prompt threats: Attacking safety alignment and unlearning in open-source LLMs through the embedding space
Schwinn, L., Dobre, D., Xhonneux, S., Gidel, G., and Günnemann, S · 2024
Later among the works it cites.
Shi, L., Ma, C., Liang, W., Ma, W., and Vosoughi, S · 2024
Later among the works it cites.
Judgebench: A benchmark for evaluating llm-based judges
Tan, S., Zhuang, S., Montgomery, K., Tang, W. Y., Cuadron, A., Wang, C., Popa, R. A., and Stoica, I · 2024
Later among the works it cites.
Flrt: Fluent student-teacher redteaming
Thompson, T. B. and Sklar, M · 2024
Later among the works it cites.
Efficient adversarial training in llms with continuous attacks
Xhonneux, S., Sordoni, A., Günnemann, S., Gidel, G., and Schwinn, L · 2024
Later among the works it cites.
Improving alignment and robustness with short circuiting
Zou, A., Phan, L., Wang, J., Duenas, D., Lin, M., Andriushchenko, M., Wang, R., Kolter, Z., Fredrikson, M., and Hendrycks, D · 2024
Later among the works it cites.
Model tampering attacks enable more rigorous evaluations of llm capabilities
Che, Z., Casper, S., Kirk, R., Satheesh, A., Slocum, S., McKinney, L. E., Gandikota, R., Ewart, A., Rosati, D., Wu, Z., et al · 2025
Closest in time.
Exploring and mitigating adversarial manipulation of voting-based leaderboards
Huang, Y., Nasr, M., Angelopoulos, A., Carlini, N., Chiang, W.-L., Choquette-Choo, C. A., Ippolito, D., Jagielski, M., Lee, K., Liu, K. Z., et al · 2025
Closest in time.
Apple says collision in child-abuse hashing system is not a concern, August 2021
Levenson, E · 2025
Closest in time.
Adversarial ml problems are getting harder to solve and to evaluate
Rando, J., Zhang, J., Carlini, N., and Tramèr, F · 2025
Closest in time.
A probabilistic perspective on unlearning and alignment for large language models
Scholten, Y., Günnemann, S., and Schwinn, L · 2025
Closest in time.
Sharma, M., Tong, M., Mu, J., Wei, J., Kruthoff, J., Goodfriend, S., Ong, E., Peng, A., Agarwal, R., Anil, C., et al · 2025
Closest in time.