Fetching the paper…
Reading the bibliography…
As ML models become increasingly complex and integral to high-stakes domains such as finance and healthcare, they also become more susceptible to sophisticated adversarial attacks.
Constructing digital signatures from a one way function
Leslie Lamport · 1979
Earlier work this paper cites.
A digital signature scheme secure against adaptive chosen-message attacks
Shafi Goldwasser, Silvio Micali, and Ronald L Rivest · 1988
Earlier work this paper cites.
Approximation by superpositions of a sigmoidal function
George Cybenko · 1989
Earlier work this paper cites.
Universal one-way hash functions and their cryptographic applications
Moni Naor and Moti Yung · 1989
Earlier work this paper cites.
One-way functions are necessary and sufficient for secure signatures
John Rompel · 1990
Earlier work this paper cites.
Universal approximation bounds for superpositions of a sigmoidal function
Andrew R Barron · 1993
Earlier work this paper cites.
Approximation and estimation bounds for artificial neural networks
Andrew R Barron · 1994
Earlier work this paper cites.
Kleptography: Using cryptography against cryptography
Adam Young and Moti Yung · 1997
Earlier work this paper cites.
On the limits of steganography
Ross J Anderson and Fabien AP Petitcolas · 1998
Earlier work this paper cites.
A pseudorandom generator from any one-way function
Johan Håstad, Russell Impagliazzo, Leonid A Levin, and Michael Luby · 1999
Earlier work this paper cites.
On the (im) possibility of obfuscating programs
Boaz Barak, Oded Goldreich, Rusell Impagliazzo, Steven Rudich, Amit Sahai, Salil Vadhan, and Ke Yang · 2001
Earlier work this paper cites.
Can we obfuscate programs
Boaz Barak · 2002
Earlier work this paper cites.
Provably secure steganography
Nicholas J Hopper, John Langford, and Luis Von Ahn · 2002
Earlier work this paper cites.
Upper and lower bounds on black-box steganography
Nenad Dedić, Gene Itkis, Leonid Reyzin, and Scott Russell · 2005
Earlier work this paper cites.
Random features for large-scale kernel machines
Ali Rahimi and Benjamin Recht · 2007
Earlier work this paper cites.
Intriguing properties of neural networks
Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus · 2013
Earlier work this paper cites.
The power of depth for feedforward neural networks
Ronen Eldan and Ohad Shamir · 2016
Earlier work this paper cites.
Why deep neural networks for function approximation?
Shiyu Liang and Rayadurgam Srikant · 2016
Earlier work this paper cites.
Protecting software through obfuscation: Can it keep pace with progress in code analysis?
Sebastian Schrittwieser, Stefan Katzenbeisser, Johannes Kinder, Georg Merzdovnik, and Edgar Weippl · 2016
Earlier work this paper cites.
Benefits of depth in neural networks
Matus Telgarsky · 2016
Earlier work this paper cites.
Targeted backdoor attacks on deep learning systems using data poisoning
Xinyun Chen, Chang Liu, Bo Li, Kimberly Lu, and Dawn Song · 2017
Earlier work this paper cites.
Badnets: Identifying vulnerabilities in the machine learning model supply chain
Tianyu Gu, Brendan Dolan-Gavitt, and Siddharth Garg · 2017
Earlier work this paper cites.
Approximating continuous functions by relu nets of minimal width
Boris Hanin and Mark Sellke · 2017
Earlier work this paper cites.
The expressive power of neural networks: A view from the width
Zhou Lu, Hongming Pu, Feicheng Wang, Zhiqiang Hu, and Liwei Wang · 2017
Earlier work this paper cites.
Digital watermarking and steganography: fundamentals and techniques
Frank Y Shih · 2017
Earlier work this paper cites.
Machine learning models that remember too much
Congzheng Song, Thomas Ristenpart, and Vitaly Shmatikov · 2017
Earlier work this paper cites.
Depth-width tradeoffs in approximating natural functions with neural networks
Itay Safran and Ohad Shamir · 2017
Earlier work this paper cites.
Error bounds for approximations with deep relu networks
Dmitry Yarotsky · 2017
Earlier work this paper cites.
Synthesizing robust adversarial examples
Anish Athalye, Logan Engstrom, Andrew Ilyas, and Kevin Kwok · 2018
Earlier work this paper cites.
Membership inference attack against differentially private deep learning model
Md Atiqur Rahman, Tanzila Rahman, Robert Laganière, Noman Mohammed, and Yang Wang · 2018
Earlier work this paper cites.
Certified defenses against adversarial examples
Aditi Raghunathan, Jacob Steinhardt, and Percy Liang · 2018
Earlier work this paper cites.
Spectral signatures in backdoor attacks
Brandon Tran, Jerry Li, and Aleksander Madry · 2018
Cited alongside, same era.
Provable defenses against adversarial examples via the convex outer adversarial polytope
Eric Wong and Zico Kolter · 2018
Cited alongside, same era.
Adversarial examples from computational constraints
Sébastien Bubeck, Yin Tat Lee, Eric Price, and Ilya Razenshteyn · 2019
Cited alongside, same era.
Badnets: Evaluating backdooring attacks on deep neural networks
Tianyu Gu, Kang Liu, Brendan Dolan-Gavitt, and Siddharth Garg · 2019
Cited alongside, same era.
Adversarial examples are not bugs, they are features
Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras, Logan Engstrom, Brandon Tran, and Aleksander Madry · 2019
Cited alongside, same era.
Comprehensive privacy analysis of deep learning: Passive and active white-box inference attacks against centralized and federated learning
Handcrafted backdoors in deep neural networks
Sanghyun Hong, Nicholas Carlini, and Alexey Kurakin · 2022
Later among the works it cites.
When does bias transfer in transfer learning?
Hadi Salman, Saachi Jain, Andrew Ilyas, Logan Engstrom, Eric Wong, and Aleksander Madry · 2022
Later among the works it cites.
Exploring the universal vulnerability of prompt-based learning paradigm
Lei Xu, Yangyi Chen, Ganqu Cui, Hongcheng Gao, and Zhiyuan Liu · 2022
Later among the works it cites.
Neurocryptography. invited plenary talk at crypto’2023
Scott Aaronson · 2023
Later among the works it cites.
Badloss: Backdoor detection via loss dynamics
Neel Alex, Shoaib Ahmed Siddiqui, Amartya Sanyal, and David Krueger · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Milad Nasr, Reza Shokri, and Amir Houmansadr · 2019
Cited alongside, same era.
How do infinite width bounded norm networks look in function space?
Pedro Savarese, Itay Evron, Daniel Soudry, and Nathan Srebro · 2019
Cited alongside, same era.
Adversarial training for free!
Ali Shafahi, Mahyar Najibi, Mohammad Amin Ghiasi, Zheng Xu, John Dickerson, Christoph Studer, Larry S Davis, Gavin Taylor, and Tom Goldstein · 2019
Cited alongside, same era.
Poly-time universality and limitations of deep learning
Emmanuel Abbe and Colin Sandon · 2020
Cited alongside, same era.
Evading deepfake-image detectors with white-and black-box attacks
Nicholas Carlini and Hany Farid · 2020
Cited alongside, same era.
Adversarially robust learning could leverage computational hardness
Sanjam Garg, Somesh Jha, Saeed Mahloujifar, and Mahmoody Mohammad · 2020
Cited alongside, same era.
Nonparametric regression using deep neural networks with relu activation function
Johannes Schmidt-Hieber · 2020
Cited alongside, same era.
Miranda Christ, Sam Gunn, and Or Zamir · 2023
Later among the works it cites.
Extracting training data from diffusion models
Nicolas Carlini, Jamie Hayes, Milad Nasr, Matthew Jagielski, Vikash Sehwag, Florian Tramer, Borja Balle, Daphne Ippolito, and Eric Wallace · 2023
Later among the works it cites.
Composite backdoor attacks against large language models
Hai Huang, Zhengyu Zhao, Michael Backes, Yun Shen, and Yang Zhang · 2023
Later among the works it cites.
A watermark for large language models
John Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz, Ian Miers, and Tom Goldstein · 2023
Later among the works it cites.
Backdoor attacks for in-context learning with language models
Nikhil Kandpal, Matthew Jagielski, Florian Tramèr, and Nicholas Carlini · 2023
Later among the works it cites.
Rethinking backdoor attacks
Alaa Khaddaj, Guillaume Leclerc, Aleksandar Makelov, Kristian Georgiev, Hadi Salman, Andrew Ilyas, and Aleksander Madry · 2023
Later among the works it cites.
Notable: Transferable backdoor attacks against prompt-based nlp models
Kai Mei, Zheng Li, Zhenting Wang, Yang Zhang, and Shiqing Ma · 2023
Later among the works it cites.
Universal jailbreak backdoors from poisoned human feedback
Javier Rando and Florian Tramèr · 2023
Later among the works it cites.
Decodingtrust: A comprehensive assessment of trustworthiness in gpt models
Boxin Wang, Weixin Chen, Hengzhi Pei, Chulin Xie, Mintong Kang, Chenhui Zhang, Chejian Xu, Zidi Xiong, Ritik Dutta, Rylan Schaeffer, et al · 2023
Later among the works it cites.
Poisoning language models during instruction tuning
Alexander Wan, Eric Wallace, Sheng Shen, and Dan Klein · 2023
Later among the works it cites.
Instructions as backdoors: Backdoor vulnerabilities of instruction tuning for large language models
Jiashu Xu, Mingyu Derek Ma, Fei Wang, Chaowei Xiao, and Muhao Chen · 2023
Later among the works it cites.
Universal and transferable adversarial attacks on aligned language models
Andy Zou, Zifan Wang, J Zico Kolter, and Matt Fredrikson · 2023
Later among the works it cites.
Prompt as triggers for backdoor attack: Examining the vulnerability in language models
Shuai Zhao, Jinming Wen, Luu Anh Tuan, Junbo Zhao, and Jie Fu · 2023
Later among the works it cites.
Foundational challenges in assuring alignment and safety of large language models
Usman Anwar, Abulhair Saparov, Javier Rando, Daniel Paleka, Miles Turpin, Peter Hase, Ekdeep Singh Lubana, Erik Jenner, Stephen Casper, Oliver Sourbut, et al · 2024
Closest in time.
Pseudorandom error-correcting codes
Miranda Christ and Sam Gunn · 2024
Closest in time.
Sleeper agents: Training deceptive llms that persist through safety training
Evan Hubinger, Carson Denison, Jesse Mu, Mike Lambert, Meg Tong, Monte MacDiarmid, Tamera Lanham, Daniel M Ziegler, Tim Maxwell, Newton Cheng, et al · 2024
Closest in time.
Talk too much: Poisoning large language models under token limit
Jiaming He, Wenbo Jiang, Guanyu Hou, Wenshu Fan, Rui Zhang, and Hongwei Li · 2024
Closest in time.
Label poisoning is all you need
Rishi Jha, Jonathan Hayase, and Sewoong Oh · 2024
Closest in time.
Badedit: Backdooring large language models by model editing
Yanzhou Li, Tianlin Li, Kangjie Chen, Jian Zhang, Shangqing Liu, Wenhan Wang, Tianwei Zhang, and Yang Liu · 2024
Closest in time.
Competition report: Finding universal jailbreak backdoors in aligned llms
Javier Rando, Francesco Croce, Kryštof Mitka, Stepan Shabalin, Maksym Andriushchenko, Nicolas Flammarion, and Florian Tramèr · 2024
Closest in time.
Privacy backdoors: Enhancing membership inference through poisoning pre-trained models
Yuxin Wen, Leo Marchyok, Sanghyun Hong, Jonas Geiping, Tom Goldstein, and Nicholas Carlini · 2024
Closest in time.
Badchain: Backdoor chain-of-thought prompting for large language models
Zhen Xiang, Fengqing Jiang, Zidi Xiong, Bhaskar Ramasubramanian, Radha Poovendran, and Bo Li · 2024
Closest in time.
A comprehensive overview of backdoor attacks in large language models within communication networks
Haomiao Yang, Kunlan Xiang, Mengyu Ge, Hongwei Li, Rongxing Lu, and Shui Yu · 2024
Closest in time.
Excuse me, sir? your language model is leaking (information)
Or Zamir · 2024
Closest in time.
Universal vulnerabilities in large language models: Backdoor attacks for in-context learning
Shuai Zhao, Meihuizi Jia, Luu Anh Tuan, Fengjun Pan, and Jinming Wen · 2024
Closest in time.