Fetching the paper…
Reading the bibliography…
The rapid proliferation of open-source language models significantly increases the risks of downstream backdoor attacks.
Hijacking Malaria Simulators with Probabilistic Programming, May 2019
B. Gram-Hansen, C. Schroeder de Witt, T. Rainforth, P. H. S. Torr, Y. W. Teh, and A. G. Baydin · 1905
Earlier work this paper cites.
HuggingFace’s Transformers: State-of-the-art Natural Language Processing, July 2020
T. Wolf, L. Debut, V. Sanh, J. Chaumond, C. Delangue, A. Moi, P. Cistac, T. Rault, R. Louf, M. Funtowicz, J. Davison, S. Shleifer, P. von Platen, C. Ma, Y. Jernite, J. Plu, C. Xu, T. L. Scao, S. Gugger, M. Drame, Q. Lhoest, and A. M. Rush · 1910
Earlier work this paper cites.
The perceptron: A probabilistic model for information storage and organization in the brain
F. Rosenblatt · 1939
Earlier work this paper cites.
Foundations of Cryptography
O. Goldreich · 2001
Earlier work this paper cites.
Reducibility among Combinatorial Problems , pages 85–103
R. M. Karp · 2001
Earlier work this paper cites.
Positive results and techniques for obfuscation
B. Lynn, M. Prabhakaran, and A. Sahai · 2004
Earlier work this paper cites.
On best-possible obfuscation
S. Goldwasser and G. N. Rothblum · 2007
Earlier work this paper cites.
Random Features for Large-Scale Kernel Machines
A. Rahimi and B. Recht · 2007
Earlier work this paper cites.
Obfuscating point functions with multibit output
R. Canetti and R. R. Dakdouk · 2008
Earlier work this paper cites.
On the secure hash algorithm family
W. Penard and T. van Werkhoven · 2008
Earlier work this paper cites.
On the (im)possibility of obfuscating programs
B. Barak, O. Goldreich, R. Impagliazzo, S. Rudich, A. Sahai, S. Vadhan, and K. Yang · 2012
Earlier work this paper cites.
Reusable fuzzy extractors for low-entropy distributions
R. Canetti, B. Fuller, O. Paneth, L. Reyzin, and A. Smith · 2016
Earlier work this paper cites.
Protecting software through obfuscation: Can it keep pace with progress in code analysis?
S. Schrittwieser, S. Katzenbeisser, J. Kinder, G. Merzdovnik, and E. Weippl · 2016
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library, 2019
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, A. Desmaison, A. Köpf, E. Yang, Z. DeVito, M. Raison, A. Tejani, S. Chilamkurthy, B. Steiner, L. Fang, J. Bai, and S. Chintala · 2019
Earlier work this paper cites.
Interpreting gpt: The logit lens
nostalgebraist · 2020
Earlier work this paper cites.
Bridging mode connectivity in loss landscapes and adversarial robustness, 2020
P. Zhao, P.-Y. Chen, P. Das, K. N. Ramamurthy, and X. Lin · 2020
Earlier work this paper cites.
The hardness of lwe and ring-lwe: A survey
D. Balbás · 2021
Earlier work this paper cites.
Eliciting latent knowledge: How to tell if your eyes deceive you, December 2021
P. Christiano, A. Cotra, and M. Xu · 2021
Earlier work this paper cites.
An overview of backdoor attacks against deep neural networks and possible defences, 2021
W. Guo, B. Tondi, and M. Barni · 2021
Earlier work this paper cites.
Thinking like transformers, 2021
G. Weiss, Y. Goldberg, and E. Yahav · 2021
Cited alongside, same era.
Nonmalleable digital lockers and robust fuzzy extractors in the plain model
D. Apon, C. Cachet, B. Fuller, P. Hall, and F.-H. Liu · 2022
Cited alongside, same era.
Planting undetectable backdoors in machine learning models
S. Goldwasser, M. P. Kim, V. Vaikuntanathan, and O. Zamir · 2022
Cited alongside, same era.
Handcrafted Backdoors in Deep Neural Networks
S. Hong, N. Carlini, and A. Kurakin · 2022
Cited alongside, same era.
Backdoor learning: A survey, 2022
Y. Li, Y. Jiang, Z. Li, and S.-T. Xia · 2022
Cited alongside, same era.
Piccolo: Exposing complex backdoors in nlp transformer models
Y. Liu, G. Shen, G. Tao, S. An, S. Ma, and X. Zhang · 2022
Cited alongside, same era.
A comprehensive overview of backdoor attacks in large language models within communication networks, 2023
H. Yang, K. Xiang, M. Ge, H. Li, R. Lu, and S. Yu · 2023
Later among the works it cites.
What algorithms can transformers learn? a study in length generalization, 2023
H. Zhou, A. Bradley, E. Littwin, N. Razin, O. Saremi, J. Susskind, S. Bengio, and P. Nakkiran · 2023
Later among the works it cites.
Universal and transferable adversarial attacks on aligned language models, 2023
A. Zou, Z. Wang, N. Carlini, M. Nasr, J. Z. Kolter, and M. Fredrikson · 2023
Later among the works it cites.
Mechanistic interpretability for ai safety – a review, 2024
L. Bereska and E. Gavves · 2024
Closest in time.
Near to mid-term risks and opportunities of open source generative ai, 2024
F. Eiras, A. Petrov, B. Vidgen, C. S. de Witt, F. Pizzati, K. Elkins, S. Mukhopadhyay, A. Bibi, B. Csaba, F. Steibel, F. Barez, G. Smith, G. Guadagni, J. Chun, J. Cabot, J. M. Imperial, J. A. Nolazco-Flores, L. Landay, M. Jackson, P. Röttger, P. H. S. Torr, T. Darrell, Y. S. Lee, and J. Foerster · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Transformerlens
N. Nanda and J. Bloom · 2022
Cited alongside, same era.
Verifying Neural Networks Against Backdoor Attacks
L. H. Pham and J. Sun · 2022
Cited alongside, same era.
Perfectly Secure Steganography Using Minimum Entropy Coupling
C. Schroeder de Witt, S. Sokota, J. Z. Kolter, J. N. Foerster, and M. Strohmeier · 2022
Cited alongside, same era.
Causality-based neural network repair, 2022
B. Sun, J. Sun, H. L. Pham, and J. Shi · 2022
Cited alongside, same era.
Interpretability in the wild: a circuit for indirect object identification in gpt-2 small, 2022
K. Wang, A. Variengien, A. Conmy, B. Shlegeris, and J. Steinhardt · 2022
Cited alongside, same era.
Fine-mixing: Mitigating backdoors in fine-tuned language models
Z. Zhang, L. Lyu, X. Ma, C. Wang, and X. Sun · 2022
Cited alongside, same era.
Teams of llm agents can exploit zero-day vulnerabilities
R. Fang, R. Bindu, A. Gupta, Q. Zhan, and D. Kang · 2024
Closest in time.
Illusory Attacks: Information-Theoretic Detectability Matters in Adversarial Attacks, May 2024
T. Franzmeyer, S. McAleer, J. F. Henriques, J. N. Foerster, P. H. S. Torr, A. Bibi, and C. Schroeder de Witt · 2024
Closest in time.
Implementing an SHA transformer by hand
A. Gritsevskiy · 2024
Closest in time.
Ten-guard: Tensor decomposition for backdoor attack detection in deep neural networks, 2024
K. M. Hossain and T. Oates · 2024
Closest in time.
Composite backdoor attacks against large language models, 2024
H. Huang, Z. Zhao, M. Backes, Y. Shen, and Y. Zhang · 2024
Closest in time.
Sleeper agents: Training deceptive llms that persist through safety training, 2024
E. Hubinger, C. Denison, J. Mu, M. Lambert, M. Tong, M. MacDiarmid, T. Lanham, D. M. Ziegler, T. Maxwell, N. Cheng, A. Jermyn, A. Askell, A. Radhakrishnan, C. Anil, D. Duvenaud, D. Ganguli, F. Barez, J. Clark, K. Ndousse, K. Sachan, M. Sellitto, M. Sharma, N. DasSarma, R. Grosse, S. Kravec, Y. Bai, Z. Witten, M. Favaro, J. Brauner, H. Karnofsky, P. Christiano, S. R. Bowman, L. Graham, J. Kaplan, S. Mindermann, R. Greenblatt, B. Shlegeris, N. Schiefer, and E. Perez · 2024
Closest in time.
Badedit: Backdooring large language models by model editing, 2024
Y. Li, T. Li, K. Chen, J. Zhang, S. Liu, W. Wang, T. Zhang, and Y. Liu · 2024
Closest in time.
In pursuit of superposition: A benchmark for mechanistic interpretability in compressed models
I. A. Moreno · 2024
Closest in time.
Secret Collusion Among Generative AI Agents, Feb. 2024
S. R. Motwani, M. Baranchuk, M. Strohmeier, V. Bolina, P. H. S. Torr, L. Hammond, and C. Schroeder de Witt · 2024
Closest in time.
FIPS PUB 180-4: Secure Hash Standard (SHS)
National Institute of Standards and Technology · 2024
Closest in time.
Competition report: Finding universal jailbreak backdoors in aligned llms, 2024
J. Rando, F. Croce, K. Mitka, S. Shabalin, M. Andriushchenko, N. Flammarion, and F. Tramèr · 2024
Closest in time.
Setting the trap: Capturing and defeating backdoors in pretrained language models through honeypots
R. R. Tang, J. Yuan, Y. Li, Z. Liu, R. Chen, and X. Hu · 2024
Closest in time.
Malicious AI models on Hugging Face backdoor users’ machines, 2024
B. Toulas · 2024
Closest in time.
Mastering symbolic operations: Augmenting language models with compiled neural networks, 2024
Y. Weng, M. Zhu, F. Xia, B. Li, S. He, K. Liu, and J. Zhao · 2024
Closest in time.