Fetching the paper…
Reading the bibliography…
The exposure of security vulnerabilities in safety-aligned language models, e.g., susceptibility to adversarial attacks, has shed light on the intricate interplay between AI safety and AI security.
Membership inference attacks from first principles
Carlini, N., Chien, S., Nasr, M., Song, S., Terzis, A., and Tramer, F · 1914
Earlier work this paper cites.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Bradley, R. A. and Terry, M. E · 1952
Earlier work this paper cites.
How to generate and exchange secrets
Yao, A. C.-C · 1986
Earlier work this paper cites.
On the meaning of safety and security
Burns, A., McDermid, J., and Dobson, J · 1992
Earlier work this paper cites.
Health Insurance Portability and Accountability Act of 1996
European Parliament and Council of the European Union · 1997
Earlier work this paper cites.
Secure multi-party computation
Goldreich, O · 1998
Earlier work this paper cites.
Studying organizational justice cross-culturally: fundamental challenges
Greenberg, J · 2001
Earlier work this paper cites.
Data privacy through optimal k-anonymization
Bayardo, R. J. and Agrawal, R · 2005
Earlier work this paper cites.
Exploiting machine learning to subvert your spam filter
Nelson, B., Barreno, M., Chi, F. J., Joseph, A. D., Rubinstein, B. I., Saini, U., Sutton, C., Tygar, J. D., and Xia, K · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
The sema referential framework: Avoiding ambiguities in the terms “security” and “safety”
Piètre-Cambacédès, L. and Chaudet, C · 2010
Earlier work this paper cites.
Understanding privacy
Solove, D. J · 2010
Earlier work this paper cites.
Poisoning attacks against support vector machines
Biggio, B., Nelson, B., and Laskov, P · 2012
Earlier work this paper cites.
The mnist database of handwritten digit images for machine learning research
Deng, L · 2012
Earlier work this paper cites.
Evasion attacks against machine learning at test time
Biggio, B., Corona, I., Maiorca, D., Nelson, B., Šrndić, N., Laskov, P., Giacinto, G., and Roli, F · 2013
Earlier work this paper cites.
The algorithmic foundations of differential privacy
Dwork, C., Roth, A., et al · 2014
Earlier work this paper cites.
Explaining and harnessing adversarial examples
Goodfellow, I. J., Shlens, J., and Szegedy, C · 2014
Earlier work this paper cites.
Intriguing properties of neural networks
Szegedy, C., Zaremba, W., Sutskever, I., Bruna, J., Erhan, D., Goodfellow, I., and Fergus, R · 2014
Earlier work this paper cites.
Hacking smart machines with smarter ones: How to extract meaningful data from machine learning classifiers
Ateniese, G., Mancini, L. V., Spognardi, A., Villani, A., Vitali, D., and Felici, G · 2015
Earlier work this paper cites.
Concrete problems in ai safety
Amodei, D., Olah, C., Steinhardt, J., Christiano, P., Schulman, J., and Mané, D · 2016
Earlier work this paper cites.
Defensive distillation is not robust to adversarial examples
Carlini, N. and Wagner, D · 2016
Earlier work this paper cites.
Regulation (EU) 2016/679 of the European Parliament and of the Council
European Parliament and Council of the European Union · 2016
Earlier work this paper cites.
Inherent trade-offs in the fair determination of risk scores
Kleinberg, J., Mullainathan, S., and Raghavan, M · 2016
Earlier work this paper cites.
Learning from tay’s introduction
Lee, P · 2016
Earlier work this paper cites.
Safely interruptible agents
Orseau, L. and Armstrong, M · 2016
Earlier work this paper cites.
Membership inference attacks against machine learning models
Shokri, R., Stronati, M., Song, C., and Shmatikov, V · 2016
Earlier work this paper cites.
Stealing machine learning models via prediction { \{ APIs } \}
Tramèr, F., Zhang, F., Juels, A., Reiter, M. K., and Ristenpart, T · 2016
Earlier work this paper cites.
Artificial intelligence safety and cybersecurity: A timeline of ai failures
Yampolskiy, R. V. and Spellchecker, M · 2016
Earlier work this paper cites.
Adversarial examples are not easily detected: Bypassing ten detection methods
Carlini, N. and Wagner, D · 2017
Earlier work this paper cites.
Targeted backdoor attacks on deep learning systems using data poisoning
Chen, X., Liu, C., Li, B., Lu, K., and Song, D · 2017
Earlier work this paper cites.
Predictably unequal? the effects of machine learning on credit markets
Fuster, A., Goldsmith-Pinkham, P., Ramadorai, T., and Walther, A · 2017
Earlier work this paper cites.
Badnets: Identifying vulnerabilities in the machine learning model supply chain
Gu, T., Dolan-Gavitt, B., and Garg, S · 2017
Earlier work this paper cites.
On calibration of modern neural networks
Guo, C., Pleiss, G., Sun, Y., and Weinberger, K. Q · 2017
Earlier work this paper cites.
Adversarial example defense: Ensembles of weak defenses are not strong
He, W., Wei, J., Chen, X., Carlini, N., and Song, D · 2017
Earlier work this paper cites.
Fault injection attack on deep neural network
Liu, Y., Wei, L., Luo, B., and Xu, Q · 2017
Earlier work this paper cites.
Towards deep learning models resistant to adversarial attacks
Madry, A., Makelov, A., Schmidt, L., Tsipras, D., and Vladu, A · 2017
Earlier work this paper cites.
Milli, S., Hadfield-Menell, D., Dragan, A., and Russell, S · 2017
Earlier work this paper cites.
Towards poisoning of deep learning algorithms with back-gradient optimization
Muñoz-González, L., Biggio, B., Demontis, A., Paudice, A., Wongrassamee, V., Lupu, E. C., and Roli, F · 2017
Earlier work this paper cites.
An introduction to information security
Nieles, M., Dempsey, K., and Pillitteri, V. Y · 2017
Earlier work this paper cites.
Practical black-box attacks against machine learning
Papernot, N., McDaniel, P., Goodfellow, I., Jha, S., Celik, Z. B., and Swami, A · 2017
Earlier work this paper cites.
A game-theoretic analysis of label flipping attacks on distributed support vector machines
Rui Zhang and Quanyan Zhu · 2017
Earlier work this paper cites.
Membership inference attacks against machine learning models
Shokri, R., Stronati, M., Song, C., and Shmatikov, V · 2017
Earlier work this paper cites.
Privacy risk in machine learning: Analyzing the connection to overfitting
Yeom, S., Giacomelli, I., Fredrikson, M., and Jha, S · 2017
Earlier work this paper cites.
Obfuscated gradients give a false sense of security: Circumventing defenses to adversarial examples
Athalye, A., Carlini, N., and Wagner, D · 2018
Earlier work this paper cites.
Practical fault attack on deep neural networks
Breier, J., Hou, X., Jap, D., Ma, L., Bhasin, S., and Liu, Y · 2018
Earlier work this paper cites.
Stealing neural networks via timing side channels
Duddu, V., Samanta, D., Rao, D. V., and Balas, V. E · 2018
Earlier work this paper cites.
Robust physical-world attacks on deep learning visual classification
Eykholt, K., Evtimov, I., Fernandes, E., Li, B., Rahmati, A., Xiao, C., Prakash, A., Kohno, T., and Song, D · 2018
Earlier work this paper cites.
Property inference attacks on fully connected neural networks using permutation invariant representations
Ganju, K., Wang, Q., Yang, W., Gunter, C. A., and Borisov, N · 2018
Earlier work this paper cites.
Human perceptions of fairness in algorithmic decision making: A case study of criminal risk prediction
Grgic-Hlaca, N., Redmiles, E. M., Gummadi, K. P., and Weller, A · 2018
Earlier work this paper cites.
Risk of re-identification from payment card histories in multiple domains
Ito, S., Harada, R., and Kikuchi, H · 2018
Earlier work this paper cites.
Scalable agent alignment via reward modeling: a research direction, 2018
Leike, J., Krueger, D., Everitt, T., Martic, M., Maini, V., and Legg, S · 2018
Earlier work this paper cites.
Trojaning attack on neural networks
Liu, Y., Ma, S., Aafer, Y., Lee, W.-C., Zhai, J., Wang, W., and Zhang, X · 2018
Earlier work this paper cites.
Framework for Improving Critical Infrastructure Cybersecurity
NIST · 2018
Earlier work this paper cites.
Poison frogs! targeted clean-label poisoning attacks on neural networks
Shafahi, A., Huang, W. R., Najibi, M., Suciu, O., Studer, C., Dumitras, T., and Goldstein, T · 2018
Earlier work this paper cites.
California consumer privacy act (ccpa)
State of California Legislative Counsel · 2018
Earlier work this paper cites.
Robustness may be at odds with accuracy
Tsipras, D., Santurkar, S., Engstrom, L., Turner, A., and Madry, A · 2018
Earlier work this paper cites.
Stealing hyperparameters in machine learning
Wang, B. and Gong, N. Z · 2018
Earlier work this paper cites.
Rules and Policies - Protecting PII - Privacy Act
Administration, U. G. S · 2019
Earlier work this paper cites.
Scibert: A pretrained language model for scientific text
Beltagy, I., Lo, K., and Cohan, A · 2019
Earlier work this paper cites.
Adversarial sensor attack on lidar-based perception in autonomous driving
Cao, Y., Xiao, C., Cyr, B., Zhou, Y., Park, W., Rampazzi, S., Chen, Q. A., Fu, K., and Mao, Z. M · 2019
Earlier work this paper cites.
The secret sharer: Evaluating and testing unintended memorization in neural networks
Carlini, N., Liu, C., Erlingsson, Ú., Kos, J., and Song, D · 2019
Earlier work this paper cites.
Certified adversarial robustness via randomized smoothing
Cohen, J., Rosenfeld, E., and Kolter, Z · 2019
Earlier work this paper cites.
Adversarial reprogramming of neural networks
Elsayed, G. F., Goodfellow, I., and Sohl-Dickstein, J · 2019
Earlier work this paper cites.
Is tricking a robot hacking?
Evtimov, I., O’Hair, D., Fernandes, E., Calo, R., and Kohno, T · 2019
Earlier work this paper cites.
Benchmarking neural network robustness to common corruptions and perturbations
Hendrycks, D. and Dietterich, T · 2019
Earlier work this paper cites.
Clinicalbert: Modeling clinical notes and predicting hospital readmission
Huang, K., Altosaar, J., and Ranganath, R · 2019
Earlier work this paper cites.
Prada: protecting against dnn model stealing attacks
Juuti, M., Szyller, S., Marchal, S., and Asokan, N · 2019
Earlier work this paper cites.
Enabling Pedestrian Safety Using Computer Vision Techniques: A Case Study of the 2018 Uber Inc. Self-driving Car Crash , pp. 261–279
Kohli, P. and Chadha, A · 2019
Earlier work this paper cites.
Thieves on sesame street! model extraction of bert-based apis
Krishna, K., Tomar, G. S., Parikh, A. P., Papernot, N., and Iyyer, M · 2019
Earlier work this paper cites.
Towards reverse-engineering black-box neural networks
Oh, S. J., Schiele, B., and Fritz, M · 2019
Earlier work this paper cites.
Knockoff nets: Stealing functionality of black-box models
Orekondy, T., Schiele, B., and Fritz, M · 2019
Earlier work this paper cites.
Bit-flip attack: Crushing neural network with progressive bit search
Rakin, A. S., He, Z., and Fan, D · 2019
Earlier work this paper cites.
Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead
Rudin, C · 2019
Earlier work this paper cites.
Explainable ai: A brief survey on history, research areas, approaches and challenges
Xu, F., Uszkoreit, H., Du, Y., Fan, W., Zhao, D., and Zhu, J · 2019
Earlier work this paper cites.
Fault sneaking attack: A stealthy framework for misleading deep neural networks
Zhao, P., Wang, S., Gongye, C., Wang, Y., Fei, Y., and Lin, X · 2019
Earlier work this paper cites.
Model extraction from counterfactual explanations
Aïvodji, U., Bolot, A., and Gambs, S · 2020
Earlier work this paper cites.
Explainability for artificial intelligence in healthcare: a multidisciplinary perspective
Amann, J., Blasimme, A., Vayena, E., Frey, D., Madai, V. I., and Consortium, P · 2020
Earlier work this paper cites.
How to backdoor federated learning
Bagdasaryan, E., Veit, A., Hua, Y., Estrin, D., and Shmatikov, V · 2020
Earlier work this paper cites.
Black-box ripper: Copying black-box models using generative evolutionary algorithms
Barbalau, A., Cosma, A., Ionescu, R. T., and Popescu, M · 2020
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Cryptanalytic extraction of neural network models
Carlini, N., Jagielski, M., and Mironov, I · 2020
Earlier work this paper cites.
On adversarial bias and the robustness of fair machine learning
Chang, H., Nguyen, T. D., Murakonda, S. K., Kazemi, E., and Shokri, R · 2020
Earlier work this paper cites.
Robustbench: a standardized adversarial robustness benchmark
Croce, F., Andriushchenko, M., Sehwag, V., Debenedetti, E., Flammarion, N., Chiang, M., Mittal, P., and Hein, M · 2020
Earlier work this paper cites.
Artificial intelligence, values, and alignment
Gabriel, I · 2020
Earlier work this paper cites.
Realtoxicityprompts: Evaluating neural toxic degeneration in language models
Gehman, S., Gururangan, S., Sap, M., Choi, Y., and Smith, N. A · 2020
Earlier work this paper cites.
Model extraction attacks and defenses on cloud-based machine learning models
Gong, X., Wang, Q., Chen, Y., Yang, W., and Jiang, X · 2020
Cited alongside, same era.
Avguardian: Detecting and mitigating publish-subscribe overprivilege for autonomous vehicle systems
Hong, D. K., Kloosterman, J., Jin, Y., Cao, Y., Chen, Q. A., Mahlke, S., and Mao, Z. M · 2020
Cited alongside, same era.
Metapoison: Practical general-purpose clean-label data poisoning
Huang, W. R., Geiping, J., Fowl, L., Taylor, G., and Goldstein, T · 2020
Cited alongside, same era.
A context-aware citation recommendation model with bert and graph convolutional networks
Jeong, C., Jang, S., Park, E., and Choi, S · 2020
Cited alongside, same era.
The NIST Privacy Framework: A Tool for Improving Privacy through Enterprise Risk Management
NIST · 2020
Cited alongside, same era.
Are diffusion models vulnerable to membership inference attacks?
Duan, J., Kong, F., Wang, S., Shi, X., and Xu, K · 2023
Later among the works it cites.
Helping or herding? reward model ensembles mitigate but do not eliminate reward hacking
Eisenstein, J., Nagpal, C., Agarwal, A., Beirami, A., D’Amour, A., Dvijotham, D., Fisch, A., Heller, K., Pfohl, S., Ramachandran, D., et al · 2023
Later among the works it cites.
Eu artificial intelligence act
European Commission · 2023
Later among the works it cites.
Gemini: A family of highly capable multimodal models
Gemini Team · 2023
Later among the works it cites.
Sparsity-preserving differentially private training of large embedding models
Ghazi, B., Huang, Y., Kamath, P., Kumar, R., Manurangsi, P., Sinha, A., and Zhang, C · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Pal, S., Gupta, Y., Shukla, A., Kanade, A., Shevade, S., and Ganapathy, V · 2020
Cited alongside, same era.
Tbt: Targeted neural network attack with bit trojan
Rakin, A. S., He, Z., and Fan, D · 2020
Cited alongside, same era.
Negligence and ai’s human users
Selbst, A. D · 2020
Cited alongside, same era.
Poisoning attacks on algorithmic fairness
Solans, D., Biggio, B., and Castillo, C · 2020
Cited alongside, same era.
Model extraction attacks on recurrent neural networks
Takemura, T., Yanai, N., and Fujiwara, T · 2020
Cited alongside, same era.
Statistical consequences of fat tails: Real world preasymptotics, epistemology, and applications
Taleb, N. N · 2020
Cited alongside, same era.
An embarrassingly simple approach for trojan attack in deep neural networks
Tang, R., Du, M., Liu, N., Yang, F., and Hu, X · 2020
Cited alongside, same era.
Later among the works it cites.
A survey on the possibilities & impossibilities of AI-generated text detection
Ghosal, S. S., Chakraborty, S., Geiping, J., Huang, F., Manocha, D., and Bedi, A · 2023
Later among the works it cites.
Ai safety summit 2023
GOV.UK · 2023
Later among the works it cites.
An overview of catastrophic ai risks, 2023
Hendrycks, D., Mazeika, M., and Woodside, T · 2023
Later among the works it cites.
Llama guard: Llm-based input-output safeguard for human-ai conversations
Inan, H., Upasani, K., Chi, J., Rungta, R., Iyer, K., Mao, Y., Tontchev, M., Hu, Q., Fuller, B., Testuggine, D., et al · 2023
Later among the works it cites.
A systems view of nuclear security and nuclear safety: Identifying interfaces and building synergies
International Atomic Energy Agency · 2023
Later among the works it cites.
Trick ChatGPT to say its SECRET PROMPT
Investor, T. P · 2023
Later among the works it cites.
Llm platform security: Applying a systematic evaluation framework to openai’s chatgpt plugins
Iqbal, U., Kohno, T., and Roesner, F · 2023
Later among the works it cites.
Gpt3.5 just adding a random dude’s photo in the reply
jutogashi · 2023
Later among the works it cites.
Toward comprehensive risk assessments and assurance of ai-based systems
Khlaaf, H · 2023
Later among the works it cites.
Propile: Probing privacy leakage in large language models
Kim, S., Yun, S., Lee, H., Gubri, M., Yoon, S., and Oh, S. J · 2023
Later among the works it cites.
A watermark for large language models
Kirchenbauer, J., Geiping, J., Wen, Y., Katz, J., Miers, I., and Goldstein, T · 2023
Later among the works it cites.
Pretraining language models with human preferences
Korbak, T., Shi, K., Chen, A., Bhalerao, R. V., Buckley, C., Phang, J., Bowman, S. R., and Perez, E · 2023
Later among the works it cites.
Paraphrasing evades detectors of AI-generated text, but retrieval is an effective defense
Krishna, K., Song, Y., Karpinska, M., Wieting, J. F., and Iyyer, M · 2023
Later among the works it cites.
Robust distortion-free watermarks for language models
Kuditipudi, R., Thickstun, J., Hashimoto, T., and Liang, P · 2023
Later among the works it cites.
Framing the fallibility of computer-aided detection aids cancer detection
Kunar, M. A. and Watson, D. G · 2023
Later among the works it cites.
Ai safety on whose terms?, 2023
Lazar, S. and Nelson, A · 2023
Later among the works it cites.
Introducing Superalignment
Leike, J. and Sutskever, I · 2023
Later among the works it cites.
Lora fine-tuning efficiently undoes safety training in llama 2-chat 70b
Lermen, S., Rogers-Smith, C., and Ladish, J · 2023
Later among the works it cites.
Quda: query-limited data-free model extraction
Lin, Z., Xu, K., Fang, C., Zheng, H., Ahmed Jaheezuddin, A., and Shi, J · 2023
Later among the works it cites.
Exploring the limits of model-targeted indiscriminate data poisoning attacks
Lu, Y., Kamath, G., and Yu, Y · 2023
Later among the works it cites.
An empirical study of catastrophic forgetting in large language models during continual fine-tuning
Luo, Y., Yang, Z., Meng, F., Li, Y., Zhou, J., and Zhang, Y · 2023
Later among the works it cites.
McIntosh, T. R., Susnjak, T., Liu, T., Watters, P., and Halgamuge, M. N · 2023
Later among the works it cites.
URL https://ai.meta.com/llama/use-policy/
Meta, 2023 · 2023
Later among the works it cites.
Silo language models: Isolating legal risk in a nonparametric datastore
Min, S., Gururangan, S., Wallace, E., Hajishirzi, H., Smith, N. A., and Zettlemoyer, L · 2023
Later among the works it cites.
Nash learning from human feedback
Munos, R., Valko, M., Calandriello, D., Azar, M. G., Rowland, M., Guo, Z. D., Tang, Y., Geist, M., Mesnard, T., Michi, A., et al · 2023
Later among the works it cites.
Model alignment protects against accidental harms, not intentional ones
Narayanan, A. and Kapoor, S · 2023
Later among the works it cites.
AI Risk Management Framework
NIST · 2023
Later among the works it cites.
Asset: Robust backdoor data detection across a multiplicity of deep learning paradigms
Pan, M., Zeng, Y., Lyu, L., Lin, X., and Jia, R · 2023
Later among the works it cites.
Pedro, R., Castro, D., Carreira, P., and Santos, N · 2023
Later among the works it cites.
A survey of bit-flip attacks on deep neural network and corresponding defense methods
Qian, C., Zhang, M., Nie, Y., Lu, S., and Cao, H · 2023
Later among the works it cites.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C. D., and Finn, C · 2023
Later among the works it cites.
Concrete problems in ai safety, revisited
Raji, I. D. and Dobbe, R · 2023
Later among the works it cites.
Universal jailbreak backdoors from poisoned human feedback
Rando, J. and Tramèr, F · 2023
Later among the works it cites.
Doubling down on dumb: Lessons from mata v. avianca inc
Rapoport, N. B., Norton, H., and Cynthia, A · 2023
Later among the works it cites.
Two us lawyers fined for submitting fake court citations from chatgpt
Reed, B · 2023
Later among the works it cites.
Smoothllm: Defending large language models against jailbreaking attacks
Robey, A., Wong, E., Hassani, H., and Pappas, G. J · 2023
Later among the works it cites.
Can ai-generated text be reliably detected?
Sadasivan, V. S., Kumar, A., Balasubramanian, S., Wang, W., and Feizi, S · 2023
Later among the works it cites.
Copyright safety for generative ai
Sag, M · 2023
Later among the works it cites.
Practices for governing agentic ai systems, 2023
Shavit, Y., Agarwal, S., Brundage, M., Adler, S., O’Keefe, C., Campbell, R., Lee, T., Mishkin, P., Eloundou, T., Hickey, A., et al · 2023
Later among the works it cites.
Detecting pretraining data from large language models
Shi, W., Ajith, A., Xia, M., Huang, Y., Liu, D., Blevins, T., Chen, D., and Zettlemoyer, L · 2023
Later among the works it cites.
On the exploitability of instruction tuning
Shu, M., Wang, J., Zhu, C., Geiping, J., Xiao, C., and Goldstein, T · 2023
Later among the works it cites.
Large language models encode clinical knowledge
Singhal, K., Azizi, S., Tu, T., Mahdavi, S. S., Wei, J., Chung, H. W., Scales, N., Tanwani, A., Cole-Lewis, H., Pfohl, S., et al · 2023
Later among the works it cites.
Privacy auditing with one (1) training run
Steinke, T., Nasr, M., and Jagielski, M · 2023
Later among the works it cites.
AI Security Has Serious Terminology Issues
Thacker, J · 2023
Later among the works it cites.
Large language models in medicine
Thirunavukarasu, A. J., Ting, D. S. J., Elangovan, K., Gutierrez, L., Tan, T. F., and Ting, D. S. W · 2023
Later among the works it cites.
Explanations can reduce overreliance on ai systems during decision-making
Vasconcelos, H., Jörke, M., Grunde-McLaughlin, M., Gerstenberg, T., Bernstein, M. S., and Krishna, R · 2023
Later among the works it cites.
Poisoning language models during instruction tuning
Wan, A., Wallace, E., Shen, S., and Klein, D · 2023
Later among the works it cites.
These are Microsoft’s Bing AI secret rules and why it says it’s named Sydney
Warren, T · 2023
Later among the works it cites.
Jailbroken: How does llm safety training fail?
Wei, A., Haghtalab, N., and Steinhardt, J · 2023
Later among the works it cites.
Learning to invert: Simple adaptive attacks for gradient inversion in federated learning
Wu, R., Chen, X., Guo, C., and Weinberger, K. Q · 2023
Later among the works it cites.
Data selection for language models via importance resampling
Xie, S. M., Santurkar, S., Ma, T., and Liang, P · 2023
Later among the works it cites.
Backdooring instruction-tuned large language models with virtual prompt injection
Yan, J., Yadav, V., Li, S., Chen, L., Tang, Z., Wang, H., Srinivasan, V., Ren, X., and Jin, H · 2023
Later among the works it cites.
Shadow alignment: The ease of subverting safely-aligned language models, 2023
Yang, X., Wang, X., Zhang, Q., Petzold, L., Wang, W. Y., Zhao, X., and Lin, D · 2023
Later among the works it cites.
A survey on large language model (llm) security and privacy: The good, the bad, and the ugly
Yao, Y., Duan, J., Xu, K., Cai, Y., Sun, E., and Zhang, Y · 2023
Later among the works it cites.
Benchmarking and defending against indirect prompt injection attacks on large language models
Yi, J., Xie, Y., Zhu, B., Hines, K., Kiciman, E., Sun, G., Xie, X., and Wu, F · 2023
Later among the works it cites.
Low-resource languages jailbreak gpt-4
Yong, Z.-X., Menghini, C., and Bach, S. H · 2023
Later among the works it cites.
Investigating the catastrophic forgetting in multimodal large language models
Zhai, Y., Tong, S., Li, X., Cai, M., Qu, Q., Lee, Y. J., and Ma, Y · 2023
Later among the works it cites.
Removing rlhf protections in gpt-4 via fine-tuning
Zhan, Q., Fang, R., Bindu, R., Gupta, A., Hashimoto, T., and Kang, D · 2023
Later among the works it cites.
Watermarks in the sand: Impossibility of strong watermarking for generative models
Zhang, H., Edelman, B. L., Francati, D., Venturi, D., Ateniese, G., and Barak, B · 2023
Later among the works it cites.
Prompts should not be seen as secrets: Systematically measuring prompt extraction attack success
Zhang, Y. and Ippolito, D · 2023
Later among the works it cites.
Breaking llama guard
Zou, A · 2023
Later among the works it cites.
The risks of expanding the definition of ’ai safety’
Albergotti, R · 2024
Closest in time.
Planning for agi and beyond
Altman, O · 2024
Closest in time.
Advancing Differential Privacy: Where We Are Now and Future Directions for Real-World Deployment
Cummings, R., Desfontaines, D., Evans, D., Geambasu, R., Huang, Y., Jagielski, M., Kairouz, P., Kamath, G., Oh, S., Ohrimenko, O., Papernot, N., Rogers, R., Shen, M., Song, S., Su, W., Terzis, A., Thakurta, A., Vassilvitskii, S., Wang, Y.-X., Xiong, L., Yekhanin, S., Yu, D., Zhang, H., and Zhang, W · 2024
Closest in time.
Keep Your Enemies Safer: Technical Cooperation and Transferring Nuclear Safety and Security Technologies
Ding, J · 2024
Closest in time.
Towards more realistic membership inference attacks on large diffusion models
Dubiński, J., Kowalczuk, A., Pawlak, S., Rokita, P., Trzciński, T., and Morawiecki, P · 2024
Closest in time.
Challenges of artificial intelligence in medicine and dermatology
Grzybowski, A., Jin, K., and Wu, H · 2024
Closest in time.
Sleeper agents: Training deceptive llms that persist through safety training, 2024
Hubinger, E., Denison, C., Mu, J., Lambert, M., Tong, M., MacDiarmid, M., Lanham, T., Ziegler, D. M., Maxwell, T., Cheng, N., Jermyn, A., Askell, A., Radhakrishnan, A., Anil, C., Duvenaud, D., Ganguli, D., Barez, F., Clark, J., Ndousse, K., Sachan, K., Sellitto, M., Sharma, M., DasSarma, N., Grosse, R., Kravec, S., Bai, Y., Witten, Z., Favaro, M., Brauner, J., Karnofsky, H., Christiano, P., Bowman, S. R., Graham, L., Kaplan, J., Mindermann, S., Greenblatt, R., Shlegeris, B., Schiefer, N., and Perez, E · 2024
Closest in time.
Gpts prompts leaked list
Leaked-GPTs · 2024
Closest in time.
A mechanistic understanding of alignment algorithms: A case study on dpo and toxicity
Lee, A., Bai, X., Pres, I., Wattenberg, M., Kummerfeld, J. K., and Mihalcea, R · 2024
Closest in time.
Yes, one-bit-flip matters! universal dnn model inference depletion with runtime code fault injection
Li, S., Wang, X., Xue, M., Zhu, H., Zhang, Z., Gao, Y., Wu, W., and Shen, X. S · 2024
Closest in time.
Harmbench: A standardized evaluation framework for automated red teaming and robust refusal
Mazeika, M., Phan, L., Yin, X., Zou, A., Wang, Z., Mu, N., Sakhaee, E., Li, N., Basart, S., Li, B., et al · 2024
Closest in time.
Terms of use
OpenAI · 2024
Closest in time.
Aws fixes data exfiltration attack angle in amazon q for business
Rehberger, J · 2024
Closest in time.
Escalation risks from language models in military and diplomatic decision-making
Rivera, J.-P., Mukobi, G., Reuel, A., Lamparth, M., Smith, C., and Schneider, J · 2024
Closest in time.
Private fine-tuning of large language models with zeroth-order optimization
Tang, X., Panda, A., Nasr, M., Mahloujifar, S., and Mittal, P · 2024
Closest in time.
Assessing the brittleness of safety alignment via pruning and low-rank modifications
Wei, B., Huang, K., Huang, Y., Xie, T., Qi, X., Xia, M., Mittal, P., Wang, M., and Henderson, P · 2024
Closest in time.
Badchain: Backdoor chain-of-thought prompting for large language models
Xiang, Z., Jiang, F., Xiong, Z., Ramasubramanian, B., Poovendran, R., and Li, B · 2024
Closest in time.
Zeng, Y., Lin, H., Zhang, J., Yang, D., Jia, R., and Shi, W · 2024
Closest in time.
Robust prompt optimization for defending language models against jailbreaking attacks
Zhou, A., Li, B., and Wang, H · 2024
Closest in time.