Fetching the paper…
Reading the bibliography…
Vulnerability of Frontier language models to misuse and jailbreaks has prompted the development of safety measures like filters and alignment training in an effort to ensure safety through robustness to adversarially crafted prompts.
An improved randomized response strategy
N. S. Mangat · 1994
Earlier work this paper cites.
Differential privacy
C. Dwork · 2006
Earlier work this paper cites.
Composition attacks and auxiliary information in data privacy
S. R. Ganta, S. P. Kasiviswanathan, and A. Smith · 2008
Earlier work this paper cites.
Robust de-anonymization of large sparse datasets
A. Narayanan and V. Shmatikov · 2008
Earlier work this paper cites.
Stealing machine learning models via prediction { \{ APIs } \}
F. Tramèr, F. Zhang, A. Juels, M. K. Reiter, and T. Ristenpart · 2016
Earlier work this paper cites.
Membership inference attacks against machine learning models
R. Shokri, M. Stronati, C. Song, and V. Shmatikov · 2017
Earlier work this paper cites.
Black-box adversarial attacks with limited queries and information
A. Ilyas, L. Engstrom, A. Athalye, and J. Lin · 2018
Earlier work this paper cites.
An overview of information-theoretic security and privacy: Metrics, limits and applications
M. Bloch, O. Günlü, A. Yener, F. Oggier, H. V. Poor, L. Sankar, and R. F. Schaefer · 2021
Earlier work this paper cites.
Decomposed prompting: A modular approach for solving complex tasks
T. Khot, H. Trivedi, M. Finlayson, Y. Fu, K. Richardson, P. Clark, and A. Sabharwal · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al · 2022
Earlier work this paper cites.
Taxonomy of risks posed by language models
L. Weidinger, J. Uesato, M. Rauh, C. Griffin, P.-S. Huang, J. Mellor, A. Glaese, M. Cheng, B. Balle, A. Kasirzadeh, C. Biles, S. Brown, Z. Kenton, W. Hawkins, T. Stepleton, A. Birhane, L. A. Hendricks, L. Rimell, W. Isaac, J. Haas, S. Legassick, G. Irving, and I. Gabriel · 2022
Earlier work this paper cites.
Sparks of artificial general intelligence: Early experiments with gpt-4
S. Bubeck, V. Chandrasekaran, R. Eldan, J. Gehrke, E. Horvitz, E. Kamar, P. Lee, Y. T. Lee, Y. Li, S. Lundberg, et al · 2023
Earlier work this paper cites.
Jailbreaking black box large language models in twenty queries
P. Chao, A. Robey, E. Dobriban, H. Hassani, G. J. Pappas, and E. Wong · 2023
Cited alongside, same era.
Privacy side channels in machine learning systems
E. Debenedetti, G. Severi, N. Carlini, C. A. Choquette-Choo, M. Jagielski, M. Nasr, E. Wallace, and F. Tramèr · 2023
Cited alongside, same era.
LLM censorship: A machine learning challenge or a computer security problem?
D. Glukhov, I. Shumailov, Y. Gal, N. Papernot, and V. Papyan · 2023
Cited alongside, same era.
Llm-blender: Ensembling large language models with pairwise ranking and generative fusion
D. Jiang, X. Ren, and B. Y. Lin · 2023
Cited alongside, same era.
Exploiting programmatic behavior of llms: Dual-use through standard security attacks, 2023
Measuring the persuasiveness of language models, 2024
E. Durmus, L. Lovitt, A. Tamkin, S. Ritchie, J. Clark, and D. Ganguli · 2024
Closest in time.
Red-teaming for generative ai: Silver bullet or security theater?, 2024
M. Feffer, A. Sinha, Z. C. Lipton, and H. Heidari · 2024
Closest in time.
Quantifying privacy via information density, 2024
L. Grosse, S. Saeidian, P. Sadeghi, T. J. Oechtering, and M. Skoglund · 2024
Closest in time.
On the societal impact of open foundation models
S. Kapoor, R. Bommasani, K. Klyman, S. Longpre, A. Ramaswami, P. Cihon, A. Hopkins, K. Bankston, S. Biderman, M. Bogen, et al · 2024
Closest in time.
The llama 3 herd of models, 2024
A. . M. Llama Team · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
D. Kang, X. Li, I. Stoica, C. Guestrin, M. Zaharia, and T. Hashimoto · 2023
Cited alongside, same era.
Jailbreaking chatgpt via prompt engineering: An empirical study
Y. Liu, G. Deng, Z. Xu, Y. Li, Y. Zheng, Y. Zhang, L. Zhao, T. Zhang, K. Wang, and Y. Liu · 2023
Cited alongside, same era.
Pufferfish privacy: An information-theoretic study, 2023
T. Nuradha and Z. Goldfeld · 2023
Cited alongside, same era.
Question decomposition improves the faithfulness of model-generated reasoning
A. Radhakrishnan, K. Nguyen, A. Chen, C. Chen, C. Denison, D. Hernandez, E. Durmus, E. Hubinger, J. Kernion, K. Lukošiūtė, et al · 2023
Cited alongside, same era.
Polynomial time cryptanalytic extraction of neural network models
A. Shamir, I. Canales-Martinez, A. Hambitzer, J. Chavez-Saab, F. Rodrigez-Henriquez, and N. Satpute · 2023
Cited alongside, same era.
On the privacy-utility trade-off with and without direct access to the private data
A. Zamani, T. J. Oechtering, and M. Skoglund · 2023
Cited alongside, same era.
https://www.anthropic.com/news/claude-3-5-sonnet , 2024
Introducing Claude 3.5 Sonnet — anthropic.com · 2024
Cited alongside, same era.
Many-shot jailbreaking
C. Anil, E. Durmus, M. Sharma, J. Benton, S. Kundu, J. Batson, N. Rimsky, M. Tong, J. Mu, D. Ford, et al · 2024
Cited alongside, same era.
M. Phuong, M. Aitchison, E. Catt, S. Cogan, A. Kaskasoli, V. Krakovna, D. Lindner, M. Rahtz, Y. Assael, S. Hodkinson, et al · 2024
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
R. Rafailov, A. Sharma, E. Mitchell, C. D. Manning, S. Ermon, and C. Finn · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
M. Reid, N. Savinov, D. Teplyashin, D. Lepikhin, T. Lillicrap, J.-b. Alayrac, R. Soricut, A. Lazaridou, O. Firat, J. Schrittwieser, et al · 2024
Closest in time.
Great, now write an article about that: The crescendo multi-turn llm jailbreak attack, 2024
M. Russinovich, A. Salem, and R. Eldan · 2024
Closest in time.
P. Slattery, A. K. Saeri, E. A. Grundy, J. Graham, M. Noetel, R. Uuk, J. Dao, S. Pour, S. Casper, and N. Thompson · 2024
Closest in time.
A strongreject for empty jailbreaks, 2024
A. Souly, Q. Lu, D. Bowen, T. Trinh, E. Hsieh, S. Pandey, P. Abbeel, J. Svegliato, S. Emmons, O. Watkins, and S. Toyer · 2024
Closest in time.
Improving alignment and robustness with circuit breakers, 2024
A. Zou, L. Phan, J. Wang, D. Duenas, M. Lin, M. Andriushchenko, R. Wang, Z. Kolter, M. Fredrikson, and D. Hendrycks · 2024
Closest in time.