Fetching the paper…
Reading the bibliography…
AI systems can take harmful actions and are highly vulnerable to adversarial attacks.
Engineering a safer world: Systems thinking applied to safety
N. Leveson · 2012
Earlier work this paper cites.
Efficient estimation of word representations in vector space
T. Mikolov, K. Chen, G. Corrado, and J. Dean · 2013
Earlier work this paper cites.
Microsoft coco: Common objects in context, 2015
T.-Y. Lin, M. Maire, S. Belongie, L. Bourdev, R. Girshick, J. Hays, P. Perona, D. Ramanan, C. L. Zitnick, and P. Dollár · 2015
Earlier work this paper cites.
Deep reinforcement learning from human preferences
P. F. Christiano, J. Leike, T. Brown, M. Martic, S. Legg, and D. Amodei · 2017
Earlier work this paper cites.
Towards deep learning models resistant to adversarial attacks
A. Madry, A. Makelov, L. Schmidt, D. Tsipras, and A. Vladu · 2017
Earlier work this paper cites.
Deep feature interpolation for image content changes
P. Upchurch, J. Gardner, G. Pleiss, R. Pless, N. Snavely, K. Bala, and K. Weinberger · 2017
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge, 2018
P. Clark, I. Cowhey, O. Etzioni, T. Khot, A. Sabharwal, C. Schoenick, and O. Tafjord · 2018
Earlier work this paper cites.
WINOGRANDE: an adversarial winograd schema challenge at scale, 2019
K. Sakaguchi, R. L. Bras, C. Bhagavatula, and Y. Choi · 2019
Earlier work this paper cites.
Robustness may be at odds with accuracy, 2019
D. Tsipras, S. Santurkar, L. Engstrom, A. Turner, and A. Madry · 2019
Earlier work this paper cites.
Hellaswag: Can a machine really finish your sentence?, 2019
R. Zellers, A. Holtzman, Y. Bisk, A. Farhadi, and Y. Choi · 2019
Earlier work this paper cites.
Semantic photo manipulation with a generative image prior
D. Bau, H. Strobelt, W. Peebles, J. Wulff, B. Zhou, J.-Y. Zhu, and A. Torralba · 2020
Earlier work this paper cites.
Measuring massive multitask language understanding
D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt · 2020
Earlier work this paper cites.
Emerging properties in self-supervised vision transformers
M. Caron, H. Touvron, I. Misra, H. Jégou, J. Mairal, P. Bojanowski, and A. Joulin · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems, 2021
K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, C. Hesse, and J. Schulman · 2021
Earlier work this paper cites.
Multimodal neurons in artificial neural networks
G. Goh, N. Cammarata, C. Voss, S. Carter, M. Petrov, L. Schubert, A. Radford, and C. Olah · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen · 2021
Earlier work this paper cites.
Editgan: High-precision semantic image editing
H. Ling, K. Kreis, D. Li, S. W. Kim, A. Torralba, and S. Fidler · 2021
Earlier work this paper cites.
E. Mitchell, C. Lin, A. Bosselut, C. Finn, and C. D. Manning · 2021
Earlier work this paper cites.
Editing models with task arithmetic
G. Ilharco, M. T. Ribeiro, M. Wortsman, S. Gururangan, L. Schmidt, H. Hajishirzi, and A. Farhadi · 2022
Earlier work this paper cites.
Truthfulqa: Measuring how models mimic human falsehoods, 2022
S. Lin, J. Hilton, and O. Evans · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al · 2022
Earlier work this paper cites.
Red teaming language models with language models
E. Perez, S. Huang, F. Song, T. Cai, R. Ring, J. Aslanides, A. Glaese, N. McAleese, and G. Irving · 2022
Earlier work this paper cites.
Detecting language model attacks with perplexity
G. Alon and M. Kamfonas · 2023
Earlier work this paper cites.
Image hijacks: Adversarial images can control generative models at runtime
L. Bailey, E. Ong, S. Russell, and S. Emmons · 2023
Cited alongside, same era.
Open llm leaderboard
E. Beeching, C. Fourrier, N. Habib, S. Han, N. Lambert, N. Rajani, O. Sanseviero, L. Tunstall, and T. Wolf · 2023
Cited alongside, same era.
Are aligned neural networks adversarially aligned?
N. Carlini, M. Nasr, C. A. Choquette-Choo, M. Jagielski, I. Gao, P. W. W. Koh, D. Ippolito, F. Tramer, and L. Schmidt · 2023
Cited alongside, same era.
Jailbreaking black box large language models in twenty queries, 2023
P. Chao, A. Robey, E. Dobriban, H. Hassani, G. J. Pappas, and E. Wong · 2023
Cited alongside, same era.
Enhancing chat language models by scaling high-quality instructional conversations
N. Ding, Y. Chen, B. Xu, Y. Qin, Z. Zheng, S. Hu, Z. Liu, M. Sun, and B. Zhou · 2023
Cited alongside, same era.
Judging llm-as-a-judge with mt-bench and chatbot arena
L. Zheng, W.-L. Chiang, Y. Sheng, S. Zhuang, Z. Wu, Y. Zhuang, Z. Lin, Z. Li, D. Li, E. Xing, et al · 2023
Later among the works it cites.
Jailbreaking leading safety-aligned LLMs with simple adaptive attacks
M. Andriushchenko, F. Croce, and N. Flammarion · 2024
Closest in time.
Red-teaming for generative ai: Silver bullet or security theater?
M. Feffer, A. Sinha, Z. C. Lipton, and H. Heidari · 2024
Closest in time.
Glaive function calling v2 dataset, 2024
GlaiveAI · 2024
Closest in time.
Groqcloud models documentation, 2024
Groq · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Llm self defense: By self examination, llms know they are being tricked
A. Helbling, M. Phute, M. Hull, and D. H. Chau · 2023
Cited alongside, same era.
Llama guard: Llm-based input-output safeguard for human-ai conversations, 2023
H. Inan, K. Upasani, J. Chi, R. Rungta, K. Iyer, Y. Mao, M. Tontchev, Q. Hu, B. Fuller, D. Testuggine, and M. Khabsa · 2023
Cited alongside, same era.
Baseline defenses for adversarial attacks against aligned language models
N. Jain, A. Schwarzschild, Y. Wen, G. Somepalli, J. Kirchenbauer, P.-y. Chiang, M. Goldblum, A. Saha, J. Geiping, and T. Goldstein · 2023
Cited alongside, same era.
Circuit breaking: Removing model behaviors with targeted ablation
N. M. Li Maximilian, Davies Xander · 2023
Cited alongside, same era.
Agentbench: Evaluating llms as agents
X. Liu, H. Yu, H. Zhang, Y. Xu, X. Lei, H. Lai, Y. Gu, H. Ding, K. Men, K. Yang, et al · 2023
Cited alongside, same era.
Tree of attacks: Jailbreaking black-box llms automatically
A. Mehrotra, M. Zampetakis, P. Kassianik, B. Nelson, H. Anderson, Y. Singer, and A. Karbasi · 2023
Cited alongside, same era.
Augmented language models: a survey
G. Mialon, R. Dessì, M. Lomeli, C. Nalmpantis, R. Pasunuru, R. Raileanu, B. Rozière, T. Schick, J. Dwivedi-Yu, A. Celikyilmaz, et al · 2023
Cited alongside, same era.
D. Hendrycks · 2024
Closest in time.
Jailbreaking is best solved by definition
T. Kim, S. Kotha, and A. Raghunathan · 2024
Closest in time.
The wmdp benchmark: Measuring and reducing malicious use with unlearning, 2024
N. Li, A. Pan, A. Gopal, S. Yue, D. Berrios, A. Gatti, J. D. Li, A.-K. Dombrowski, S. Goel, L. Phan, G. Mukobi, N. Helm-Burger, R. Lababidi, L. Justen, A. B. Liu, M. Chen, I. Barrass, O. Zhang, X. Zhu, R. Tamirisa, B. Bharathi, A. Khoja, Z. Zhao, A. Herbert-Voss, C. B. Breuer, S. Marks, O. Patel, A. Zou, M. Mazeika, Z. Wang, P. Oswal, W. Liu, A. A. Hunt, J. Tienken-Harder, K. Y. Shih, K. Talley, J. Guan, R. Kaplan, I. Steneker, D. Campbell, B. Jokubaitis, A. Levinson, J. Wang, W. Qian, K. K. Karmakar, S. Basart, S. Fitz, M. Levine, P. Kumaraguru, U. Tupakula, V. Varadharajan, Y. Shoshitaishvili, J. Ba, K. M. Esvelt, A. Wang, and D. Hendrycks · 2024
Closest in time.
Llava-next: Improved reasoning, ocr, and world knowledge, January 2024a
H. Liu, C. Li, Y. Li, B. Li, Y. Zhang, S. Shen, and Y. J. Lee · 2024
Closest in time.
Prp: Propagating universal perturbations to attack large language model guard-rails
N. Mangaokar, A. Hooda, J. Choi, S. Chandrashekaran, K. Fawaz, S. Jha, and A. Prakash · 2024
Closest in time.
Harmbench: A standardized evaluation framework for automated red teaming and robust refusal
M. Mazeika, L. Phan, X. Yin, A. Zou, Z. Wang, N. Mu, E. Sakhaee, N. Li, S. Basart, B. Li, D. Forsyth, and D. Hendrycks · 2024
Closest in time.
Llama-3 8b instruct
Meta AI · 2024
Closest in time.
Mistral 7b v0.2
Mistral · 2024
Closest in time.
Jailbreaking chatgpt on release day
Z. Mowshowitz · 2024
Closest in time.
Chat completions ( tool_choice
OpenAI · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
M. Reid, N. Savinov, D. Teplyashin, D. Lepikhin, T. Lillicrap, J.-b. Alayrac, R. Soricut, A. Lazaridou, O. Firat, J. Schrittwieser, et al · 2024
Closest in time.
L. Schwinn, D. Dobre, S. Xhonneux, G. Gidel, and S. Gunnemann · 2024
Closest in time.
Berkeley function calling leaderboard
F. Yan, H. Mao, C. C.-J. Ji, T. Zhang, S. G. Patil, I. Stoica, and J. E. Gonzalez · 2024
Closest in time.
Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi
X. Yue, Y. Ni, K. Zhang, T. Zheng, R. Liu, G. Zhang, S. Stevens, D. Jiang, W. Ren, Y. Sun, C. Wei, B. Yu, R. Yuan, R. Sun, M. Yin, B. Zheng, Z. Yang, Y. Liu, W. Huang, H. Sun, Y. Su, and W. Chen · 2024
Closest in time.
Y. Zeng, H. Lin, J. Zhang, D. Yang, R. Jia, and W. Shi · 2024
Closest in time.
Wildchat: 1m chatGPT interaction logs in the wild
W. Zhao, X. Ren, J. Hessel, C. Cardie, Y. Choi, and Y. Deng · 2024
Closest in time.
Robust prompt optimization for defending language models against jailbreaking attacks
A. Zhou, B. Li, and H. Wang · 2024
Closest in time.