Fetching the paper…
Reading the bibliography…
Content Warning: This paper may contain unsafe or harmful content generated by LLMs that may be offensive to readers.
Why Should Adversarial Perturbations be Imperceptible? Rethink the Research Paradigm in Adversarial NLP. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing . 11222–11237
Yangyi Chen, Hongcheng Gao, Ganqu Cui, Fanchao Qi, Longtao Huang, Zhiyuan Liu, and Maosong Sun. 2022 · 2022
Earlier work this paper cites.
Synchromesh: Reliable code generation from pre-trained language models
Gabriel Poesia, Oleksandr Polozov, Vu Le, Ashish Tiwari, Gustavo Soares, Christopher Meek, and Sumit Gulwani. 2022 · 2022
Earlier work this paper cites.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Earlier work this paper cites.
LLaMA: Open and Efficient Foundation Language Models
Meta AI. 2023b · 2023
Earlier work this paper cites.
Are aligned neural networks adversarially aligned?. In Thirty-seventh Conference on Neural Information Processing Systems
Nicholas Carlini, Milad Nasr, Christopher A. Choquette-Choo, Matthew Jagielski, Irena Gao, Pang Wei Koh, Daphne Ippolito, Florian Tramèr, and Ludwig Schmidt. 2023 · 2023
Earlier work this paper cites.
Gemini: A Family of Highly Capable Multimodal Models
Gemini Team Google. 2023 · 2023
Earlier work this paper cites.
llama.cpp: LLM inference in C/C++
Georgi Gerganov. 2023 · 2023
Earlier work this paper cites.
Beavertails: Towards improved safety alignment of llm via a human-preference dataset
Jiaming Ji, Mickel Liu, Josef Dai, Xuehai Pan, Chi Zhang, Ce Bian, Boyuan Chen, Ruiyang Sun, Yizhou Wang, and Yaodong Yang. 2023 · 2023
Earlier work this paper cites.
Efficient Memory Management for Large Language Model Serving with PagedAttention. In Proceedings of the 29th Symposium on Operating Systems Principles . 611–626
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, and Ion Stoica. 2023 · 2023
Earlier work this paper cites.
A holistic approach to undesired content detection in the real world. In Proceedings of the Thirty-Seventh AAAI Conference on Artificial Intelligence and Thirty-Fifth Conference on Innovative Applications of Artificial Intelligence and Thirteenth Symposium on Educational Advances in Artificial Intelligence . Article 1683, 10 pages
Todor Markov, Chong Zhang, Sandhini Agarwal, Florentine Eloundou Nekoul, Theodore Lee, Steven Adler, Angela Jiang, and Lilian Weng. 2023 · 2023
Earlier work this paper cites.
Llama 2: Open Foundation and Fine-Tuned Chat Models
Meta AI. 2023 · 2023
Earlier work this paper cites.
Evaluating GPT-3 generated explanations for hateful content moderation. In Proceedings of the Thirty-Second International Joint Conference on Artificial Intelligence . Article 694, 9 pages
Han Wang, Ming Shan Hee, Md Rabiul Awal, Kenny Tsu Wei Choo, and Roy Ka-Wei Lee. 2023 · 2023
Earlier work this paper cites.
Jailbroken: How Does LLM Safety Training Fail?. In Thirty-seventh Conference on Neural Information Processing Systems
Alexander Wei, Nika Haghtalab, and Jacob Steinhardt. 2023 · 2023
Earlier work this paper cites.
Efficient Guided Generation for Large Language Models
Brandon T. Willard and Rémi Louf. 2023 · 2023
Earlier work this paper cites.
Universal and Transferable Adversarial Attacks on Aligned Language Models
Andy Zou, Zifan Wang, Nicholas Carlini, Milad Nasr, J. Zico Kolter, and Matt Fredrikson. 2023 · 2023
Earlier work this paper cites.
Introducing the next generation of Claude
Anthropic. 2024 · 2024
Earlier work this paper cites.
Reducing hallucination in structured outputs via Retrieval-Augmented Generation. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 6: Industry Track) . 228–238
Orlando Ayala and Patrice Bechard. 2024 · 2024
Earlier work this paper cites.
Output Scouting: Auditing Large Language Models for Catastrophic Responses
Andrew Bell and Joao Fonseca. 2024 · 2024
Earlier work this paper cites.
Play Guessing Game with LLM: Indirect Jailbreak Attack with Implicit Clues. In Findings of the Association for Computational Linguistics: ACL 2024 . 5135–5147
Zhiyuan Chang, Mingyang Li, Yi Liu, Junjie Wang, Qing Wang, and Yang Liu. 2024 · 2024
Earlier work this paper cites.
MASTERKEY: Automated Jailbreaking of Large Language Model Chatbots. In NDSS
Gelei Deng, Yi Liu, Yuekang Li, Kailong Wang, Ying Zhang, Zefeng Li, Haoyu Wang, Tianwei Zhang, and Yang Liu. 2024 · 2024
Earlier work this paper cites.
Phi-3 Safety Post-Training: Aligning Language Models with a "Break-Fix" Cycle
Emman Haider, Daniel Perez-Becker, Thomas Portet, Piyush Madan, Amit Garg, Atabak Ashfaq, David Majercak, Wen Wen, Dongwoo Kim, Ziyi Yang, Jianwen Zhang, Hiteshi Sharma, Blake Bullwinkel, Martin Pouliot, Amanda Minnich, Shiven Chawla, Solianna Herrera, Shahed Warreth, Maggie Engler, Gary Lopez, Nina Chikanov, Raja Sekhar Rao Dheekonda, Bolor-Erdene Jagdagdorj, Roman Lutz, Richard Lundeen, Tori Westerhoff, Pete Bryan, Christian Seifert, Ram Shankar Siva Kumar, Andrew Berkley, and Alex Kessler. 2024 · 2024
Earlier work this paper cites.
Catastrophic Jailbreak of Open-source LLMs via Exploiting Generation. In The Twelfth International Conference on Learning Representations
Yangsibo Huang, Samyak Gupta, Mengzhou Xia, Kai Li, and Danqi Chen. 2024 · 2024
Earlier work this paper cites.
Pku-saferlhf: A safety alignment preference dataset for llama family models
Jiaming Ji, Donghai Hong, Borong Zhang, Boyuan Chen, Josef Dai, Boren Zheng, Tianyi Qiu, Boxun Li, and Yaodong Yang. 2024 · 2024
Cited alongside, same era.
Automata-based constraints for language model decoding. In First Conference on Language Modeling
Terry Koo, Frederick Liu, and Luheng He. 2024 · 2024
Cited alongside, same era.
DrAttack: Prompt Decomposition and Reconstruction Makes Powerful LLMs Jailbreakers. In Findings of the Association for Computational Linguistics: EMNLP 2024 . 13891–13913
Xirui Li, Ruochen Wang, Minhao Cheng, Tianyi Zhou, and Cho-Jui Hsieh. 2024b · 2024
Cited alongside, same era.
HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal. In Proceedings of the 41st International Conference on Machine Learning , Vol. 235. 35181–35224
Mantas Mazeika, Long Phan, Xuwang Yin, Andy Zou, Zifan Wang, Norman Mu, Elham Sakhaee, Nathaniel Li, Steven Basart, Bo Li, David Forsyth, and Dan Hendrycks. 2024 · 2024
Cited alongside, same era.
Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks. In The Thirteenth International Conference on Learning Representations
Maksym Andriushchenko, Francesco Croce, and Nicolas Flammarion. 2025 · 2025
Closest in time.
Cursor: The AI-first Code Editor
Anysphere, Inc. 2023 · 2025
Closest in time.
LangChain
Harrison Chase. 2022 · 2025
Closest in time.
Xgrammar: Flexible and efficient structured generation engine for large language models
Yixin Dong, Charlie F Ruan, Yaxing Cai, Ziyi Xu, Yilong Zhao, Ruihang Lai, and Tianqi Chen. 2025 · 2025
Closest in time.
JSONSchemaBench: A Rigorous Benchmark of Structured Outputs for Language Models
Saibo Geng, Hudson Cooper, Michał Moskal, Samuel Jenkins, Julian Berman, Nathan Ranchin, Robert West, Eric Horvitz, and Harsha Nori. 2025 · 2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tree of Attacks: Jailbreaking Black-Box LLMs Automatically. In The Thirty-eighth Annual Conference on Neural Information Processing Systems
Anay Mehrotra, Manolis Zampetakis, Paul Kassianik, Blaine Nelson, Hyrum S Anderson, Yaron Singer, and Amin Karbasi. 2024 · 2024
Cited alongside, same era.
Introducing Meta Llama 3: The most capable openly available LLM to date
Meta AI. 2024 · 2024
Cited alongside, same era.
Mistral NeMo: A State-of-the-Art 12B Model with 128k Context Length
Mistral AI team. 2024 · 2024
Cited alongside, same era.
Bypassing OpenAI’s Structured Outputs: Another Simple Jailbreak
Aman Priyanshu. 2024 · 2024
Cited alongside, same era.
Visual adversarial examples jailbreak aligned large language models. In Proceedings of the AAAI conference on artificial intelligence , Vol. 38. 21527–21536
Xiangyu Qi, Kaixuan Huang, Ashwinee Panda, Peter Henderson, Mengdi Wang, and Prateek Mittal. 2024 · 2024
Cited alongside, same era.
CodeAttack: Revealing Safety Generalization Challenges of Large Language Models via Code Completion. In Findings of the Association for Computational Linguistics: ACL 2024 . 11437–11452
Qibing Ren, Chang Gao, Jing Shao, Junchi Yan, Xin Tan, Wai Lam, and Lizhuang Ma. 2024 · 2024
Cited alongside, same era.
"Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models. In Proceedings of the 2024 on ACM SIGSAC Conference on Computer and Communications Security . 1671–1685
Xinyue Shen, Zeyuan Chen, Michael Backes, Yun Shen, and Yang Zhang. 2024 · 2024
Cited alongside, same era.
A StrongREJECT for Empty Jailbreaks. In The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track
Alexandra Souly, Qingyuan Lu, Dillon Bowen, Tu Trinh, Elvis Hsieh, Sana Pandey, Pieter Abbeel, Justin Svegliato, Scott Emmons, Olivia Watkins, and Sam Toyer. 2024 · 2024
Cited alongside, same era.
AutoDAN-Turbo: A Lifelong Agent for Strategy Self-Exploration to Jailbreak LLMs. In The Thirteenth International Conference on Learning Representations
Xiaogeng Liu, Peiran Li, G. Edward Suh, Yevgeniy Vorobeychik, Zhuoqing Mao, Somesh Jha, Patrick McDaniel, Huan Sun, Bo Li, and Chaowei Xiao. 2025 · 2025
Closest in time.
Learning to Generate Structured Output with Schema Reinforcement Learning. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . 4905–4918
Yaxi Lu, Haolun Li, Xin Cong, Zhong Zhang, Yesai Wu, Yankai Lin, Zhiyuan Liu, Fangming Liu, and Maosong Sun. 2025 · 2025
Closest in time.
You Can’t Eat Your Cake and Have It Too: The Performance Degradation of LLMs with Jailbreak Defense. In THE WEB CONFERENCE 2025
Wuyuao Mai, Geng Hong, Pei Chen, Xudong Pan, Baojun Liu, Yuan Zhang, Haixin Duan, and Min Yang. 2025 · 2025
Closest in time.
LightLLM: A Python-based LLM inference and serving framework
ModelTC. 2025 · 2025
Closest in time.
GPT-4o System Card
OpenAI. 2024 · 2025
Closest in time.
Introducing Structured Outputs in the API
OpenAI. 2024 · 2025
Closest in time.
gpt-oss-120b & gpt-oss-20b Model Card
OpenAI. 2025 · 2025
Closest in time.
OpenRouter
OpenRouter Team. 2023 · 2025
Closest in time.
Safety Alignment Should be Made More Than Just a Few Tokens Deep. In The Thirteenth International Conference on Learning Representations
Xiangyu Qi, Ashwinee Panda, Kaifeng Lyu, Xiao Ma, Subhrajit Roy, Ahmad Beirami, Prateek Mittal, and Peter Henderson. 2025 · 2025
Closest in time.
Qwen. 2025 · 2025
Closest in time.
Introduction - Model Context Protocol
Model Context Protocol Team. 2025 · 2025
Closest in time.
LangGraph
The LangChain Team. 2023 · 2025
Closest in time.
SELFDEFEND: LLMs can defend themselves against jailbreaking in a practical manner. In Proceedings of the 34th USENIX Conference on Security Symposium . Article 126, 20 pages
Xunguang Wang, Daoyuan Wu, Zhenlan Ji, Zongjie Li, Pingchuan Ma, Shuai Wang, Yingjiu Li, Yang Liu, Ning Liu, and Juergen Rahmel. 2025 · 2025
Closest in time.
SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal. In The Thirteenth International Conference on Learning Representations
Tinghao Xie, Xiangyu Qi, Yi Zeng, Yangsibo Huang, Udari Madhushani Sehwag, Kaixuan Huang, Luxi He, Boyi Wei, Dacheng Li, Ying Sheng, Ruoxi Jia, Bo Li, Kai Li, Danqi Chen, Peter Henderson, and Prateek Mittal. 2025 · 2025
Closest in time.
StructTransform: A Scalable Attack Surface for Safety-Aligned Large Language Models. In Computer Security – ESORICS 2025: 30th European Symposium on Research in Computer Security, Toulouse, France, September 22–24, 2025, Proceedings, Part I . 488–507
Shehel Yoosuf, Temoor Ali, Ahmed Lekssays, Mashael AlSabah, and Issa Khalil. 2025 · 2025
Closest in time.
JBShield: Defending Large Language Models from Jailbreak Attacks through Activated Concept Analysis and Manipulation. In 34th USENIX Security Symposium (USENIX Security 25) . 8215–8234
Shenyi Zhang, Yuchen Zhai, Keyan Guo, Hongxin Hu, Shengnan Guo, Zheng Fang, Lingchen Zhao, Chao Shen, Cong Wang, and Qian Wang. 2025 · 2025
Closest in time.
Improving LLM Safety Alignment with Dual-Objective Optimization
Xuandong Zhao, Will Cai, Tianneng Shi, David Huang, Licong Lin, Song Mei, and Dawn Song. 2025 · 2025
Closest in time.