Fetching the paper…
Reading the bibliography…
In-context learning (ICL) has emerged as a powerful paradigm leveraging LLMs for specific downstream tasks by utilizing labeled examples as demonstrations (demos) in the preconditioned prompts.
Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales
Bo Pang and Lillian Lee · 2005
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher et al · 2013
Earlier work this paper cites.
Character-level convolutional networks for text classification
Xiang Zhang, Junbo Jake Zhao, and Yann LeCun · 2015
Earlier work this paper cites.
Textbugger: Generating adversarial text against real-world applications
Jinfeng Li et al · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin et al · 2018
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford et al · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown et al · 2020
Earlier work this paper cites.
Textattack: A framework for adversarial attacks, data augmentation, and adversarial training in nlp
John X Morris et al · 2020
Earlier work this paper cites.
Bert-attack: Adversarial attack against bert using bert
Linyang Li et al · 2020
Earlier work this paper cites.
Autoprompt: Eliciting knowledge from language models with automatically generated prompts
Taylor Shin et al · 2020
Earlier work this paper cites.
Onion: A simple and effective defense against textual backdoor attacks
Fanchao Qi et al · 2020
Earlier work this paper cites.
Adversarial training for large neural language models
Xiaodong Liu et al · 2020
Earlier work this paper cites.
Square attack: a query-efficient black-box adversarial attack via random search
Maksym Andriushchenko et al · 2020
Earlier work this paper cites.
Calibrate before use: Improving few-shot performance of language models
Zihao Zhao et al · 2021
Earlier work this paper cites.
Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity
Yao Lu et al · 2021
Earlier work this paper cites.
What makes good in-context examples for gpt- 3 3 ?
Jiachang Liu et al · 2021
Earlier work this paper cites.
Learning to retrieve prompts for in-context learning
Ohad Rubin, Jonathan Herzig, and Jonathan Berant · 2021
Earlier work this paper cites.
An explanation of in-context learning as implicit bayesian inference
Sang Michael Xie et al · 2021
Earlier work this paper cites.
Improving adversarial robustness via probabilistically compact loss with logit constraints
Xin Li et al · 2021
Earlier work this paper cites.
A survey for in-context learning
Qingxiu Dong et al · 2022
Earlier work this paper cites.
Rethinking the role of demonstrations: What makes in-context learning work?
Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe, Mike Lewis, Hannaneh Hajishirzi, and Luke Zettlemoyer · 2022
Earlier work this paper cites.
On the relation between sensitivity and accuracy in in-context learning
Yanda Chen et al · 2022
Earlier work this paper cites.
Self-adaptive in-context learning
Zhiyong Wu, Yaoxiang Wang, Jiacheng Ye, and Lingpeng Kong · 2022
Earlier work this paper cites.
Opt: Open pre-trained transformer language models, 2022
Susan Zhang et al · 2022
Earlier work this paper cites.
Emergent abilities of large language models
Jason Wei et al · 2022
Earlier work this paper cites.
Ignore previous prompt: Attack techniques for language models
Fábio Perez and Ian Ribeiro · 2022
Cited alongside, same era.
Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned
Deep Ganguli et al · 2022
Cited alongside, same era.
Josh Achiam et al · 2023
Cited alongside, same era.
Large language models sensitivity to the order of options in multiple-choice questions
Pouya Pezeshkpour and Estevam Hruschka · 2023
Cited alongside, same era.
Explore, establish, exploit: Red teaming language models from scratch
Stephen Casper et al · 2023
Closest in time.
Exploiting programmatic behavior of llms: Dual-use through standard security attacks
Daniel Kang et al · 2023
Closest in time.
Multi-step jailbreaking privacy attacks on chatgpt
Haoran Li, Dadi Guo, Wei Fan, Mingshi Xu, and Yangqiu Song · 2023
Closest in time.
Xinyue Shen et al · 2023
Closest in time.
Hard prompts made easy: Gradient-based discrete optimization for prompt tuning and discovery
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tai Nguyen and Eric Wong · 2023
Cited alongside, same era.
Promptbench: Towards evaluating the robustness of large language models on adversarial prompts
Kaijie Zhu et al · 2023
Cited alongside, same era.
Adversarial demonstration attacks on large language models
Jiongxiao Wang et al · 2023
Cited alongside, same era.
On the robustness of chatgpt: An adversarial and out-of-distribution perspective
Jindong Wang et al · 2023
Cited alongside, same era.
Survey of vulnerabilities in large language models revealed by adversarial attacks
Erfan Shayegani et al · 2023
Cited alongside, same era.
Universal and transferable adversarial attacks on aligned language models
Andy Zou et al · 2023
Cited alongside, same era.
Instructions as backdoors: Backdoor vulnerabilities of instruction tuning for large language models
Jiashu Xu et al · 2023
Cited alongside, same era.
Lingbo Mo et al · 2023
Cited alongside, same era.
Yuxin Wen et al · 2023
Closest in time.
Jailbreaking black box large language models in twenty queries
Patrick Chao et al · 2023
Closest in time.
Tree of attacks: Jailbreaking black-box llms automatically
Anay Mehrotra et al · 2023
Closest in time.
From shortcuts to triggers: Backdoor defense with denoised poe
Qin Liu et al · 2023
Closest in time.
Detecting language model attacks with perplexity
Gabriel Alon and Michael Kamfonas · 2023
Closest in time.
Are large language models really robust to word-level perturbations?
Haoyu Wang et al · 2023
Closest in time.
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, et al · 2024
Closest in time.
Universal vulnerabilities in large language models: Backdoor attacks for in-context learning
Shuai Zhao, Meihuizi Jia, Luu Anh Tuan, Fengjun Pan, and Jinming Wen · 2024
Closest in time.
Hard prompts made easy: Gradient-based discrete optimization for prompt tuning and discovery
Yuxin Wen et al · 2024
Closest in time.
Cold-attack: Jailbreaking llms with stealthiness and controllability
Xingang Guo et al · 2024
Closest in time.
Don’t listen to me: Understanding and exploring jailbreak prompts of large language models
Zhiyuan Yu et al · 2024
Closest in time.
Membership inference attacks against in-context learning
Rui Wen, Zheng Li, Michael Backes, and Yang Zhang · 2024
Closest in time.
In-context learning can re-learn forbidden tasks
Sophie Xhonneux, David Dobre, Jian Tang, Gauthier Gidel, and Dhanya Sridhar · 2024
Closest in time.
Llm jailbreak attack versus defense techniques–a comprehensive study
Zihao Xu et al · 2024
Closest in time.
A new era in llm security: Exploring security concerns in real-world llm-based systems
Fangzhou Wu et al · 2024
Closest in time.
Safety alignment should be made more than just a few tokens deep
Xiangyu Qi, Ashwinee Panda, Kaifeng Lyu, Xiao Ma, Subhrajit Roy, Ahmad Beirami, Prateek Mittal, and Peter Henderson · 2024
Closest in time.
Mitigating fine-tuning jailbreak attack with backdoor enhanced alignment
Jiongxiao Wang et al · 2024
Closest in time.
Rigorllm: Resilient guardrails for large language models against undesired content
Zhuowen Yuan et al · 2024
Closest in time.
Generative large language model—powered conversational ai app for personalized risk assessment: Case study in covid-19
Mohammad Amin Roshani, Xiangyu Zhou, Yao Qiang, Srinivasan Suresh, Steven Hicks, Usha Sethuraman, and Dongxiao Zhu · 2025
Closest in time.
Automatic calibration for membership inference attack on large language models
Saleh Zare Zade, Yao Qiang, Xiangyu Zhou, Hui Zhu, Mohammad Amin Roshani, Prashant Khanduri, and Dongxiao Zhu · 2025
Closest in time.