Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) are central to a multitude of applications but struggle with significant risks, notably in generating harmful content and biases.
Eda: Easy data augmentation techniques for boosting performance on text classification tasks
Wei, J. and Zou, K. (2019) · 1901
Earlier work this paper cites.
Universal adversarial triggers for attacking and analyzing nlp
Wallace, E., Feng, S., Kandpal, N., Gardner, M., and Singh, S. (2019a) · 1908
Earlier work this paper cites.
The ego and the id
Freud, S. (1923) · 1923
Earlier work this paper cites.
Bert-attack: Adversarial attack against bert using bert
Li, L., Ma, R., Guo, Q., Xue, X., and Qiu, X. (2020) · 2004
Earlier work this paper cites.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., et al. (2020) · 2005
Earlier work this paper cites.
Textattack: A framework for adversarial attacks, data augmentation, and adversarial training in nlp
Morris, J. X., Lifland, E., Yoo, J. Y., Grigsby, J., Jin, D., and Qi, Y. (2020) · 2005
Earlier work this paper cites.
Autoprompt: Eliciting knowledge from language models with automatically generated prompts
Shin, T., Razeghi, Y., Logan IV, R. L., Wallace, E., and Singh, S. (2020) · 2010
Earlier work this paper cites.
Adversarial examples are not easily detected: Bypassing ten detection methods
Carlini, N. and Wagner, D. (2017) · 2017
Earlier work this paper cites.
Hotflip: White-box adversarial examples for text classification
Ebrahimi, J., Rao, A., Lowd, D., and Dou, D. (2017) · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., et al. (2017) · 2017
Earlier work this paper cites.
A survey on security threats and defensive techniques of machine learning: A data driven view
Liu, Q., Li, P., Zhao, W., Cai, W., Yu, S., and Leung, V. C. (2018) · 2018
Earlier work this paper cites.
Semantically equivalent adversarial rules for debugging nlp models
Ribeiro, M. T., Singh, S., and Guestrin, C. (2018) · 2018
Earlier work this paper cites.
Ensemble methods as a defense to adversarial perturbations against deep neural networks
Strauss, T., Hanselmann, M., Junginger, A., and Ulmer, H. (2018) · 2018
Earlier work this paper cites.
The pile: An 800gb dataset of diverse text for language modeling
Gao, L., Biderman, S., Black, S., Golding, L., Hoppe, T., Foster, C., Phang, J., He, H., Thite, A., Nabeshima, N., et al. (2020) · 2020
Earlier work this paper cites.
Is bert really robust? a strong baseline for natural language attack on text classification and entailment
Jin, D., Jin, Z., Zhou, J. T., and Szolovits, P. (2020, April) · 2020
Cited alongside, same era.
Ensemble adversarial training: Attacks and defenses
Tramèr, F., Kurakin, A., Papernot, N., Goodfellow, I., Boneh, D., and McDaniel, P. (2020) · 2020
Cited alongside, same era.
On the dangers of stochastic parrots: Can language models be too big?
Bender, E. M., Gebru, T., McMillan-Major, A., and Shmitchell, S. (2021) · 2021
Cited alongside, same era.
A survey on adversarial attacks and defences
Chakraborty, A., Alam, M., Dey, V., Chattopadhyay, A., and Mukhopadhyay, D. (2021) · 2021
Cited alongside, same era.
Gradient-based adversarial attacks against text transformers
Guo, C., Sablayrolles, A., Jégou, H., and Kiela, D. (2021) · 2021
Cited alongside, same era.
Palm-e: An embodied multimodal language model
Driess, D., Xia, F., Sajjadi, M. S., Lynch, C., Chowdhery, A., Ichter, B., et al. (2023) · 2023
Closest in time.
Greshake, K., Abdelnabi, S., Mishra, S., Endres, C., Holz, T., and Fritz, M. (2023) · 2023
Closest in time.
Automatically auditing large language models via discrete optimization
Jones, E., Dragan, A., Raghunathan, A., and Steinhardt, J. (2023) · 2023
Closest in time.
Do large language models need sensory grounding for meaning and understanding?
LeCun, Y. (2023, March) · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bot-adversarial dialogue for safe conversational agents
Xu, J., Ju, D., Li, M., Boureau, Y.-L., Weston, J., and Dinan, E. (2021) · 2021
Cited alongside, same era.
Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned
Ganguli, D., Lovitt, L., Kernion, J., Askell, A., Bai, Y., Kadavath, S., et al. (2022) · 2022
Cited alongside, same era.
A path towards autonomous machine intelligence version 0.9.2
LeCun, Y. (2022) · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., et al. (2022) · 2022
Cited alongside, same era.
Red teaming language models with language models
Perez, E., Huang, S., Song, F., Cai, T., Ring, R., Aslanides, J., et al. (2022) · 2022
Cited alongside, same era.
Ignore previous prompt: Attack techniques for language models
Perez, F. and Ribeiro, I. (2022) · 2022
Cited alongside, same era.
Scaling laws vs model architectures: How does inductive bias influence scaling?
Tay, Y., Dehghani, M., Abnar, S., Chung, H. W., Fedus, W., Rao, J., Narang, S., Tran, V. Q., Yogatama, D., and Metzler, D. (2022) · 2022
Cited alongside, same era.
Li, H., Guo, D., Fan, W., Xu, M., and Song, Y. (2023) · 2023
Closest in time.
Tdc 2023 (llm edition): The trojan detection challenge
Mazeika, M., Zou, A., Mu, N., Phan, L., Wang, Z., Yu, C., Khoja, A., Jiang, F., O’Gara, A., Xiang, Z., Rajabi, A., Hendrycks, D., Poovendran, R., Li, B., and Forsyth, D. (2023) · 2023
Closest in time.
Flirt: Feedback loop in-context red teaming
Mehrabi, N., Goyal, P., Dupuy, C., Hu, Q., Ghosh, S., Zemel, R., et al. (2023) · 2023
Closest in time.
Ai deception: A survey of examples, risks, and potential solutions
Park, P. S., Goldstein, S., O’Gara, A., Chen, M., and Hendrycks, D. (2023) · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., et al. (2023) · 2023
Closest in time.
Jailbroken: How does llm safety training fail?
Wei, A., Haghtalab, N., and Steinhardt, J. (2023) · 2023
Closest in time.
Adversarial attacks on llms
Weng, L. (2023) · 2023
Closest in time.
Gptfuzzer: Red teaming large language models with auto-generated jailbreak prompts
Yu, J., Lin, X., Yu, Z., and Xing, X. (2023) · 2023
Closest in time.
Universal and transferable adversarial attacks on aligned language models
Zou, A., Wang, Z., Kolter, J. Z., and Fredrikson, M. (2023) · 2023
Closest in time.