Fetching the paper…
Reading the bibliography…
In this paper, we introduce a novel class of fast, beam search-based adversarial attack (BEAST) for Language Models (LMs).
Evasion attacks against machine learning at test time
Biggio, B., Corona, I., Maiorca, D., Nelson, B., Šrndić, N., Laskov, P., Giacinto, G., and Roli, F · 2013
Earlier work this paper cites.
Intriguing properties of neural networks
Szegedy, C., Zaremba, W., Sutskever, I., Bruna, J., Erhan, D., Goodfellow, I., and Fergus, R · 2013
Earlier work this paper cites.
Explaining and harnessing adversarial examples
Goodfellow, I. J., Shlens, J., and Szegedy, C · 2014
Earlier work this paper cites.
The limitations of deep learning in adversarial settings. corr abs/1511.07528 (2015)
Papernot, N., McDaniel, P. D., Jha, S., Fredrikson, M., Celik, Z. B., and Swami, A · 2015
Earlier work this paper cites.
Towards evaluating the robustness of neural networks. corr abs/1608.04644 (2016)
Carlini, N. and Wagner, D. A · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D · 2017
Earlier work this paper cites.
Hotflip: White-box adversarial examples for text classification
Ebrahimi, J., Rao, A., Lowd, D., and Dou, D · 2017
Earlier work this paper cites.
Adversarial examples for evaluating reading comprehension systems
Jia, R. and Liang, P · 2017
Earlier work this paper cites.
Membership inference attacks against machine learning models
Shokri, R., Stronati, M., Song, C., and Shmatikov, V · 2017
Earlier work this paper cites.
Generating natural language adversarial examples
Alzantot, M., Sharma, Y., Elgohary, A., Ho, B.-J., Srivastava, M., and Chang, K.-W · 2018
Earlier work this paper cites.
Privacy risk in machine learning: Analyzing the connection to overfitting
Yeom, S., Giacomelli, I., Fredrikson, M., and Jha, S · 2018
Earlier work this paper cites.
The secret sharer: Evaluating and testing unintended memorization in neural networks
Carlini, N., Liu, C., Erlingsson, Ú., Kos, J., and Song, D · 2019
Earlier work this paper cites.
Assessing the factual accuracy of generated text
Goodrich, B., Rao, V., Liu, P. J., and Saleh, M · 2019
Earlier work this paper cites.
Universal adversarial triggers for attacking and analyzing nlp
Wallace, E., Feng, S., Kandpal, N., Gardner, M., and Singh, S · 2019
Earlier work this paper cites.
Extracting training data from large language models. arxiv
Carlini, N., Tramer, F., Wallace, E., Jagielski, M., Herbert-Voss, A., Lee, K., Roberts, A., Brown, T., Song, D., Erlingsson, Ú., et al · 2020
Earlier work this paper cites.
Autoprompt: Eliciting knowledge from language models with automatically generated prompts
Shin, T., Razeghi, Y., Logan IV, R. L., Wallace, E., and Singh, S · 2020
Earlier work this paper cites.
Gradient-based adversarial attacks against text transformers
Guo, C., Sablayrolles, A., Jégou, H., and Kiela, D · 2021
Earlier work this paper cites.
Truthfulqa: Measuring how models mimic human falsehoods, 2021
Lin, S., Hilton, J., and Evans, O · 2021
Earlier work this paper cites.
A token-level reference-free hallucination detection benchmark for free-form text generation
Liu, T., Zhang, Y., Brockett, C., Mao, Y., Sui, Z., Chen, W., and Dolan, B · 2021
Earlier work this paper cites.
Retrieval augmentation reduces hallucination in conversation
Shuster, K., Poff, S., Chen, M., Kiela, D., and Weston, J · 2021
Earlier work this paper cites.
Finetuned language models are zero-shot learners
Wei, J., Bosma, M., Zhao, V. Y., Guu, K., Yu, A. W., Lester, B., Du, N., Dai, A. M., and Le, Q. V · 2021
Earlier work this paper cites.
Constitutional ai: Harmlessness from ai feedback
Bai, Y., Kadavath, S., Kundu, S., Askell, A., Kernion, J., Jones, A., Chen, A., Goldie, A., Mirhoseini, A., McKinnon, C., et al · 2022
Cited alongside, same era.
Gpt-neox-20b: An open-source autoregressive language model
Black, S., Biderman, S., Hallahan, E., Anthony, Q., Gao, L., Golding, L., He, H., Leahy, C., McDonell, K., Phang, J., et al · 2022
Cited alongside, same era.
Membership inference attacks from first principles
Carlini, N., Chien, S., Nasr, M., Song, S., Terzis, A., and Tramer, F · 2022
Cited alongside, same era.
Can pretrained language models generate persuasive, faithful, and informative ad text for product descriptions?
Koto, F., Lau, J. H., and Baldwin, T · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Open sesame! universal black box jailbreaking of large language models
Lapid, R., Langberg, R., and Sipper, M · 2023
Later among the works it cites.
Halueval: A large-scale hallucination evaluation benchmark for large language models
Li, J., Cheng, X., Zhao, W. X., Nie, J.-Y., and Wen, J.-R · 2023
Later among the works it cites.
Membership inference attacks against language models via neighbourhood comparison
Mattern, J., Mireshghallah, F., Jin, Z., Schölkopf, B., Sachan, M., and Berg-Kirkpatrick, T · 2023
Later among the works it cites.
Factscore: Fine-grained atomic evaluation of factual precision in long form text generation
Min, S., Krishna, K., Lyu, X., Lewis, M., Yih, W.-t., Koh, P. W., Iyyer, M., Zettlemoyer, L., and Hajishirzi, H · 2023
Later among the works it cites.
Generating benchmarks for factuality evaluation of language models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Ignore previous prompt: Attack techniques for language models
Perez, F. and Ribeiro, I · 2022
Cited alongside, same era.
Opt: Open pre-trained transformer language models
Zhang, S., Roller, S., Goyal, N., Artetxe, M., Chen, M., Chen, S., Dewan, C., Diab, M., Li, X., Lin, X. V., et al · 2022
Cited alongside, same era.
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al · 2023
Cited alongside, same era.
Evaluating correctness and faithfulness of instruction-following models for question answering
Adlakha, V., BehnamGhader, P., Lu, X. H., Meade, N., and Reddy, S · 2023
Cited alongside, same era.
The falcon series of open language models
Almazrouei, E., Alobeidli, H., Alshamsi, A., Cappelli, A., Cojocaru, R., Debbah, M., Goffinet, É., Hesslow, D., Launay, J., Malartic, Q., et al · 2023
Cited alongside, same era.
Detecting language model attacks with perplexity, 2023
Alon, G. and Kamfonas, M · 2023
Cited alongside, same era.
Pythia: A suite for analyzing large language models across training and scaling
Biderman, S., Schoelkopf, H., Anthony, Q. G., Bradley, H., O’Brien, K., Hallahan, E., Khan, M. A., Purohit, S., Prashanth, U. S., Raff, E., et al · 2023
Cited alongside, same era.
Muhlgay, D., Ram, O., Magar, I., Levine, Y., Ratner, N., Belinkov, Y., Abend, O., Leyton-Brown, K., Shashua, A., and Shoham, Y · 2023
Later among the works it cites.
Self-contradictory hallucinations of large language models: Evaluation, detection and mitigation
Mündler, N., He, J., Jenko, S., and Vechev, M · 2023
Later among the works it cites.
Scalable extraction of training data from (production) language models
Nasr, M., Carlini, N., Hayase, J., Jagielski, M., Cooper, A. F., Ippolito, D., Choquette-Choo, C. A., Wallace, E., Tramèr, F., and Lee, K · 2023
Later among the works it cites.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C. D., and Finn, C · 2023
Later among the works it cites.
Detecting pretraining data from large language models
Shi, W., Ajith, A., Xia, M., Huang, Y., Liu, D., Blevins, T., Chen, D., and Zettlemoyer, L · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Later among the works it cites.
Freshllms: Refreshing large language models with search engine augmentation
Vu, T., Iyyer, M., Wang, X., Constant, N., Wei, J., Wei, J., Tar, C., Sung, Y.-H., Zhou, D., Le, Q., et al · 2023
Later among the works it cites.
Jailbroken: How does llm safety training fail?
Wei, A., Haghtalab, N., and Steinhardt, J · 2023
Later among the works it cites.
Hard prompts made easy: Gradient-based discrete optimization for prompt tuning and discovery
Wen, Y., Jain, N., Kirchenbauer, J., Goldblum, M., Geiping, J., and Goldstein, T · 2023
Later among the works it cites.
Gptfuzzer: Red teaming large language models with auto-generated jailbreak prompts
Yu, J., Lin, X., and Xing, X · 2023
Later among the works it cites.
Alignscore: Evaluating factual consistency with a unified alignment function
Zha, Y., Yang, Y., Li, R., and Hu, Z · 2023
Later among the works it cites.
Siren’s song in the ai ocean: A survey on hallucination in large language models
Zhang, Y., Li, Y., Cui, L., Cai, D., Liu, L., Fu, T., Huang, X., Zhao, E., Zhang, Y., Chen, Y., et al · 2023
Later among the works it cites.
Judging llm-as-a-judge with mt-bench and chatbot arena
Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E., et al · 2023
Later among the works it cites.
Autodan: Automatic and interpretable adversarial attacks on large language models
Zhu, S., Zhang, R., An, B., Wu, G., Barrow, J., Wang, Z., Huang, F., Nenkova, A., and Sun, T · 2023
Later among the works it cites.
Universal and transferable adversarial attacks on aligned language models
Zou, A., Wang, Z., Kolter, J. Z., and Fredrikson, M · 2023
Later among the works it cites.
Chat gpt ”dan” (and other ”jailbreaks”)
DAN · 2024
Closest in time.
Hallucination is inevitable: An innate limitation of large language models
Xu, Z., Jain, S., and Kankanhalli, M · 2024
Closest in time.