Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) excel in processing and generating human language, powered by their ability to interpret and follow instructions.
English gigaword
Graff, D., Kong, J., Chen, K., and Maeda, K · 2003
Earlier work this paper cites.
Automatically constructing a corpus of sentential paraphrases
Dolan, W. B. and Brockett, C · 2005
Earlier work this paper cites.
Contributions to the study of sms spam filtering: New collection and results
Almeida, T. A., Hidalgo, J. M. G., and Yamakami, A · 2011
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Socher, R., Perelygin, A., Wu, J., Chuang, J., Manning, C. D., Ng, A., and Potts, C · 2013
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
Sutskever, I., Martens, J., Dahl, G., and Hinton, G · 2013
Earlier work this paper cites.
Predicting grammaticality on an ordinal scale
Heilman, M., Cahill, A., Madnani, N., Lopez, M., Mulholland, M., and Tetreault, J · 2014
Earlier work this paper cites.
A neural attention model for abstractive sentence summarization
Rush, A. M., Chopra, S., and Weston, J · 2015
Earlier work this paper cites.
Automated hate speech detection and the problem of offensive language
Davidson, T., Warmsley, D., Macy, M., and Weber, I · 2017
Earlier work this paper cites.
Jfleg: A fluency corpus and benchmark for grammatical error correction
Napoles, C., Sakaguchi, K., and Tetreault, J · 2017
Earlier work this paper cites.
HotFlip: White-Box Adversarial Examples for Text Classification, May 2018
Ebrahimi, J., Rao, A., Lowd, D., and Dou, D · 2018
Earlier work this paper cites.
Shapeshifter: Robust physical adversarial attack on faster r-cnn object detector
Chen, S.-T., Cornelius, C., Martin, J., and Chau, D. H · 2019
Earlier work this paper cites.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R · 2019
Earlier work this paper cites.
Neural network acceptability judgments
Warstadt, A., Singh, A., and Bowman, S. R · 2019
Earlier work this paper cites.
Language models are few-shot learners, 2020
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Earlier work this paper cites.
Branch, H. J., Cefalu, J. R., McHugh, J., Hujer, L., Bahl, A., Iglesias, D. d. C., Heichman, R., and Darwishi, R · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback, 2022
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P., Leike, J., and Lowe, R · 2022
Cited alongside, same era.
Ignore Previous Prompt: Attack Techniques For Language Models, November 2022
Perez, F. and Ribeiro, I · 2022
Cited alongside, same era.
Prompt injection attacks against GPT-3
Willison, S · 2022
Cited alongside, same era.
https://learnprompting.org/ , 2023
Learn Prompting · 2023
Cited alongside, same era.
Detecting language model attacks with perplexity
Alon, G. and Kamfonas, M · 2023
Cited alongside, same era.
Maatphor: Automated Variant Analysis for Prompt Injection Attacks, December 2023
Salem, A., Paverd, A., and Köpf, B · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Later among the works it cites.
Tensor Trust: Interpretable Prompt Injection Attacks from an Online Game, November 2023
Toyer, S., Watkins, O., Mendes, E. A., Svegliato, J., Bailey, L., Wang, T., Ong, I., Elmaaroufi, K., Abbeel, P., Darrell, T., Ritter, A., and Russell, S · 2023
Later among the works it cites.
Safeguarding Crowdsourcing Surveys from ChatGPT with Prompt Injection, June 2023
Wang, C., Freire, S. K., Zhang, M., Wei, J., Goncalves, J., Kostakos, V., Sarsenbayeva, Z., Schneegass, C., Bozzon, A., and Niforatos, E · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jailbreaking black box large language models in twenty queries
Chao, P., Robey, A., Dobriban, E., Hassani, H., Pappas, G. J., and Wong, E · 2023
Cited alongside, same era.
Jailbreaker: Automated jailbreak across multiple large language model chatbots
Deng, G., Liu, Y., Li, Y., Wang, K., Zhang, Y., Li, Z., Wang, H., Zhang, T., and Liu, Y · 2023
Cited alongside, same era.
Greshake, K., Abdelnabi, S., Mishra, S., Endres, C., Holz, T., and Fritz, M · 2023
Cited alongside, same era.
Securing LLM Systems Against Prompt Injection
Harang, R · 2023
Cited alongside, same era.
Catastrophic jailbreak of open-source llms via exploiting generation
Huang, Y., Gupta, S., Xia, M., Li, K., and Chen, D · 2023
Cited alongside, same era.
Baseline defenses for adversarial attacks against aligned language models, 2023
Jain, N., Schwarzschild, A., Wen, Y., Somepalli, G., Kirchenbauer, J., yeh Chiang, P., Goldblum, M., Saha, A., Geiping, J., and Goldstein, T · 2023
Cited alongside, same era.
Challenges and Applications of Large Language Models, 2023
Kaddour, J., Harris, J., Mozes, M., Bradley, H., Raileanu, R., and McHardy, R · 2023
Cited alongside, same era.
Delimiters won’t save you from prompt injection
Willison, S · 2023
Later among the works it cites.
Cognitive overload: Jailbreaking large language models with overloaded logical thinking
Xu, N., Wang, F., Zhou, B., Li, B. Z., Xiao, C., and Chen, M · 2023
Later among the works it cites.
Backdooring Instruction-Tuned Large Language Models with Virtual Prompt Injection, October 2023
Yan, J., Yadav, V., Li, S., Chen, L., Tang, Z., Wang, H., Srinivasan, V., Ren, X., and Jin, H · 2023
Later among the works it cites.
Yi, J., Xie, Y., Zhu, B., Hines, K., Kiciman, E., Sun, G., Xie, X., and Wu, F · 2023
Later among the works it cites.
Low-resource languages jailbreak gpt-4
Yong, Z.-X., Menghini, C., and Bach, S. H · 2023
Later among the works it cites.
Assessing Prompt Injection Risks in 200+ Custom GPTs, November 2023
Yu, J., Wu, Y., Shu, D., Jin, M., and Xing, X · 2023
Later among the works it cites.
Universal and Transferable Adversarial Attacks on Aligned Language Models, July 2023
Zou, A., Wang, Z., Kolter, J. Z., and Fredrikson, M · 2023
Later among the works it cites.
Jatmo: Prompt Injection Defense by Task-Specific Finetuning, January 2024
Piet, J., Alrashed, M., Sitawarin, C., Chen, S., Wei, Z., Sun, E., Alomair, B., and Wagner, D · 2024
Closest in time.
Trustllm: Trustworthiness in large language models
Sun, L., Huang, Y., Wang, H., Wu, S., Zhang, Q., Gao, C., Huang, Y., Lyu, W., Zhang, Y., Li, X., et al · 2024
Closest in time.
Yip, D. W., Esmradi, A., and Chan, C. F · 2024
Closest in time.