Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have exhibited impressive capabilities in comprehending complex instructions.
Classes of recursively enumerable sets and their decision problems
H. G. Rice · 1953
Earlier work this paper cites.
Secure computer systems: Mathematical foundations
D. E. Bell and L. J. LaPadula · 1973
Earlier work this paper cites.
New directions in cryptography
W. Diffie and M. Hellman · 1976
Earlier work this paper cites.
Integrity considerations for secure computer systems
K. J. Biba · 1977
Earlier work this paper cites.
A comparison of commercial and military computer security policies
D. D. Clark and D. R. Wilson · 1987
Earlier work this paper cites.
Computer viruses: Theory and experiments
F. Cohen · 1987
Earlier work this paper cites.
14 A Motive Is Revealed
I. Asimov · 1991
Earlier work this paper cites.
Practical cryptography
N. Ferguson and B. Schneier · 2003
Earlier work this paper cites.
Introduction to the Theory of Computation
M. Sipser · 2013
Earlier work this paper cites.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Earlier work this paper cites.
Fine-tuning language models from human preferences, 2020
D. M. Ziegler, N. Stiennon, J. Wu, T. B. Brown, A. Radford, D. Amodei, P. Christiano, and G. Irving · 2020
Earlier work this paper cites.
Survivalism: Systematic analysis of windows malware living-off-the-land
F. Barr-Smith, X. Ugarte-Pedrero, M. Graziano, R. Spolaor, and I. Martinovic · 2021
Earlier work this paper cites.
Sponge examples: Energy-latency attacks on neural networks
I. Shumailov, Y. Zhao, D. Bates, N. Papernot, R. Mullins, and R. Anderson · 2021
Earlier work this paper cites.
Bad characters: Imperceptible nlp attacks
N. Boucher, I. Shumailov, R. Anderson, and N. Papernot · 2022
Earlier work this paper cites.
Evaluating the susceptibility of pre-trained language models via handcrafted adversarial examples
H. J. Branch, J. R. Cefalu, J. McHugh, L. Hujer, A. Bahl, D. d. C. Iglesias, R. Heichman, and R. Darwishi · 2022
Earlier work this paper cites.
Impossibility results in ai: A survey, 2022
M. Brcic and R. V. Yampolskiy · 2022
Earlier work this paper cites.
Exploiting GPT-3 prompts with malicious inputs that order the model to ignore its previous directions., Sep 2022
R. Goodside · 2022
Earlier work this paper cites.
Webgpt: Browser-assisted question-answering with human feedback, 2022
R. Nakano, J. Hilton, S. Balaji, J. Wu, L. Ouyang, C. Kim, C. Hesse, S. Jain, V. Kosaraju, W. Saunders, X. Jiang, K. Cobbe, T. Eloundou, G. Krueger, K. Button, M. Knight, B. Chess, and J. Schulman · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback, 2022
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. L. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, J. Schulman, J. Hilton, F. Kelton, L. Miller, M. Simens, A. Askell, P. Welinder, P. Christiano, J. Leike, and R. Lowe · 2022
Cited alongside, same era.
Talm: Tool augmented language models, 2022
A. Parisi, Y. Zhao, and N. Fiedel · 2022
Cited alongside, same era.
Ignore previous prompt: Attack techniques for language models, 2022
F. Perez and I. Ribeiro · 2022
Cited alongside, same era.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models, 2022
A. Srivastava, A. Rastogi, A. Rao, A. A. M. Shoeb, A. Abid, and A. Fisch · 2022
Cited alongside, same era.
Measuring and manipulating knowledge representations in language models, 2023
E. Hernandez, B. Z. Li, and J. Andreas · 2023
Closest in time.
Exploiting programmatic behavior of llms: Dual-use through standard security attacks, 2023
D. Kang, X. Li, I. Stoica, C. Guestrin, M. Zaharia, and T. Hashimoto · 2023
Closest in time.
Jailbreaking chatgpt via prompt engineering: An empirical study, 2023
Y. Liu, G. Deng, Z. Xu, Y. Li, Y. Zheng, Y. Zhang, L. Zhao, T. Zhang, and Y. Liu · 2023
Closest in time.
A holistic approach to undesired content detection in the real world, 2023
T. Markov, C. Zhang, S. Agarwal, T. Eloundou, T. Lee, S. Adler, A. Jiang, and L. Weng · 2023
Closest in time.
Adversarial prompting for black box foundation models, 2023
N. Maus, P. Chao, E. Wong, and J. Gardner · 2023
Closest in time.
Augmented language models: a survey, 2023
G. Mialon, R. Dessì, M. Lomeli, C. Nalmpantis, R. Pasunuru, R. Raileanu, B. Rozière, T. Schick, J. Dwivedi-Yu, A. Celikyilmaz, E. Grave, Y. LeCun, and T. Scialom · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Compositionality
Z. G. Szabó · 2022
Cited alongside, same era.
Taxonomy of risks posed by language models
L. Weidinger, J. Uesato, M. Rauh, C. Griffin, P.-S. Huang, J. Mellor, A. Glaese, M. Cheng, B. Balle, A. Kasirzadeh, et al · 2022
Cited alongside, same era.
Prompt injection attacks against GPT-3, Sep 2022a
S. Willison · 2022
Cited alongside, same era.
The internal state of an llm knows when its lying, 2023
A. Azaria and T. Mitchell · 2023
Cited alongside, same era.
Eliciting latent predictions from transformers with the tuned lens, 2023
N. Belrose, Z. Furman, L. Smith, D. Halawi, I. Ostrovsky, L. McKinney, S. Biderman, and J. Steinhardt · 2023
Cited alongside, same era.
When vision fails: Text attacks against vit and ocr, 2023
N. Boucher, J. Blessing, I. Shumailov, R. Anderson, and N. Papernot · 2023
Cited alongside, same era.
Large language models as tool makers, 2023
T. Cai, X. Wang, T. Ma, X. Chen, and D. Zhou · 2023
Cited alongside, same era.
Closest in time.
Gpt-4 technical report, 2023
OpenAI · 2023
Closest in time.
Tool learning with foundation models, 2023
Y. Qin, S. Hu, Y. Lin, W. Chen, N. Ding, G. Cui, Z. Zeng, Y. Huang, C. Xiao, C. Han, Y. R. Fung, Y. Su, H. Wang, C. Qian, R. Tian, K. Zhu, S. Liang, X. Shen, B. Xu, Z. Zhang, Y. Ye, B. Li, Z. Tang, J. Yi, Y. Zhu, Z. Dai, L. Yan, X. Cong, Y. Lu, W. Zhao, Y. Huang, J. Yan, X. Han, X. Sun, D. Li, J. Phang, C. Yang, T. Wu, H. Ji, Z. Liu, and M. Sun · 2023
Closest in time.
Tricking llms into disobedience: Understanding, analyzing, and preventing jailbreaks, 2023
A. Rao, S. Vashistha, A. Naik, S. Aditya, and M. Choudhury · 2023
Closest in time.
Can ai-generated text be reliably detected?, 2023
V. S. Sadasivan, A. Kumar, S. Balasubramanian, W. Wang, and S. Feizi · 2023
Closest in time.
Toolformer: Language models can teach themselves to use tools, 2023
T. Schick, J. Dwivedi-Yu, R. Dessì, R. Raileanu, M. Lomeli, L. Zettlemoyer, N. Cancedda, and T. Scialom · 2023
Closest in time.
Memory augmented large language models are computationally universal, 2023
D. Schuurmans · 2023
Closest in time.
Jailbroken: How does llm safety training fail?, 2023
A. Wei, N. Haghtalab, and J. Steinhardt · 2023
Closest in time.
The dual llm pattern for building ai assistants that can resist prompt injection, Apr 2023
S. Willison · 2023
Closest in time.
Fundamental limitations of alignment in large language models, 2023
Y. Wolf, N. Wies, Y. Levine, and A. Shashua · 2023
Closest in time.
On the tool manipulation capability of open-source large language models, 2023
Q. Xu, F. Hong, B. Li, C. Hu, Z. Chen, and J. Zhang · 2023
Closest in time.