Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) are seeing significant adoption in every type of organization due to their exceptional generative capabilities.
On the early history of the singular value decomposition,
G. W. Stewart, · 1993
Earlier work this paper cites.
Random forests,
L. Breiman, · 2001
Earlier work this paper cites.
Sqlrand: Preventing sql injection attacks,
S. W. Boyd, A. D. Keromytis, · 2004
Earlier work this paper cites.
Visualizing data using t-sne.,
L. Van der Maaten, G. Hinton, · 2008
Earlier work this paper cites.
Principal component analysis,
H. Abdi, L. J. Williams, · 2010
Earlier work this paper cites.
Scikit-learn: Machine learning in Python,
F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, E. Duchesnay, · 2011
Earlier work this paper cites.
S. Mac Lane, Categories for the working mathematician, volume 5, Springer Science & Business Media, 2013
2013
Earlier work this paper cites.
How to use t-sne effectively,
M. Wattenberg, F. Viégas, I. Johnson, · 2016
Earlier work this paper cites.
Xgboost: A scalable tree boosting system,
T. Chen, C. Guestrin, · 2016
Earlier work this paper cites.
Cross-site scripting (xss) attacks and defense mechanisms: classification and state-of-the-art,
S. Gupta, B. B. Gupta, · 2017
Earlier work this paper cites.
Minimizing finite sums with the stochastic average gradient,
M. Schmidt, N. Le Roux, F. Bach, · 2017
Earlier work this paper cites.
Umap: Uniform manifold approximation and projection for dimension reduction,
L. McInnes, J. Healy, J. Melville, · 2018
Earlier work this paper cites.
Linguistic knowledge and transferability of contextual representations,
N. F. Liu, M. Gardner, Y. Belinkov, M. E. Peters, N. A. Smith, · 2019
Earlier work this paper cites.
Deberta: Decoding-enhanced bert with disentangled attention,
P. He, X. Liu, J. Gao, W. Chen, · 2020
Earlier work this paper cites.
Holistic evaluation of language models,
P. Liang, R. Bommasani, T. Lee, D. Tsipras, D. Soylu, M. Yasunaga, Y. Zhang, D. Narayanan, Y. Wu, A. Kumar, et al., · 2022
Earlier work this paper cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback,
Y. Bai, A. Jones, K. Ndousse, A. Askell, A. Chen, N. DasSarma, D. Drain, S. Fort, D. Ganguli, T. Henighan, et al., · 2022
Earlier work this paper cites.
Ignore previous prompt: Attack techniques for language models,
F. Perez, I. Ribeiro, · 2022
Cited alongside, same era.
Text and code embeddings by contrastive pre-training,
A. Neelakantan, T. Xu, R. Puri, A. Radford, J. M. Han, J. Tworek, Q. Yuan, N. Tezak, J. W. Kim, C. Hallacy, et al., · 2022
Cited alongside, same era.
Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned,
D. Ganguli, L. Lovitt, J. Kernion, A. Askell, Y. Bai, S. Kadavath, B. Mann, E. Perez, N. Schiefer, K. Ndousse, et al., · 2022
Cited alongside, same era.
Using gpt-eliezer against chatgpt jailbreaking. 2022,
S. Armstrong, R. Gorman, · 2022
Cited alongside, same era.
Efficient few-shot learning without prompts,
L. Tunstall, N. Reimers, U. E. S. Jo, L. Bates, D. Korat, M. Wasserblat, O. Pereg, · 2022
X. Shen, Z. Chen, M. Backes, Y. Shen, Y. Zhang, · 2023
Later among the works it cites.
Detecting language model attacks with perplexity,
G. Alon, M. Kamfonas, · 2023
Later among the works it cites.
Baseline defenses for adversarial attacks against aligned language models,
N. Jain, A. Schwarzschild, Y. Wen, G. Somepalli, J. Kirchenbauer, P.-y. Chiang, M. Goldblum, A. Saha, J. Geiping, T. Goldstein, · 2023
Later among the works it cites.
Benchmarking and defending against indirect prompt injection attacks on large language models,
J. Yi, Y. Xie, B. Zhu, K. Hines, E. Kiciman, G. Sun, X. Xie, F. Wu, · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
The science of detecting llm-generated texts,
R. Tang, Y.-N. Chuang, X. Hu, · 2023
Cited alongside, same era.
Booookscore: A systematic exploration of book-length summarization in the era of llms,
Y. Chang, K. Lo, T. Goyal, M. Iyyer, · 2023
Cited alongside, same era.
Enhancing financial sentiment analysis via retrieval augmented large language models,
B. Zhang, H. Yang, T. Zhou, M. Ali Babar, X.-Y. Liu, · 2023
Cited alongside, same era.
Chatgpt and large language model (llm) chatbots: The current state of acceptability and a proposal for guidelines on utilization in academic medicine,
J. K. Kim, M. Chua, M. Rickard, A. Lorenzo, · 2023
Cited alongside, same era.
Exploring ai ethics of chatgpt: A diagnostic analysis,
T. Y. Zhuo, Y. Huang, C. Chen, Z. Xing, · 2023
Cited alongside, same era.
Assessing prompt injection risks in 200+ custom gpts,
J. Yu, Y. Wu, D. Shu, M. Jin, X. Xing, · 2023
Cited alongside, same era.
Prompt injection attack against llm-integrated applications,
Y. Liu, G. Deng, Y. Li, K. Wang, T. Zhang, Y. Liu, H. Wang, Y. Zheng, Y. Liu, · 2023
Cited alongside, same era.
Ignore this title and hackaprompt: Exposing systemic vulnerabilities of llms through a global prompt hacking competition,
S. Schulhoff, J. Pinto, A. Khan, L.-F. Bouchard, C. Si, S. Anati, V. Tagliabue, A. Kost, C. Carnahan, J. Boyd-Graber, · 2023
Later among the works it cites.
J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al., · 2023
Later among the works it cites.
Eliciting latent predictions from transformers with the tuned lens,
N. Belrose, Z. Furman, L. Smith, D. Halawi, I. Ostrovsky, L. McKinney, S. Biderman, J. Steinhardt, · 2023
Later among the works it cites.
Is your code generated by chatgpt really correct? rigorous evaluation of large language models for code generation,
J. Liu, C. S. Xia, Y. Wang, L. Zhang, · 2024
Closest in time.
Exploring human-like translation strategy with large language models,
Z. He, T. Liang, W. Jiao, Z. Zhang, Y. Yang, R. Wang, Z. Tu, S. Shi, X. Wang, · 2024
Closest in time.
Jailbroken: How does llm safety training fail?,
A. Wei, N. Haghtalab, J. Steinhardt, · 2024
Closest in time.
Rainbow teaming: Open-ended generation of diverse adversarial prompts,
M. Samvelyan, S. C. Raparthy, A. Lupu, E. Hambro, A. H. Markosyan, M. Bhatt, Y. Mao, M. Jiang, J. Parker-Holder, J. Foerster, et al., · 2024
Closest in time.
Special characters attack: Toward scalable training data extraction from large language models,
Y. Bai, G. Pei, J. Gu, Y. Yang, X. Ma, · 2024
Closest in time.
Struq: Defending against prompt injection with structured queries,
S. Chen, J. Piet, C. Sitawarin, D. Wagner, · 2024
Closest in time.
ProtectAI.com, Fine-tuned deberta-v3-base for prompt injection detection, 2024. URL: https://huggingface.co/ProtectAI/deberta-v3-base-prompt-injection-v2
2024
Closest in time.
Meta, Model card - prompt guard, 2024. URL: https://huggingface.co/meta-llama/Prompt-Guard-86M
2024
Closest in time.