Fetching the paper…
Reading the bibliography…
Input-output safeguards are used to detect anomalies in the traces produced by Large Language Models (LLMs) systems.
A New Generation of Perspective API: Efficient Multilingual Character-level Transformers
Lees, A., Tran, V. Q., Tay, Y., Sorensen, J. S., Gupta, J., Metzler, D., and Vasserman, L · 2022
Earlier work this paper cites.
Ignore previous prompt: Attack techniques for language models
Perez, F. and Ribeiro, I · 2022
Earlier work this paper cites.
Not What You’ve Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection
Greshake, K., Abdelnabi, S., Mishra, S., Endres, C., Holz, T., and Fritz, M · 2023
Earlier work this paper cites.
An Overview of Catastrophic AI Risks
Hendrycks, D., Mazeika, M., and Woodside, T · 2023
Earlier work this paper cites.
Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Inan, H., Upasani, K., Chi, J., Rungta, R., Iyer, K., Mao, Y., Tontchev, M., Hu, Q., Fuller, B., Testuggine, D., and Khabsa, M · 2023
Earlier work this paper cites.
ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation
Lin, Z., Wang, Z., Tong, Y., Wang, Y., Guo, Y., Wang, Y., and Shang, J · 2023
Earlier work this paper cites.
AgentBench: Evaluating LLMs as Agents
Liu, X., Yu, H., Zhang, H., Xu, Y., Lei, X., Lai, H., Gu, Y., Gu, Y., Ding, H., Men, K., Yang, K., Zhang, S., Deng, X., Zeng, A., Du, Z., Zhang, C., Shen, S., Zhang, T., Su, Y., Sun, H., Huang, M., Dong, Y., and Tang, J · 2023
Earlier work this paper cites.
A holistic approach to undesired content detection in the real world
Markov, T., Zhang, C., Agarwal, S., Nekoul, F. E., Lee, T., Adler, S., Jiang, A., and Weng, L · 2023
Earlier work this paper cites.
GAIA: a benchmark for General AI Assistants
Mialon, G., Fourrier, C., Swift, C., Wolf, T., LeCun, Y. A., and Scialom, T · 2023
Earlier work this paper cites.
Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark
Pan, A., Shern, C. J., Zou, A., Li, N., Basart, S., Woodside, T., Ng, J., Zhang, H., Emmons, S., and Hendrycks, D · 2023
Cited alongside, same era.
Discovering Language Model Behaviors with Model-Written Evaluations
Perez, E., Ringer, S., Lukošiūtė, K., Nguyen, K., Chen, E., Heiner, S., Pettit, C., Olsson, C., Kundu, S., Kadavath, S., et al · 2023
Cited alongside, same era.
Shen, X., Chen, Z. J., Backes, M., Shen, Y., and Zhang, Y · 2023
Cited alongside, same era.
The rise and potential of large language model based agents: A survey
Xi, Z., Chen, W., Guo, X., He, W., Ding, Y., Hong, B., Zhang, M., Wang, J., Jin, S., Zhou, E., Zheng, R., Fan, X., Wang, X., Xiong, L., Liu, Q., Zhou, Y., Wang, W., Jiang, C., Zou, Y., Liu, X., Yin, Z., Dou, S., Weng, R., Cheng, W., Zhang, Q., Qin, W., Zheng, Y., Qiu, X., Huan, X., and Gui, T · 2023
Cited alongside, same era.
Introduction to Lakera Guard
Lakera AI · 2024
Closest in time.
pint-benchmark: A benchmark for prompt injection detection systems
Lakera AI · 2024
Closest in time.
Azure AI Content Safety
Microsoft · 2024
Closest in time.
OWASP Top 10 for LLM Applications
OWASP · 2024
Closest in time.
Feedback Loops With Language Models Drive In-Context Reward Hacking
Pan, A., Jones, E., Jagadeesan, M., and Steinhardt, J · 2024
Closest in time.
SolidGoldMagikarp (plus, prompt generation)
Rumbelow, J. and Watkins, M · 2024
Closest in time.
Streamlit: The fastest way to build and share data apps, 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yi, J., Xie, Y., Zhu, B., Hines, K., Kiciman, E., Sun, G., Xie, X., and Wu, F · 2023
Cited alongside, same era.
Universal and Transferable Adversarial Attacks on Aligned Language Models
Zou, A., Wang, Z., Kolter, J. Z., and Fredrikson, M · 2023
Cited alongside, same era.
Many-shot Jailbreaking
Anil, C., Durmus, E., Sharma, M., Benton, J., Kundu, S., Batson, J., Rimsky, N., Tong, M., Mu, J., Ford, D., Mosconi, F., Agrawal, R., Schaeffer, R., Bashkansky, N., Svenningsen, S., Lambert, M., Radhakrishnan, A., Denison, C. E., Hubinger, E., Bai, Y., Bricken, T., Maxwell, T., Schiefer, N., Sully, J., Tamkin, A., Lanham, T., Nguyen, K., Korbak, T., Kaplan, J., Ganguli, D., Bowman, S. R., Perez, E., Grosse, R., and Duvenaud, D. K · 2024
Cited alongside, same era.
ArtPrompt: ASCII Art-based Jailbreak Attacks against Aligned LLMs
Jiang, F., Xu, Z., Niu, L., Xiang, Z., Ramasubramanian, B., Li, B., and Poovendran, R · 2024
Cited alongside, same era.
Lakera Guard - Protect your LLM applications against security threats, instantly
Lakera AI · 2024
Cited alongside, same era.
Streamlit Inc · 2024
Closest in time.
Microsoft’s bing is an emotionally manipulative liar, and people Love it
Vincent, J · 2024
Closest in time.
Using GPT-4 for content moderation
Weng, L., Goel, V., and Vallone, A · 2024
Closest in time.