Fetching the paper…
Reading the bibliography…
In this paper, we introduce a novel technique for content safety and prompt injection classification for Large Language Models.
“Visualizing Attention in Transformer-Based Language Representation Models” arXiv:1904.02679 [cs]
Jesse Vig · 1904
Earlier work this paper cites.
Alejandro Arrieta et al · 1910
Earlier work this paper cites.
Benjamin Hoover, Hendrik Strobelt and Sebastian Gehrmann · 1910
Earlier work this paper cites.
“A Unified Approach to Interpreting Model Predictions” arXiv:1705.07874 [cs]
Scott Lundberg and Su-In Lee · 2017
Earlier work this paper cites.
“Understanding intermediate layers using linear classifier probes” arXiv:1610.01644 [stat]
Guillaume Alain and Yoshua Bengio · 2018
Earlier work this paper cites.
“Evolutionary Fuzzy Systems for Explainable Artificial Intelligence: Why, When, What for, and Where to?”
Alberto Fernandez et al · 2018
Earlier work this paper cites.
Nicolas Papernot and Patrick McDaniel · 2018
Earlier work this paper cites.
“Seq2Seq-Vis: A Visual Debugging Tool for Sequence-to-Sequence Models” arXiv:1804.09299 [cs]
Hendrik Strobelt et al · 2018
Earlier work this paper cites.
“Extending Knowledge Graphs with Subjective Influence Networks for Personalized Fashion”
Kurt Bollacker, Natalia Díaz-Rodríguez and Xian Li · 2019
Earlier work this paper cites.
“Model Explainability in Deep Learning Based Natural Language Processing” arXiv:2106.07410 [cs]
Shafie Gholizadeh and Nengfeng Zhou · 2021
Earlier work this paper cites.
“Interactive Visualization and Manipulation of Attention-based Neural Machine Translation”
Jaesong Lee, Joong-Hwi Shin and Jun-Seok Kim · 2021
Earlier work this paper cites.
Kevin Wang et al · 2022
Cited alongside, same era.
Hyunsoo Cho et al · 2023
Cited alongside, same era.
“Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations” arXiv:2312.06674 [cs]
Hakan Inan et al · 2023
Cited alongside, same era.
Shuyu Jiang, Xingshu Chen and Rui Tang · 2023
Cited alongside, same era.
“LLM-Pruner: On the Structural Pruning of Large Language Models” arXiv:2305.11627 [cs]
“The Llama 3 Herd of Models” arXiv:2407.21783 [cs]
Aaron Grattafiori et al · 2024
Closest in time.
“The Unreasonable Ineffectiveness of the Deeper Layers” arXiv:2403.17887 [cs]
Andrey Gromov et al · 2024
Closest in time.
“Attention Tracker: Detecting Prompt Injection Attacks in LLMs” arXiv:2411.00348 [cs]
Kuo-Han Hung et al · 2024
Closest in time.
“Large Language Models Are Overparameterized Text Encoders” arXiv:2410.14578 [cs]
Thennal. K, Tim Fischer and Chris Biemann · 2024
Closest in time.
Lakera, 2024
“lakeraai/pint-benchmark” original-date: 2024-03-27T19:04:05Z · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Xinyin Ma, Gongfan Fang and Xinchao Wang · 2023
Cited alongside, same era.
“Towards Agile Text Classifiers for Everyone” arXiv:2302.06541 [cs]
Maximilian Mozes et al · 2023
Cited alongside, same era.
“The geometry of hidden representations of large transformer models” arXiv:2302.00294 [cs]
Lucrezia Valeriani et al · 2023
Cited alongside, same era.
“On the Explainability of Natural Language Processing Deep Models” arXiv:2210.06929 [cs]
Julia Zini and Mariette Awad · 2023
Cited alongside, same era.
Marcus Buckmann and Edward Hill · 2024
Cited alongside, same era.
“MINI-LLM: Memory-Efficient Structured Pruning for Large Language Models” arXiv:2407.11681 [cs]
Hongrong Cheng, Miao Zhang and Javen Shi · 2024
Cited alongside, same era.
Shaona Ghosh, Prasoon Varshney, Erick Galinkin and Christopher Parisien · 2024
Cited alongside, same era.
URL: https://huggingface.co/blog/leaderboard-decodingtrust
“An Introduction to AI Secure LLM Safety Leaderboard”
Cited in the paper.
Lijun Li et al · 2024
Closest in time.
Haoyan Luo and Lucia Specia · 2024
Closest in time.
“SPML: A DSL for Defending Language Models Against Prompt Attacks” arXiv:2402.11755 [cs]
Reshabh. Sharma, Vinayak Gupta and Dan Grossman · 2024
Closest in time.
“Does Representation Matter? Exploring Intermediate Layers in Large Language Models”
Oscar Skean, Md Arefin and Ravid Shwartz-Ziv · 2024
Closest in time.
“Diff-eRank: A Novel Rank-Based Metric for Evaluating Large Language Models” arXiv:2401.17139 [cs]
Lai Wei et al · 2024
Closest in time.
“LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation Dataset” arXiv:2309.11998 [cs]
Lianmin Zheng et al · 2024
Closest in time.