Fetching the paper…
Reading the bibliography…
Glitch tokens, inputs that trigger unpredictable or anomalous behavior in Large Language Models (LLMs), pose significant challenges to model reliability and safety.
Analysis of a complex of statistical variables into principal components
H. Hotelling · 1933
Earlier work this paper cites.
Support-vector networks
C. Cortes · 1995
Earlier work this paper cites.
Hotflip: White-box adversarial examples for text classification
J. Ebrahimi, A. Rao, D. Lowd, and D. Dou · 2017
Earlier work this paper cites.
From louvain to leiden: guaranteeing well-connected communities
V. A. Traag, L. Waltman, and N. J. Van Eck · 2019
Earlier work this paper cites.
Autoprompt: Eliciting knowledge from language models with automatically generated prompts
T. Shin, Y. Razeghi, R. L. Logan IV, E. Wallace, and S. Singh · 2020
Earlier work this paper cites.
Evaluating large language models trained on code
M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. D. O. Pinto, J. Kaplan, H. Edwards, Y. Burda, N. Joseph, G. Brockman, et al · 2021
Earlier work this paper cites.
Codegen: An open large language model for code with multi-turn program synthesis
E. Nijkamp, B. Pang, H. Hayashi, L. Tu, H. Wang, Y. Zhou, S. Savarese, and C. Xiong · 2022
Earlier work this paper cites.
Llms accelerate annotation for medical information extraction
A. Goel, A. Gueta, O. Gilon, C. Liu, S. Erell, L. H. Nguyen, X. Hao, B. Jaber, S. Reddy, R. Kartha, et al · 2023
Earlier work this paper cites.
A. Q. Jiang, A. Sablayrolles, A. Mensch, C. Bamford, D. S. Chaplot, D. d. l. Casas, F. Bressand, G. Lengyel, G. Lample, L. Saulnier, et al · 2023
Earlier work this paper cites.
Solidgoldmagikarp (plus, prompt generation)
LessWrong Community · 2023
Earlier work this paper cites.
X. Shen, Z. Chen, M. Backes, Y. Shen, and Y. Zhang · 2023
Cited alongside, same era.
Llama 2: Open foundation and fine-tuned chat models
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, et al · 2023
Cited alongside, same era.
Augmenting black-box llms with medical textbooks for clinical question answering
Y. Wang, X. Ma, and W. Chen · 2023
Cited alongside, same era.
Universal and transferable adversarial attacks on aligned language models
A. Zou, Z. Wang, N. Carlini, M. Nasr, J. Z. Kolter, and M. Fredrikson · 2023
Cited alongside, same era.
Phi-3 technical report: A highly capable language model locally on your phone
A survey on large language models for code generation
J. Jiang, F. Wang, J. Shen, S. Kim, and S. Kim · 2024
Closest in time.
Evaluating llm-generated worked examples in an introductory programming course
B. Jury, A. Lorusso, J. Leinonen, P. Denny, and A. Luxton-Reilly · 2024
Closest in time.
Fishing for magikarp: Automatically detecting under-trained tokens in large language models
S. Land and M. Bartolo · 2024
Closest in time.
Glitch tokens in large language models: categorization taxonomy and effective detection
Y. Li, Y. Liu, G. Deng, Y. Zhang, W. Song, L. Shi, K. Wang, Y. Li, Y. Liu, and H. Wang · 2024
Closest in time.
Gemma 2: Improving open language models at a practical size
G. Team, M. Riviere, S. Pathak, P. G. Sessa, C. Hardin, S. Bhupatiraju, L. Hussenot, T. Mesnard, B. Shahriari, A. Ramé, et al · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M. Abdin, S. A. Jacobs, A. A. Awan, J. Aneja, A. Awadallah, H. Awadalla, N. Bach, A. Bahree, A. Bakhtiari, H. Behl, et al · 2024
Cited alongside, same era.
Introducing llama 3.1: Our most capable models to date
M. AI · 2024
Cited alongside, same era.
Mistral nemo: Advancing the capabilities of large language models
M. AI · 2024
Cited alongside, same era.
Qwen 2.5: Advancing ai for everyone
Q. AI · 2024
Cited alongside, same era.
Coercing llms to do and reveal (almost) anything
J. Geiping, A. Stein, M. Shu, K. Saifullah, Y. Wen, and T. Goldstein · 2024
Cited alongside, same era.
Closest in time.
Large language models for education: A survey and outlook
S. Wang, T. Xu, H. Li, C. Zhang, J. Liang, J. Tang, P. S. Yu, and Q. Wen · 2024
Closest in time.
Hard prompts made easy: Gradient-based discrete optimization for prompt tuning and discovery
Y. Wen, N. Jain, J. Kirchenbauer, M. Goldblum, J. Geiping, and T. Goldstein · 2024
Closest in time.
A. Yang, B. Yang, B. Hui, B. Zheng, B. Yu, C. Zhou, C. Li, C. Li, D. Liu, F. Huang, et al · 2024
Closest in time.
Glitchprober: Advancing effective detection and mitigation of glitch tokens in large language models
Z. Zhang, W. Bai, Y. Li, M. H. Meng, K. Wang, L. Shi, L. Li, J. Wang, and H. Wang · 2024
Closest in time.