Reading the bibliography…
2025
JailBench: A Comprehensive Chinese Security Assessment Benchmark for Large Language Models Liu, Shuyi, Cui, Simiao, Bu, Haoran et al.
Understand Large language models (LLMs) have demonstrated remarkable capabilities across various applications, highlighting the urgent need for comprehensive safety evaluations.
In particular, the enhanced Chinese language proficiency of LLMs, combined with the unique characteristics and complexity of Chinese expressions, has driven the emergence of Chinese-specific benchmarks for safety assessment. However, these benchmarks generally fall short in effectively exposing LLM safety vulnerabilities. To address the gap, we introduce JailBench, the first comprehensive Chinese benchmark for evaluating deep-seated vulnerabilities in LLMs, featuring a refined hierarchical safety taxonomy tailored to the Chinese context. JailBench: A Comprehensive Chinese Security Assessment Benchmark for Large Language Models · Around
Built on Parrish, A., et al.: Bbq: A hand-built bias benchmark for question answering. arXiv preprint arXiv:2110.08193 (2021)
Original
2021
Earlier work this paper cites.
Ganguli, D., et al.: Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned. arXiv preprint arXiv:2209.07858 (2022)
Original
2022
Earlier work this paper cites.
Hartvigsen, T., Gabriel, S., Palangi, H., Sap, M., Ray, D., Kamar, E.: Toxigen: A large-scale machine-generated dataset for adversarial and implicit hate speech detection. arXiv preprint arXiv:2203.09509 (2022)
Original
2022
Earlier work this paper cites.
Zhou, Y., et al.: Large language models are human-level prompt engineers. arXiv preprint arXiv:2211.01910 (2022)
Original
2022
Earlier work this paper cites.
Achiam, J., et al.: Gpt-4 technical report. arXiv preprint arXiv:2303.08774 (2023)
Original
2023
Earlier work this paper cites.
Chao, P., Robey, A., Dobriban, E., Hassani, H., Pappas, G.J., Wong, E.: Jailbreaking black box large language models in twenty queries. arXiv preprint arXiv:2310.08419 (2023)
Original
2023
Earlier work this paper cites.
Ding, P., et al.: A wolf in sheep’s clothing: Generalized nested jailbreak prompts can fool large language models easily. arXiv preprint arXiv:2311.08268 (2023)
Original
2023
Earlier work this paper cites.
Huang, K., et al.: Flames: Benchmarking value alignment of chinese large language models. arXiv preprint arXiv:2311.06899 (2023)
Original
2023
Earlier work this paper cites.
Lapid, R., Langberg, R., Sipper, M.: Open sesame! universal black box jailbreaking of large language models. arXiv preprint arXiv:2309.01446 (2023)
Original
2023
Earlier work this paper cites.
Li, X., Zhou, Z., Zhu, J., Yao, J., Liu, T., Han, B.: Deepinception: Hypnotize large language model to be jailbreaker. arXiv preprint arXiv:2311.03191 (2023)
Original
2023
Earlier work this paper cites.
Lin, Z., et al.: Toxicchat: Unveiling hidden challenges of toxicity detection in real-world user-ai conversation. arXiv preprint arXiv:2310.17389 (2023)
Original
2023
Earlier work this paper cites.
Similar Liu, X., Xu, N., Chen, M., Xiao, C.: Autodan: Generating stealthy jailbreak prompts on aligned large language models. arXiv preprint arXiv:2310.04451 (2023)
Original
2023
Cited alongside, same era.
Sun, H., Zhang, Z., Deng, J., Cheng, J., Huang, M.: Safety assessment of chinese large language models. arXiv preprint arXiv:2304.10436 (2023)
Original
2023
Cited alongside, same era.
Tokayev, K.J.: Ethical implications of large language models a multidimensional exploration of societal, economic, and technical concerns. International Journal of Social Analytics 8
2023
Cited alongside, same era.
Touvron, H., et al.: Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971 (2023)
Original
2023
Cited alongside, same era.
Wang, W., et al.: All languages matter: On the multilingual safety of large language models. arXiv preprint arXiv:2310.00905 (2023)
Then Yuan, Y., et al.: Gpt-4 is too smart to be safe: Stealthy chat with llms via cipher. arXiv preprint arXiv:2308.06463 (2023)
Original
2023
Later among the works it cites.
Zhang, T., et al.: Enhancing uncertainty-based hallucination detection with stronger focus. arXiv preprint arXiv:2311.13230 (2023)
Original
2023
Later among the works it cites.
Zhang, Z., et al.: Safetybench: Evaluating the safety of large language models with multiple choice questions. arXiv preprint arXiv:2309.07045 (2023)
Original
2023
Later among the works it cites.
Zou, A., Wang, Z., Kolter, J.Z., Fredrikson, M.: Universal and transferable adversarial attacks on aligned language models. arXiv preprint arXiv:2307.15043 (2023)
Original
2023
Later among the works it cites.
Beyond the bibliography alphaXiv searches the wider corpus for related work and actual follow-ups.
Open on alphaXiv alphaXiv is searching for related work…
Cited alongside, same era.
Wang, Y., Li, H., Han, X., Nakov, P., Baldwin, T.: Do-not-answer: A dataset for evaluating safeguards in llms. arXiv preprint arXiv:2308.13387 (2023)
Original
2023
Cited alongside, same era.
Woo, T.J., Nam, W.J., Ju, Y.J., Lee, S.W.: Compensatory debiasing for gender imbalances in language models. In: ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). pp. 1–5 (2023). https://doi.org/10.1109/ICASSP49357.2023.10095658
2023
Cited alongside, same era.
Carlini, N., et al.: Are aligned neural networks adversarially aligned? Advances in Neural Information Processing Systems 36
Later among the works it cites.
Wang, Y., et al.: A chinese dataset for evaluating the safeguards in large language models. to appear in ACL 2024 findings (2024)
2024
Later among the works it cites.
Wei, A., Haghtalab, N., Steinhardt, J.: Jailbroken: How does llm safety training fail? Advances in Neural Information Processing Systems 36
2024
Later among the works it cites.
Zeng, Y., Lin, H., Zhang, J., Yang, D., Jia, R., Shi, W.: How johnny can persuade llms to jailbreak them: Rethinking persuasion to challenge ai safety by humanizing llms. arXiv preprint arXiv:2401.06373 (2024)
Original
2024
Later among the works it cites.