Fetching the paper…
Reading the bibliography…
AI safety is a rapidly growing area of research that seeks to prevent the harm and misuse of frontier AI technology, particularly with respect to generative AI (GenAI) tools that are capable of creating realistic and high-quality content through text prompts.
On large-batch training for deep learning: Generalization gap and sharp minima
Keskar, N. S., Mudigere, D., Nocedal, J., Smelyanskiy, M., and Tang, P. T. P · 2017
Earlier work this paper cites.
Visualizing the loss landscape of neural nets
Li, H., Xu, Z., Taylor, G., Studer, C., and Goldstein, T · 2018
Earlier work this paper cites.
Robust subspace learning: Robust pca, robust subspace tracking, and robust subspace recovery
Vaswani, N., Bouwmans, T., Javed, S., and Narayanamurthy, P · 2018
Earlier work this paper cites.
A primer on zeroth-order optimization in signal processing and machine learning
Liu, S., Chen, P.-Y., Kailkhura, B., Zhang, G., Hero, A., and Varshney, P. K · 2020
Earlier work this paper cites.
Media forensics and deepfakes: an overview
Verdoliva, L · 2020
Earlier work this paper cites.
Diffusion models beat gans on image synthesis
Dhariwal, P. and Nichol, A · 2021
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Earlier work this paper cites.
Adversarial Robustness for Machine Learning
Chen, P.-Y. and Hsieh, C.-J · 2023
Earlier work this paper cites.
RADAR: Robust AI-text detection via adversarial learning
Hu, X., Chen, P.-Y., and Ho, T.-Y · 2023
Earlier work this paper cites.
Llama guard: Llm-based input-output safeguard for human-ai conversations
Inan, H., Upasani, K., Chi, J., Rungta, R., Iyer, K., Mao, Y., Tontchev, M., Hu, Q., Fuller, B., Testuggine, D., et al · 2023
Earlier work this paper cites.
Baseline defenses for adversarial attacks against aligned language models
Jain, N., Schwarzschild, A., Wen, Y., Somepalli, G., Kirchenbauer, J., Chiang, P.-y., Goldblum, M., Saha, A., Geiping, J., and Goldstein, T · 2023
Earlier work this paper cites.
AlpacaEval: An automatic evaluator of instruction-following models
Li, X., Zhang, T., Dubois, Y., Taori, R., Gulrajani, I., Guestrin, C., Liang, P., and Hashimoto, T. B · 2023
Earlier work this paper cites.
DetectGPT: Zero-shot machine-generated text detection using probability curvature
Mitchell, E., Lee, Y., Khazatsky, A., Manning, C. D., and Finn, C · 2023
Earlier work this paper cites.
DINOv2: Learning robust visual features without supervision
Oquab, M., Darcet, T., Moutakanni, T., Vo, H., Szafraniec, M., Khalidov, V., Fernandez, P., Haziza, D., Massa, F., El-Nouby, A., et al · 2023
Earlier work this paper cites.
SmoothLLM: Defending large language models against jailbreaking attacks
Robey, A., Wong, E., Hassani, H., and Pappas, G. J · 2023
Earlier work this paper cites.
Cybersecurity for AI systems: A survey
Sangwan, R. S., Badr, Y., and Srinivasan, S. M · 2023
Cited alongside, same era.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Cited alongside, same era.
Defending chatgpt against jailbreak attack via self-reminders
Xie, Y., Yi, J., Shao, J., Curl, J., Lyu, L., Chen, Q., Xie, X., and Wu, F · 2023
Cited alongside, same era.
Universal and transferable adversarial attacks on aligned language models
Zou, A., Wang, Z., Carlini, N., Nasr, M., Kolter, J. Z., and Fredrikson, M · 2023
Cited alongside, same era.
Many-shot jailbreaking
Anil, C., Durmus, E., Rimsky, N., Sharma, M., Benton, J., Kundu, S., Batson, J., Tong, M., Mu, J., Ford, D. J., et al · 2024
Cited alongside, same era.
Navigating the safety landscape: Measuring risks in finetuning large language models
Peng, S., Chen, P.-Y., Hull, M., and Chau, D. H · 2024
Later among the works it cites.
AI risk management should incorporate both safety and security
Qi, X., Huang, Y., Zeng, Y., Debenedetti, E., Geiping, J., He, L., Huang, K., Madhushani, U., Sehwag, V., Shi, W., et al · 2024
Later among the works it cites.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R., Sharma, A., Mitchell, E., Manning, C. D., Ermon, S., and Finn, C · 2024
Later among the works it cites.
AEROBLADE: Training-free detection of latent diffusion images using autoencoder reconstruction error
Ricker, J., Lukovnikov, D., and Fischer, A · 2024
Later among the works it cites.
Defining and evaluating physical safety for large language models
Tang, Y.-C., Chen, P.-Y., and Ho, T.-Y · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bereska, L. and Gavves, E · 2024
Cited alongside, same era.
Model reprogramming: Resource-efficient cross-domain machine learning
Chen, P.-Y · 2024
Cited alongside, same era.
AI safety in generative AI large language models: A survey
Chua, J., Li, Y., Yang, S., Wang, C., and Yao, L · 2024
Cited alongside, same era.
RIGID: A training-free and model-agnostic framework for robust ai-generated image detection
He, Z., Chen, P.-Y., and Ho, T.-Y · 2024
Cited alongside, same era.
Safe LoRA: the silver lining of reducing safety risks when fine-tuning large language models
Hsu, C.-Y., Tsai, Y.-L., Lin, C.-H., Chen, P.-Y., Yu, C.-M., and Huang, C.-Y · 2024
Cited alongside, same era.
Gradient cuff: Detecting jailbreak attacks on large language models by exploring refusal loss landscapes
Hu, X., Chen, P.-Y., and Ho, T.-Y · 2024
Cited alongside, same era.
Defending large language models against jailbreak attacks via semantic smoothing
Ji, J., Hou, B., Robey, A., Pappas, G. J., Hassani, H., Zhang, Y., Wong, E., and Chang, S · 2024
Cited alongside, same era.
Later among the works it cites.
Tsai, C.-T., Ko, C.-Y., Chung, I., Wang, Y.-C. F., Chen, P.-Y., et al · 2024
Later among the works it cites.
AI safety assurance for automated vehicles: A survey on research, standardization, regulation
Ullrich, L., Buchholz, M., Dietmayer, K., and Graichen, K · 2024
Later among the works it cites.
Defensive prompt patch: A robust and interpretable defense of llms against jailbreak attacks
Xiong, C., Qi, X., Chen, P.-Y., and Ho, T.-Y · 2024
Later among the works it cites.
DF40: Toward next-generation deepfake detection
Yan, Z., Yao, T., Chen, S., Zhao, Y., Fu, X., Zhu, J., Luo, D., Wang, C., Ding, S., Wu, Y., and Yuan, L · 2024
Later among the works it cites.
Judging llm-as-a-judge with mt-bench and chatbot arena
Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E., et al · 2024
Later among the works it cites.
Jailbreaking black box large language models in twenty queries
Chao, P., Robey, A., Dobriban, E., Hassani, H., Pappas, G. J., and Wong, E · 2025
Closest in time.
Token highlighter: Inspecting and mitigating jailbreak prompts for large language models
Hu, X., Chen, P.-Y., and Ho, T.-Y · 2025
Closest in time.
Safety at scale: A comprehensive survey of large model safety
Ma, X., Gao, Y., Wang, Y., Wang, R., Wang, X., Sun, Y., Ding, Y., Xu, H., Chen, Y., Zhao, Y., et al · 2025
Closest in time.
Can AI-generated text be reliably detected? stress testing AI text detectors under various attacks
Sadasivan, V. S., Kumar, A., Balasubramanian, S., Wang, W., and Feizi, S · 2025
Closest in time.
Justice or prejudice? quantifying biases in LLM-as-a-judge
Ye, J., Wang, Y., Huang, Y., Chen, D., Zhang, Q., Moniz, N., Gao, T., Geyer, W., Huang, C., Chen, P.-Y., et al · 2025
Closest in time.