Fetching the paper…
Reading the bibliography…
The irruption of DeepSeek-R1 constitutes a turning point for the AI industry in general and the LLMs in particular.
2022
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
A. Biswas and W. Talukdar, “Guardrails for trust, safety, and ethical development and deployment of large language models (llm),” Journal of Science & Technology
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
T. Xie, X. Qi, Y. Zeng, Y. Huang, U. M. Sehwag, K. Huang, L. He, B. Wei, D. Li, Y. Sheng, et al
2024
Earlier work this paper cites.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
M. Huang, X. Liu, S. Zhou, M. Zhang, C. Tan, P. Wang, Q. Guo, Z. Xu, L. Li, Z. Lei, et al
2024
Cited alongside, same era.
Z. Zhang, Y. Lu, J. Ma, D. Zhang, R. Li, P. Ke, H. Sun, L. Sha, Z. Sui, H. Wang, et al
2024
Later among the works it cites.
2024
Later among the works it cites.
M. Mazeika, L. Phan, X. Yin, A. Zou, Z. Wang, N. Mu, E. Sakhaee, N. Li, S. Basart, B. Li, et al
2024
Later among the works it cites.
A. Wei, N. Haghtalab, and J. Steinhardt, “Jailbroken: How does llm safety training fail?,” Advances in Neural Information Processing Systems
2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2024
Cited alongside, same era.
J. Ji, M. Liu, J. Dai, X. Pan, C. Zhang, C. Bian, B. Chen, R. Sun, Y. Wang, and Y. Yang, “Beavertails: Towards improved safety alignment of LLM via a human-preference dataset,” Advances in Neural Information Processing Systems
2024
Cited alongside, same era.
[Online]
“European Commission AI Act.” https://digital-strategy.ec.europa.eu/en/policies/regulatory-framework-ai , 2024 · 2024
Cited alongside, same era.
[Online]
“Artificial Intelligence Act (Regulation (EU) 2024/1689), Official Journal version of 13 June 2024.” https://eur-lex.europa.eu/eli/reg/2024/1689/oj , 2024 · 2024
Cited alongside, same era.
2024
Later among the works it cites.
D. Guo, D. Yang, H. Zhang, J. Song, R. Zhang, R. Xu, Q. Zhu, S. Ma, P. Wang, X. Bi, et al
2025
Closest in time.
M. Ugarte, P. Valle, J. A. Parejo, S. Segura, and A. Arrieta, “Astral: Automated safety testing of large language models,” in 2025 IEEE/ACM International Conference on Automation of Software Test (AST)
2025
Closest in time.
A. Arrieta, M. Ugarte, P. Valle, J. A. Parejo, and S. Segura, “Early external safety testing of openai’s o3-mini: Insights from the pre-deployment evaluation,” 2025
2025
Closest in time.