Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have achieved remarkable progress across a wide range of applications.
Cervone, D., Shoda, Y.: Social-cognitive theories and the coherence of personality. The coherence of personality: Social-cognitive bases of consistency, variability, and organization pp. 3–33 (1999)
1999
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2022
Earlier work this paper cites.
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al.: Training language models to follow instructions with human feedback. Advances in neural information processing systems 35
2022
Earlier work this paper cites.
2023
Earlier work this paper cites.
Chao, P., Robey, A., Dobriban, E., Hassani, H., Pappas, G.J., Wong, E.: Jailbreaking black box large language models in twenty queries. In: R0-FoMo: Robustness of Few-shot and Zero-shot Learning in Large Foundation Models (2023)
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
OpenAI: Openai moderation endpoint api (2023), https://platform.openai.com/docs/guides/moderation
2023
Cited alongside, same era.
Rafailov, R., Sharma, A., Mitchell, E., Manning, C.D., Ermon, S., Finn, C.: Direct preference optimization: Your language model is secretly a reward model. Advances in neural information processing systems 36
2023
Cited alongside, same era.
Liu, T., Zhang, Y., Zhao, Z., Dong, Y., Meng, G., Chen, K.: Making them ask and answer: Jailbreaking large language models in few queries via disguise and reconstruction. In: 33rd USENIX Security Symposium (USENIX Security 24). pp. 4711–4728 (2024)
2024
Closest in time.
Llama Team, A..M.: The llama 3 herd of models (2024), https://arxiv.org/abs/2407.21783
2024
Closest in time.
2024
Closest in time.
Sun, Z., Shen, Y., Zhou, Q., Zhang, H., Chen, Z., Cox, D., Yang, Y., Gan, C.: Principle-driven self-alignment of language models from scratch with minimal human supervision. Advances in Neural Information Processing Systems 36
2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2023
Cited alongside, same era.
Chung, H.W., Hou, L., Longpre, S., Zoph, B., Tay, Y., Fedus, W., Li, Y., Wang, X., Dehghani, M., Brahma, S., et al.: Scaling instruction-finetuned language models. Journal of Machine Learning Research 25
2024
Cited alongside, same era.
2024
Cited alongside, same era.
Li, X., Zhou, Z., Zhu, J., Yao, J., Liu, T., Han, B.: Deepinception: Hypnotize large language model to be jailbreaker. In: Neurips Safe Generative AI Workshop 2024 (2024)
2024
Cited alongside, same era.
Wei, A., Haghtalab, N., Steinhardt, J.: Jailbroken: How does llm safety training fail? Advances in Neural Information Processing Systems 36
2024
Closest in time.
Yu, J., Lin, X., Yu, Z., Xing, X.: LLM-Fuzzer: Scaling assessment of large language model jailbreaks. In: 33rd USENIX Security Symposium (USENIX Security 24). pp. 4657–4674. USENIX Association, Philadelphia, PA (Aug 2024), https://www.usenix.org/conference/usenixsecurity24/presentation/yu-jiahao
2024
Closest in time.
Yuan, Y., Jiao, W., Wang, W., Huang, J.t., He, P., Shi, S., Tu, Z.: Gpt-4 is too smart to be safe: Stealthy chat with llms via cipher. In: ICLR (2024)
2024
Closest in time.