Fetching the paper…
Reading the bibliography…
As LLMs increasingly impact safety-critical applications, ensuring their safety using guardrails remains a key challenge.
Fine-tuning language models from human preferences
Ziegler, D. M., Stiennon, N., Wu, J., Brown, T. B., Radford, A., Amodei, D., Christiano, P., and Irving, G · 2019
Earlier work this paper cites.
Challenges in automated debiasing for toxic language detection
Zhou, X · 2020
Earlier work this paper cites.
A general language assistant as a laboratory for alignment
Askell, A., Bai, Y., Chen, A., Drain, D., Ganguli, D., Henighan, T., Jones, A., Joseph, N., Mann, B., DasSarma, N., et al · 2021
Earlier work this paper cites.
Process for adapting language models to society (palms) with values-targeted datasets
Solaiman, I. and Dennison, C · 2021
Earlier work this paper cites.
Bert-beta: A proactive probabilistic approach to text moderation
Tan, F., Hu, Y., Yen, K., and Hu, C · 2021
Earlier work this paper cites.
Challenges in detoxifying language models
Welbl, J., Glaese, A., Uesato, J., Dathathri, S., Mellor, J., Hendricks, L. A., Anderson, K., Kohli, P., Coppin, B., and Huang, P.-S · 2021
Earlier work this paper cites.
Recursively summarizing books with human feedback
Wu, J., Ouyang, L., Ziegler, D. M., Stiennon, N., Lowe, R., Leike, J., and Christiano, P · 2021
Earlier work this paper cites.
Promptsource: An integrated development environment and repository for natural language prompts
Bach, S. H., Sanh, V., Yong, Z.-X., Webson, A., Raffel, C., Nayak, N. V., Sharma, A., Kim, T., Bari, M. S., Fevry, T., et al · 2022
Earlier work this paper cites.
Understanding dataset difficulty with mathcal v-usable information
Ethayarajh, K., Choi, Y., and Swayamdipta, S · 2022
Earlier work this paper cites.
Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned
Ganguli, D., Lovitt, L., Kernion, J., Askell, A., Bai, Y., Kadavath, S., Mann, B., Perez, E., Schiefer, N., Ndousse, K., et al · 2022
Earlier work this paper cites.
Large language models are zero-shot reasoners
Kojima, T., Gu, S. S., Reid, M., Matsuo, Y., and Iwasawa, Y · 2022
Earlier work this paper cites.
A new generation of perspective api: Efficient multilingual character-level transformers
Lees, A., Tran, V. Q., Tay, Y., Sorensen, J., Gupta, J., Metzler, D., and Vasserman, L · 2022
Earlier work this paper cites.
Introducing chatgpt
OpenAI · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al · 2022
Earlier work this paper cites.
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al · 2023
Earlier work this paper cites.
Like a good nearest neighbor: Practical content moderation and text classification
Bates, L. and Gurevych, I · 2023
Earlier work this paper cites.
Black-box prompt optimization: Aligning large language models without model training
Cheng, J., Liu, X., Zheng, K., Ke, P., Wang, H., Dong, Y., Tang, J., and Huang, M · 2023
Earlier work this paper cites.
Safe rlhf: Safe reinforcement learning from human feedback
Dai, J., Pan, X., Sun, R., Ji, J., Xu, X., Liu, M., Wang, Y., and Yang, Y · 2023
Earlier work this paper cites.
Improving factuality and reasoning in language models through multiagent debate
Du, Y., Li, S., Torralba, A., Tenenbaum, J. B., and Mordatch, I · 2023
Earlier work this paper cites.
Using punctuation as an adversarial attack on deep learning-based nlp systems: An empirical study
Formento, B., Foo, C. S., Tuan, L. A., and Ng, S. K · 2023
Earlier work this paper cites.
Think before you speak: Training language models with pause tokens
Goyal, S., Ji, Z., Rawat, A. S., Menon, A. K., Kumar, S., and Nagarajan, V · 2023
Earlier work this paper cites.
Llama guard: Llm-based input-output safeguard for human-ai conversations
Inan, H., Upasani, K., Chi, J., Rungta, R., Iyer, K., Mao, Y., Tontchev, M., Hu, Q., Fuller, B., Testuggine, D., et al · 2023
Cited alongside, same era.
Ai alignment: A comprehensive survey
Ji, J., Qiu, T., Chen, B., Zhang, B., Lou, H., Wang, K., Duan, Y., He, Z., Zhou, J., Zhang, Z., et al · 2023
Cited alongside, same era.
Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., Casas, D. d. l., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., et al · 2023
Cited alongside, same era.
Ke, P., Wen, B., Feng, Z., Liu, X., Lei, X., Cheng, J., Wang, S., Zeng, A., Dong, Y., Wang, H., et al · 2023
Cited alongside, same era.
Semrode: Macro adversarial training to learn representations that are robust to word-level attacks
Formento, B., Feng, W., Foo, C. S., Tuan, L. A., and Ng, S.-K · 2024
Later among the works it cites.
Aegis2. 0: A diverse ai safety dataset and risks taxonomy for alignment of llm guardrails
Ghosh, S., Varshney, P., Sreedhar, M. N., Padmakumar, A., Rebedea, T., Varghese, J. R., and Parisien, C · 2024
Later among the works it cites.
Deliberative alignment: Reasoning enables safer language models
Guan, M. Y., Joglekar, M., Wallace, E., Jain, S., Barak, B., Heylar, A., Dias, R., Vallone, A., Ren, H., Wei, J., et al · 2024
Later among the works it cites.
Cold-attack: Jailbreaking llms with stealthiness and controllability
Guo, X., Yu, F., Zhang, H., Qin, L., and Hu, B · 2024
Later among the works it cites.
Wildguard: Open one-stop moderation tools for safety risks, jailbreaks, and refusals of llms
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Pretraining language models with human preferences
Korbak, T., Shi, K., Chen, A., Bhalerao, R. V., Buckley, C., Phang, J., Bowman, S. R., and Perez, E · 2023
Cited alongside, same era.
Watch your language: large language models and content moderation
Kumar, D., AbuHashem, Y., and Durumeric, Z · 2023
Cited alongside, same era.
Efficient memory management for large language model serving with pagedattention
Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., Gonzalez, J., Zhang, H., and Stoica, I · 2023
Cited alongside, same era.
Encouraging divergent thinking in large language models through multi-agent debate
Liang, T., He, Z., Jiao, W., Wang, X., Wang, Y., Wang, R., Yang, Y., Shi, S., and Tu, Z · 2023
Cited alongside, same era.
Toxicchat: Unveiling hidden challenges of toxicity detection in real-world user-ai conversation
Lin, Z., Wang, Z., Tong, Y., Wang, Y., Guo, Y., Wang, Y., and Shang, J · 2023
Cited alongside, same era.
Inference-time policy adapters (ipa): Tailoring extreme-scale lms without fine-tuning
Lu, X., Brahman, F., West, P., Jang, J., Chandu, K., Ravichander, A., Qin, L., Ammanabrolu, P., Jiang, L., Ramnath, S., et al · 2023
Cited alongside, same era.
A holistic approach to undesired content detection in the real world
Markov, T., Zhang, C., Agarwal, S., Nekoul, F. E., Lee, T., Adler, S., Jiang, A., and Weng, L · 2023
Cited alongside, same era.
Nemo guardrails: A toolkit for controllable and safe llm applications with programmable rails
Rebedea, T., Dinu, R., Sreedhar, M., Parisien, C., and Cohen, J · 2023
Cited alongside, same era.
Han, S., Rao, K., Ettinger, A., Jiang, L., Lin, B. Y., Lambert, N., Choi, Y., and Dziri, N · 2024
Later among the works it cites.
Training large language models to reason in a continuous latent space
Hao, S., Sukhbaatar, S., Su, D., Li, X., Hu, Z., Weston, J., and Tian, Y · 2024
Later among the works it cites.
Qwen2. 5-coder technical report
Hui, B., Yang, J., Cui, Z., Yang, J., Liu, D., Zhang, L., Liu, T., Zhang, J., Yu, B., Dang, K., et al · 2024
Later among the works it cites.
Jiang, A. Q., Sablayrolles, A., Roux, A., Mensch, A., Savary, B., Bamford, C., Chaplot, D. S., Casas, D. d. l., Hanna, E. B., Bressand, F., et al · 2024
Later among the works it cites.
R2-guard: Robust reasoning enabled llm guardrail via knowledge-enhanced logical reasoning
Kang, M. and Li, B · 2024
Later among the works it cites.
Training language models to self-correct via reinforcement learning
Kumar, A., Zhuang, V., Agarwal, R., Su, Y., Co-Reyes, J. D., Singh, A., Baumli, K., Iqbal, S., Bishop, C., Roelofs, R., et al · 2024
Later among the works it cites.
Salad-bench: A hierarchical and comprehensive safety benchmark for large language models
Li, L., Dong, B., Wang, R., Hu, X., Zuo, W., Lin, D., Qiao, Y., and Shao, J · 2024
Later among the works it cites.
Harmbench: A standardized evaluation framework for automated red teaming and robust refusal
Mazeika, M., Phan, L., Yin, X., Zou, A., Wang, Z., Mu, N., Sakhaee, E., Li, N., Basart, S., Li, B., et al · 2024
Later among the works it cites.
Guardformer: Guardrail instruction pretraining for efficient safeguarding
O’Neill, J., Subramanian, S., Lin, E., Satish, A., and Mugunthan, V · 2024
Later among the works it cites.
Searchgpt prototype
OpenAI · 2024
Later among the works it cites.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R., Sharma, A., Mitchell, E., Manning, C. D., Ermon, S., and Finn, C · 2024
Later among the works it cites.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reid, M., Savinov, N., Teplyashin, D., Lepikhin, D., Lillicrap, T., Alayrac, J.-b., Soricut, R., Lazaridou, A., Firat, O., Schrittwieser, J., et al · 2024
Later among the works it cites.
Lightweight safety classification using pruned language models
Sawtell, M., Masterman, T., Besen, S., and Brown, J · 2024
Later among the works it cites.
Gemma 2: Improving open language models at a practical size
Team, G., Riviere, M., Pathak, S., Sessa, P. G., Hardin, C., Bhupatiraju, S., Hussenot, L., Mesnard, T., Shahriari, B., Ramé, A., et al · 2024
Later among the works it cites.
detoxify
UnitaryAI · 2024
Later among the works it cites.
Guardagent: Safeguard llm agents by a guard agent via knowledge-enabled reasoning
Xiang, Z., Zheng, L., Li, Y., Hong, J., Li, Q., Xie, H., Zhang, J., Xiong, Z., Xie, C., Yang, C., et al · 2024
Later among the works it cites.
Rigorllm: Resilient guardrails for large language models against undesired content
Yuan, Z., Xiong, Z., Zeng, Y., Yu, N., Jia, R., Song, D., and Li, B · 2024
Later among the works it cites.
Shieldgemma: Generative ai content moderation based on gemma
Zeng, W., Liu, Y., Mullins, R., Peran, L., Fernandez, J., Harkous, H., Narasimhan, K., Proud, D., Kumar, P., Radharapu, B., et al · 2024
Later among the works it cites.