Fetching the paper…
Reading the bibliography…
Large language models (LLMs) are expected to follow instructions from users and engage in conversations.
Build it break it fix it for dialogue safety: Robustness from adversarial human attack
Dinan, E.; Humeau, S.; Chintagunta, B.; and Weston, J. 2019 · 1908
Earlier work this paper cites.
Adversarial machine learning at scale
Kurakin, A.; Goodfellow, I.; and Bengio, S. 2016 · 2016
Earlier work this paper cites.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Wu, Y.; Schuster, M.; Chen, Z.; Le, Q. V.; Norouzi, M.; Macherey, W.; Krikun, M.; Cao, Y.; Gao, Q.; Macherey, K.; et al. 2016 · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Christiano, P. F.; Leike, J.; Brown, T.; Martic, M.; Legg, S.; and Amodei, D. 2017 · 2017
Earlier work this paper cites.
Human-Machine Collaboration for Content Regulation: The Case of Reddit Automoderator
Jhaver, S.; Birman, I.; Gilbert, E.; and Bruckman, A. 2019 · 2019
Earlier work this paper cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Bai, Y.; Jones, A.; Ndousse, K.; Askell, A.; Chen, A.; DasSarma, N.; Drain, D.; Fort, S.; Ganguli, D.; Henighan, T.; et al. 2022 · 2022
Earlier work this paper cites.
Fine-tuning language models to find agreement among humans with diverse preferences
Bakker, M.; Chadwick, M.; Sheahan, H.; Tessler, M.; Campbell-Gillingham, L.; Balaguer, J.; McAleese, N.; Glaese, A.; Aslanides, J.; Botvinick, M.; et al. 2022 · 2022
Earlier work this paper cites.
LoRA: Low-Rank Adaptation of Large Language Models
Hu, E. J.; yelong shen; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; and Chen, W. 2022 · 2022
Earlier work this paper cites.
A new generation of perspective API: Efficient multilingual character-level transformers
Lees, A.; Tran, V. Q.; Tay, Y.; Sorensen, J.; Gupta, J.; Metzler, D.; and Vasserman, L. 2022 · 2022
Earlier work this paper cites.
Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order Sensitivity
Lu, Y.; Bartolo, M.; Moore, A.; Riedel, S.; and Stenetorp, P. 2022 · 2022
Earlier work this paper cites.
Ignore previous prompt: Attack techniques for language models
Perez, F.; and Ribeiro, I. 2022 · 2022
Earlier work this paper cites.
Finetuned Language Models are Zero-Shot Learners
Wei, J.; Bosma, M.; Zhao, V.; Guu, K.; Yu, A. W.; Lester, B.; Du, N.; Dai, A. M.; and Le, Q. V. 2022 · 2022
Earlier work this paper cites.
Achiam, J.; Adler, S.; Agarwal, S.; Ahmad, L.; Akkaya, I.; Aleman, F. L.; Almeida, D.; Altenschmidt, J.; Altman, S.; Anadkat, S.; et al. 2023 · 2023
Earlier work this paper cites.
Introducing Claude 2.1
Anthropic. 2023 · 2023
Earlier work this paper cites.
Jailbreaking black box large language models in twenty queries
Chao, P.; Robey, A.; Dobriban, E.; Hassani, H.; Pappas, G. J.; and Wong, E. 2023 · 2023
Earlier work this paper cites.
Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90%* ChatGPT Quality
Chiang, W.-L.; Li, Z.; Lin, Z.; Sheng, Y.; Wu, Z.; Zhang, H.; Zheng, L.; Zhuang, S.; Zhuang, Y.; Gonzalez, J. E.; Stoica, I.; and Xing, E. P. 2023 · 2023
Cited alongside, same era.
Multilingual jailbreak challenges in large language models
Deng, Y.; Zhang, W.; Pan, S. J.; and Bing, L. 2023 · 2023
Cited alongside, same era.
MART: Improving LLM Safety with Multi-round Automatic Red-Teaming
Ge, S.; Zhou, C.; Hou, R.; Khabsa, M.; Wang, Y.-C.; Wang, Q.; Han, J.; and Mao, Y. 2023 · 2023
Cited alongside, same era.
Demystifying Prompts in Language Models via Perplexity Estimation
Gonen, H.; Iyer, S.; Blevins, T.; Smith, N.; and Zettlemoyer, L. 2023 · 2023
Cited alongside, same era.
Build with the Gemini API
Google. 2023 · 2023
Cited alongside, same era.
Autodan: Automatic and interpretable adversarial attacks on large language models
Zhu, S.; Zhang, R.; An, B.; Wu, G.; Barrow, J.; Wang, Z.; Huang, F.; Nenkova, A.; and Sun, T. 2023 · 2023
Later among the works it cites.
Universal and transferable adversarial attacks on aligned language models
Zou, A.; Wang, Z.; Kolter, J. Z.; and Fredrikson, M. 2023 · 2023
Later among the works it cites.
Jailbreaking leading safety-aligned llms with simple adaptive attacks
Andriushchenko, M.; Croce, F.; and Flammarion, N. 2024 · 2024
Closest in time.
Introducing the next generation of Claude
Anthropic. 2024 · 2024
Closest in time.
Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Chiang, W.-L.; Zheng, L.; Sheng, Y.; Angelopoulos, A. N.; Li, T.; Li, D.; Zhang, H.; Zhu, B.; Jordan, M.; Gonzalez, J. E.; and Stoica, I. 2024 · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Huang, Y.; Gupta, S.; Xia, M.; Li, K.; and Chen, D. 2023 · 2023
Cited alongside, same era.
Jiang, A. Q.; Sablayrolles, A.; Mensch, A.; Bamford, C.; Chaplot, D. S.; Casas, D. d. l.; Bressand, F.; Lengyel, G.; Lample, G.; Saulnier, L.; et al. 2023 · 2023
Cited alongside, same era.
Automatically Auditing Large Language Models via Discrete Optimization
Jones, E.; Dragan, A.; Raghunathan, A.; and Steinhardt, J. 2023 · 2023
Cited alongside, same era.
Models-OpenAI API
OpenAI. 2023 · 2023
Cited alongside, same era.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H.; Martin, L.; Stone, K.; Albert, P.; Almahairi, A.; Babaei, Y.; Bashlykov, N.; Batra, S.; Bhargava, P.; Bhosale, S.; et al. 2023 · 2023
Cited alongside, same era.
Bypassing the Safety Training of Open-Source LLMs with Priming Attacks
Vega, J.; Chaudhary, I.; Xu, C.; and Singh, G. 2023 · 2023
Cited alongside, same era.
Jailbroken: How does LLM safety training fail?
Wei, A.; Haghtalab, N.; and Steinhardt, J. 2023 · 2023
Cited alongside, same era.
Closest in time.
Templates for Chat Models
HuggingFace. 2024 · 2024
Closest in time.
Single character perturbations break llm alignment
Lin, L.; Brown, H.; Kawaguchi, K.; and Shieh, M. 2024 · 2024
Closest in time.
Meta Llama Guard 2
Llama-Team. 2024 · 2024
Closest in time.
Harmbench: A standardized evaluation framework for automated red teaming and robust refusal
Mazeika, M.; Phan, L.; Yin, X.; Zou, A.; Wang, Z.; Mu, N.; Sakhaee, E.; Li, N.; Basart, S.; Li, B.; et al. 2024 · 2024
Closest in time.
Advprompter: Fast adaptive adversarial prompting for LLMs
Paulus, A.; Zharmagambetov, A.; Guo, C.; Amos, B.; and Tian, Y. 2024 · 2024
Closest in time.
Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Qi, X.; Zeng, Y.; Xie, T.; Chen, P.-Y.; Jia, R.; Mittal, P.; and Henderson, P. 2024 · 2024
Closest in time.
Quantifying Language Models’ Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Sclar, M.; Choi, Y.; Tsvetkov, Y.; and Suhr, A. 2024 · 2024
Closest in time.
BadChain: Backdoor Chain-of-Thought Prompting for Large Language Models
Xiang, Z.; Jiang, F.; Xiong, Z.; Ramasubramanian, B.; Poovendran, R.; and Li, B. 2024 · 2024
Closest in time.
SafeDecoding: Defending against Jailbreak Attacks via Safety-Aware Decoding
Xu, Z.; Jiang, F.; Niu, L.; Jia, J.; Lin, B. Y.; and Poovendran, R. 2024 · 2024
Closest in time.
Zeng, Y.; Lin, H.; Zhang, J.; Yang, D.; Jia, R.; and Shi, W. 2024 · 2024
Closest in time.