Fetching the paper…
Reading the bibliography…
The robustness of large language models (LLMs) against adversarial manipulations, such as jailbreak attacks, remains a significant challenge.
Simple and scalable predictive uncertainty estimation using deep ensembles
Lakshminarayanan, B., Pritzel, A., and Blundell, C · 2017
Earlier work this paper cites.
Towards deep learning models resistant to adversarial attacks
Madry, A., Makelov, A., Schmidt, L., Tsipras, D., and Vladu, A · 2017
Earlier work this paper cites.
Loss surfaces, mode connectivity, and fast ensembling of dnns
Garipov, T., Izmailov, P., Podoprikhin, D., Vetrov, D. P., and Wilson, A. G · 2018
Earlier work this paper cites.
Averaging weights leads to wider optima and better generalization
Izmailov, P., Wilson, A., Podoprikhin, D., Vetrov, D., and Garipov, T · 2018
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2021
Earlier work this paper cites.
Constitutional ai: Harmlessness from ai feedback
Bai, Y., Kadavath, S., Kundu, S., Askell, A., Kernion, J., Jones, A., Chen, A., Goldie, A., Mirhoseini, A., McKinnon, C., et al · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback, 2022
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P., Leike, J., and Lowe, R · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al · 2022
Earlier work this paper cites.
Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time
Wortsman, M., Ilharco, G., Gadre, S. Y., Roelofs, R., Gontijo-Lopes, R., Morcos, A. S., Namkoong, H., Farhadi, A., Carmon, Y., Kornblith, S., et al · 2022
Earlier work this paper cites.
Open llm leaderboard
Beeching, E., Fourrier, C., Habib, N., Han, S., Lambert, N., Rajani, N., Sanseviero, O., Tunstall, L., and Wolf, T · 2023
Earlier work this paper cites.
It’s mbr all the way down: Modern generation techniques through the lens of minimum bayes risk
Bertsch, A., Xie, A., Neubig, G., and Gormley, M. R · 2023
Earlier work this paper cites.
Are aligned neural networks adversarially aligned?
Carlini, N., Nasr, M., Choquette-Choo, C. A., Jagielski, M., Gao, I., Awadalla, A., Koh, P. W., Ippolito, D., Lee, K., Tramer, F., et al · 2023
Earlier work this paper cites.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, March 2023
Chiang, W.-L., Li, Z., Lin, Z., Sheng, Y., Wu, Z., Zhang, H., Zheng, L., Zhuang, S., Zhuang, Y., Gonzalez, J. E., Stoica, I., and Xing, E. P · 2023
Earlier work this paper cites.
Rephrase and respond: Let large language models ask better questions for themselves
Deng, Y., Zhang, W., Chen, Z., and Gu, Q · 2023
Cited alongside, same era.
Toxicity in chatgpt: Analyzing persona-assigned language models
Deshpande, A., Murahari, V., Rajpurohit, T., Kalyan, A., and Narasimhan, K · 2023
Cited alongside, same era.
Mistral 7b, 2023
Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., de las Casas, D., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., Lavaud, L. R., Lachaux, M.-A., Stock, P., Scao, T. L., Lavril, T., Wang, T., Lacroix, T., and Sayed, W. E · 2023
Cited alongside, same era.
Automatically auditing large language models via discrete optimization
Jones, E., Dragan, A., Raghunathan, A., and Steinhardt, J · 2023
Cited alongside, same era.
Exploiting asymmetry for synthetic training data generation: SynthIE and the case of information extraction
Josifoski, M., Sakota, M., Peyrard, M., and West, R · 2023
Stanford alpaca: An instruction-following llama model
Taori, R., Gulrajani, I., Zhang, T., Dubois, Y., Li, X., Guestrin, C., Liang, P., and Hashimoto, T. B · 2023
Later among the works it cites.
Jailbroken: How does llm safety training fail?
Wei, A., Haghtalab, N., and Steinhardt, J · 2023
Later among the works it cites.
TIES-merging: Resolving interference when merging models
Yadav, P., Tam, D., Choshen, L., Raffel, C., and Bansal, M · 2023
Later among the works it cites.
Autodan: Automatic and interpretable adversarial attacks on large language models
Zhu, S., Zhang, R., An, B., Wu, G., Barrow, J., Wang, Z., Huang, F., Nenkova, A., and Sun, T · 2023
Later among the works it cites.
Universal and transferable adversarial attacks on aligned language models, 2023
Zou, A., Wang, Z., Kolter, J. Z., and Fredrikson, M · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Prometheus: Inducing fine-grained evaluation capability in language models
Kim, S., Shin, J., Cho, Y., Jang, J., Longpre, S., Lee, H., Yun, S., Shin, S., Kim, S., Thorne, J., et al · 2023
Cited alongside, same era.
Textbooks are all you need ii: phi-1.5
Li, Y., Bubeck, S., Eldan, R., Del Giorno, A., Gunasekar, S., and Lee, Y. T · 2023
Cited alongside, same era.
Self-refine: Iterative refinement with self-feedback
Madaan, A., Tandon, N., Gupta, P., Hallinan, S., Gao, L., Wiegreffe, S., Alon, U., Dziri, N., Prabhumoye, S., Yang, Y., et al · 2023
Cited alongside, same era.
Generative agents: Interactive simulacra of human behavior
Park, J. S., O’Brien, J., Cai, C. J., Morris, M. R., Liang, P., and Bernstein, M. S · 2023
Cited alongside, same era.
Hijacking large language models via adversarial in-context learning
Qiang, Y., Zhou, X., and Zhu, D · 2023
Cited alongside, same era.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C. D., and Finn, C · 2023
Cited alongside, same era.
Role play with large language models
Shanahan, M., McDonell, K., and Reynolds, L · 2023
Cited alongside, same era.
Bge m3-embedding: Multi-lingual, multi-functionality, multi-granularity text embeddings through self-knowledge distillation, 2024
Chen, J., Xiao, S., Zhang, P., Luo, K., Lian, D., and Liu, Z · 2024
Closest in time.
Distilled self-critique of LLMs with synthetic data: a bayesian perspective
Gallego, V · 2024
Closest in time.
Jiang, A. Q., Sablayrolles, A., Roux, A., Mensch, A., Savary, B., Bamford, C., Chaplot, D. S., Casas, D. d. l., Hanna, E. B., Bressand, F., et al · 2024
Closest in time.
Prometheus 2: An open source language model specialized in evaluating other language models
Kim, S., Suk, J., Longpre, S., Lin, B. Y., Shin, J., Welleck, S., Neubig, G., Lee, M., Lee, K., and Seo, M · 2024
Closest in time.
Meta llama guard 2
LlamaTeam · 2024
Closest in time.
Online merging optimizers for boosting rewards and mitigating tax in alignment, 2024
Lu, K., Yu, B., Huang, F., Fan, Y., Lin, R., and Zhou, C · 2024
Closest in time.
Defending llms against jailbreaking attacks via backtranslation
Wang, Y., Shi, Z., Bai, A., and Hsieh, C.-J · 2024
Closest in time.