Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) are trained on massive text corpora, which are encoded with diverse personality traits.
On the problem of the most efficient tests of statistical hypotheses
Neyman, J. and Pearson, E. S · 1933
Earlier work this paper cites.
Adversarial training for large neural language models
Liu, X., Cheng, H., He, P., Chen, W., Wang, Y., Poon, H., and Gao, J · 2020
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Hu, E. J., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., Chen, W., et al · 2021
Earlier work this paper cites.
Mesh-Transformer-JAX: Model-Parallel Implementation of Transformer Language Model with JAX
Wang, B · 2021
Earlier work this paper cites.
An explanation of in-context learning as implicit bayesian inference
Xie, S. M., Raghunathan, A., Liang, P., and Ma, T · 2021
Earlier work this paper cites.
Prompt injection: Parameterization of fixed inputs
Choi, E., Jo, Y., Jang, J., and Seo, M · 2022
Earlier work this paper cites.
Scaling instruction-finetuned language models
Chung, H. W., Hou, L., Longpre, S., Zoph, B., Tay, Y., Fedus, W., Li, Y., Wang, X., Dehghani, M., Brahma, S., et al · 2022
Earlier work this paper cites.
Improving alignment of dialogue agents via targeted human judgements
Glaese, A., McAleese, N., Trebacz, M., Aslanides, J., Firoiu, V., Ewalds, T., Rauh, M., Weidinger, L., Chadwick, M., Thacker, P., et al · 2022
Earlier work this paper cites.
Selective annotation makes language models better few-shot learners
Hongjin, S., Kasai, J., Wu, C. H., Shi, W., Wang, T., Xin, J., Zhang, R., Ostendorf, M., Zettlemoyer, L., Smith, N. A., et al · 2022
Earlier work this paper cites.
Self-generated in-context learning: Leveraging auto-regressive language models as a demonstration generator
Kim, H. J., Cho, H., Kim, J., Kim, T., Yoo, K. M., and Lee, S.-g · 2022
Earlier work this paper cites.
What makes good in-context examples for gpt-3?
Liu, J., Shen, D., Zhang, Y., Dolan, W. B., Carin, L., and Chen, W · 2022
Earlier work this paper cites.
Teaching language models to support answers with verified quotes
Menick, J., Trebacz, M., Mikulik, V., Aslanides, J., Song, F., Chadwick, M., Glaese, M., Young, S., Campbell-Gillingham, L., Irving, G., et al · 2022
Earlier work this paper cites.
Memory-based model editing at scale
Mitchell, E., Lin, C., Bosselut, A., Manning, C. D., and Finn, C · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Earlier work this paper cites.
Discovering language model behaviors with model-written evaluations
Perez, E., Ringer, S., Lukošiūtė, K., Nguyen, K., Chen, E., Heiner, S., Pettit, C., Olsson, C., Kundu, S., Kadavath, S., Jones, A., Chen, A., Mann, B., Israel, B., Seethor, B., McKinnon, C., Olah, C., Yan, D., Amodei, D., Amodei, D., Drain, D., Li, D., Tran-Johnson, E., Khundadze, G., Kernion, J., Landis, J., Kerr, J., Mueller, J., Hyun, J., Landau, J., Ndousse, K., Goldberg, L., Lovitt, L., Lucas, M., Sellitto, M., Zhang, M., Kingsland, N., Elhage, N., Joseph, N., Mercado, N., DasSarma, N., Rausch, O., Larson, R., McCandlish, S., Johnston, S., Kravec, S., El Showk, S., Lanham, T., Telleen-Lawton, T., Brown, T., Henighan, T., Hume, T., Bai, Y., Hatfield-Dodds, Z., Clark, J., Bowman, S. R., Askell, A., Grosse, R., Hernandez, D., Ganguli, D., Hubinger, E., Schiefer, N., and Kaplan, J · 2022
Earlier work this paper cites.
Learning to retrieve prompts for in-context learning
Rubin, O., Herzig, J., and Berant, J · 2022
Cited alongside, same era.
Finetuned language models are zero-shot learners
Wei, J., Bosma, M., Zhao, V., Guu, K., Yu, A. W., Lester, B., Du, N., Dai, A. M., and Le, Q. V · 2022
Cited alongside, same era.
Opt: Open pre-trained transformer language models
Zhang, S., Roller, S., Goyal, N., Artetxe, M., Chen, M., Chen, S., Dewan, C., Diab, M., Li, X., Lin, X. V., et al · 2022
Cited alongside, same era.
Direct preference-based policy optimization without reward modeling
An, G., Lee, J., Zuo, X., Kosaka, N., Kim, K.-M., and Song, H. O · 2023
Cited alongside, same era.
Jailbreaking language models at scale via persona modulation
Anonymous · 2023
Cited alongside, same era.
Marked personas: Using natural language prompts to measure stereotypes in language models
OpenAI · 2023
Later among the works it cites.
Do llms possess a personality? making the mbti test an amazing evaluation for large language models
Pan, K. and Zeng, Y · 2023
Later among the works it cites.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C. D., and Finn, C · 2023
Later among the works it cites.
Personality traits in large language models
Safdari, M., Serapio-García, G., Crepy, C., Fitz, S., Romero, P., Sun, L., Abdulhai, M., Faust, A., and Matarić, M · 2023
Later among the works it cites.
In-context impersonation reveals large language models’ strengths and biases
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cheng, M., Durmus, E., and Jurafsky, D · 2023
Cited alongside, same era.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, March 2023
Chiang, W.-L., Li, Z., Lin, Z., Sheng, Y., Wu, Z., Zhang, H., Zheng, L., Zhuang, S., Zhuang, Y., Gonzalez, J. E., Stoica, I., and Xing, E. P · 2023
Cited alongside, same era.
Language instructed reinforcement learning for human-ai coordination
Hu, H. and Sadigh, D · 2023
Cited alongside, same era.
Evaluating and inducing personality in pre-trained language models
Jiang, G., Xu, M., Zhu, S.-C., Han, W., Zhang, C., and Zhu, Y · 2023
Cited alongside, same era.
Pretraining language models with human preferences
Korbak, T., Shi, K., Chen, A., Bhalerao, R. V., Buckley, C., Phang, J., Bowman, S. R., and Perez, E · 2023
Cited alongside, same era.
Kossen, J., Rainforth, T., and Gal, Y · 2023
Cited alongside, same era.
Unified demonstration retriever for in-context learning
Li, X., Lv, K., Yan, H., Lin, T., Zhu, W., Ni, Y., Xie, G., Wang, X., and Qiu, X · 2023
Cited alongside, same era.
Salewski, L., Alaniz, S., Rio-Torto, I., Schulz, E., and Akata, Z · 2023
Later among the works it cites.
Evaluating the moral beliefs encoded in llms
Scherrer, N., Shi, C., Feder, A., and Blei, D · 2023
Later among the works it cites.
Shen, X., Chen, Z., Backes, M., Shen, Y., and Zhang, Y · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Later among the works it cites.
Adversarial demonstration attacks on large language models
Wang, J., Liu, Z., Park, K. H., Chen, M., and Xiao, C · 2023
Later among the works it cites.
Fundamental limitations of alignment in large language models
Wolf, Y., Wies, N., Levine, Y., and Shashua, A · 2023
Later among the works it cites.
Small models are valuable plug-ins for large language models
Xu, C., Xu, Y., Wang, S., Liu, Y., Zhu, C., and McAuley, J · 2023
Later among the works it cites.
Shadow alignment: The ease of subverting safely-aligned language models
Yang, X., Wang, X., Zhang, Q., Petzold, L., Wang, W. Y., Zhao, X., and Lin, D · 2023
Later among the works it cites.
Universal and transferable adversarial attacks on aligned language models
Zou, A., Wang, Z., Kolter, J. Z., and Fredrikson, M · 2023
Later among the works it cites.
Args: Alignment as reward-guided search
Khanov, M., Burapacheep, J., and Li, Y · 2024
Closest in time.