Fetching the paper…
Reading the bibliography…
Aligning large language models (LLMs) with human values is imperative to mitigate potential adverse effects resulting from their misuse.
Symbolic interaction
Hall, P. M · 2007
Earlier work this paper cites.
Critical argument and writer identity: Social constructivism as a theoretical framework for efl academic writing
McKinley, J · 2015
Earlier work this paper cites.
A general language assistant as a laboratory for alignment
Askell, A., Bai, Y., Chen, A., Drain, D., Ganguli, D., Henighan, T., Jones, A., Joseph, N., Mann, B., DasSarma, N., et al · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2021
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Earlier work this paper cites.
Bbq: A hand-built bias benchmark for question answering
Parrish, A., Chen, A., Nangia, N., Padmakumar, V., Phang, J., Thompson, J., Htut, P. M., and Bowman, S · 2022
Earlier work this paper cites.
Identifying and mitigating the security risks of generative ai
Barrett, C., Boyd, B., Bursztein, E., Carlini, N., Chen, B., Choi, J., Chowdhury, A. R., Christodorescu, M., Datta, A., Feizi, S., et al · 2023
Earlier work this paper cites.
Red-teaming large language models using chain of utterances for safety-alignment
Bhardwaj, R. and Poria, S · 2023
Earlier work this paper cites.
Weak-to-strong generalization: Eliciting strong capabilities with weak supervision
Burns, C., Izmailov, P., Kirchner, J. H., Baker, B., Gao, L., Aschenbrenner, L., Chen, Y., Ecoffet, A., Joglekar, M., Leike, J., et al · 2023
Earlier work this paper cites.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
Chiang, W.-L., Li, Z., Lin, Z., Sheng, Y., Wu, Z., Zhang, H., Zheng, L., Zhuang, S., Zhuang, Y., Gonzalez, J. E., et al · 2023
Earlier work this paper cites.
Toxicity in chatgpt: Analyzing persona-assigned language models
Deshpande, A., Murahari, V., Rajpurohit, T., Kalyan, A., and Narasimhan, K · 2023
Earlier work this paper cites.
Qlora: Efficient finetuning of quantized llms
Dettmers, T., Pagnoni, A., Holtzman, A., and Zettlemoyer, L · 2023
Earlier work this paper cites.
Improving factuality and reasoning in language models through multiagent debate
Du, Y., Li, S., Torralba, A., Tenenbaum, J. B., and Mordatch, I · 2023
Earlier work this paper cites.
Large language models empowered agent-based modeling and simulation: A survey and perspectives
Gao, C., Lan, X., Li, N., Yuan, Y., Ding, J., Zhou, Z., Xu, F., and Li, Y · 2023
Earlier work this paper cites.
Large language models can be used to effectively scale spear phishing campaigns
Hazell, J · 2023
Earlier work this paper cites.
War and peace (waragent): Large language model-based multi-agent simulation of world wars
Hua, W., Fan, L., Li, L., Mei, K., Ji, J., Ge, Y., Hemphill, L., and Zhang, Y · 2023
Earlier work this paper cites.
Beavertails: Towards improved safety alignment of llm via a human-preference dataset
Ji, J., Liu, M., Dai, J., Pan, X., Zhang, C., Bian, C., Chen, B., Sun, R., Wang, Y., and Yang, Y · 2023
Earlier work this paper cites.
Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., Casas, D. d. l., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., et al · 2023
Earlier work this paper cites.
Exploiting programmatic behavior of llms: Dual-use through standard security attacks
Kang, D., Li, X., Stoica, I., Guestrin, C., Zaharia, M., and Hashimoto, T · 2023
Cited alongside, same era.
Rlaif: Scaling reinforcement learning from human feedback with ai feedback
Lee, H., Phatale, S., Mansoor, H., Lu, K., Mesnard, T., Bishop, C., Carbune, V., and Rastogi, A · 2023
Cited alongside, same era.
Camel: Communicative agents for” mind” exploration of large language model society
Li, G., Hammoud, H., Itani, H., Khizbullin, D., and Ghanem, B · 2023
Cited alongside, same era.
Bolaa: Benchmarking and orchestrating llm-augmented autonomous agents
Liu, Z., Yao, W., Zhang, J., Xue, L., Heinecke, S., Murthy, R., Feng, Y., Chen, Z., Niebles, J. C., Arpit, D., et al · 2023
Cited alongside, same era.
Our approach to ai safety
OpenAI · 2023
Sotopia: Interactive evaluation for social intelligence in language agents
Zhou, X., Zhu, H., Mathur, L., Zhang, R., Yu, H., Qi, Z., Morency, L.-P., Bisk, Y., Fried, D., Neubig, G., et al · 2023
Later among the works it cites.
Universal and transferable adversarial attacks on aligned language models, 2023
Zou, A., Wang, Z., Kolter, J. Z., and Fredrikson, M · 2023
Later among the works it cites.
URL https://en.wikipedia.org/wiki/Social_constructivism
Social constructivism, 2024 · 2024
Closest in time.
URL https://en.wikipedia.org/wiki/Symbolic_interactionism
Symbolic interactionism, 2024 · 2024
Closest in time.
Foundational challenges in assuring alignment and safety of large language models
Anwar, U., Saparov, A., Rando, J., Paleka, D., Turpin, M., Hase, P., Lubana, E. S., Jenner, E., Casper, S., Sourbut, O., et al · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
OpenAI · 2023
Cited alongside, same era.
Openassistant/reward-model-deberta-v3-large-v2
OpenAssistant · 2023
Cited alongside, same era.
Generative agents: Interactive simulacra of human behavior
Park, J. S., O’Brien, J., Cai, C. J., Morris, M. R., Liang, P., and Bernstein, M. S · 2023
Cited alongside, same era.
Communicative agents for software development
Qian, C., Cong, X., Yang, C., Chen, W., Su, Y., Xu, J., Liu, Z., and Sun, M · 2023
Cited alongside, same era.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R., Sharma, A., Mitchell, E., Manning, C. D., Ermon, S., and Finn, C · 2023
Cited alongside, same era.
In-context impersonation reveals large language models’ strengths and biases
Salewski, L., Alaniz, S., Rio-Torto, I., Schulz, E., and Akata, Z · 2023
Cited alongside, same era.
Principle-driven self-alignment of language models from scratch with minimal human supervision
Sun, Z., Shen, Y., Zhou, Q., Zhang, H., Chen, Z., Cox, D. D., Yang, Y., and Gan, C · 2023
Cited alongside, same era.
Bengio, Y., Hinton, G., Yao, A., Song, D., Abbeel, P., Darrell, T., Harari, Y. N., Zhang, Y.-Q., Xue, L., Shalev-Shwartz, S., et al · 2024
Closest in time.
Wizard-vicuna-30b-uncensored
Cognitivecomputations · 2024
Closest in time.
Bias runs deep: Implicit reasoning biases in persona-assigned llms
Gupta, S., Shrivastava, V., Deshpande, A., Kalyan, A., Clark, P., Sabharwal, A., and Khot, T · 2024
Closest in time.
Metagpt: Meta programming for multi-agent collaborative framework
Hong, S., Zhuge, M., Chen, J., Zheng, X., Cheng, Y., Wang, J., Zhang, C., Wang, Z., Yau, S. K. S., Lin, Z., et al · 2024
Closest in time.
Alignment as reward-guided search
Khanov, M., Burapacheep, J., and Li, Y · 2024
Closest in time.
Rain: Your language models can align themselves without finetuning
Li, Y., Wei, F., Zhao, J., Zhang, C., and Zhang, H · 2024
Closest in time.
Urial: Aligning untuned llms with just the’write’amount of in-context learning
Lin, B. Y., Ravichander, A., Lu, X., Dziri, N., Sclar, M., Chandu, K., Bhagavatula, C., and Choi, Y · 2024
Closest in time.
Training socially aligned language models on simulated social interactions
Liu, R., Yang, R., Jia, C., Zhang, G., Yang, D., and Vosoughi, S · 2024
Closest in time.
Salmon: Self-alignment with principle-following reward models
Sun, Z., Shen, Y., Zhang, H., Zhou, Q., Chen, Z., Cox, D. D., Yang, Y., and Gan, C · 2024
Closest in time.
Wizardlm: Empowering large pre-trained language models to follow complex instructions
Xu, C., Sun, Q., Zheng, K., Geng, X., Zhao, P., Feng, J., Tao, C., Lin, Q., and Jiang, D · 2024
Closest in time.
Rlcd: Reinforcement learning from contrastive distillation for lm alignment
Yang, K., Klein, D., Celikyilmaz, A., Peng, N., and Tian, Y · 2024
Closest in time.
Openfedllm: Training large language models on decentralized private data via federated learning
Ye, R., Wang, W., Chai, J., Li, D., Li, Z., Xu, Y., Du, Y., Wang, Y., and Chen, S · 2024
Closest in time.
Open-source can be dangerous: On the vulnerability of value alignment in open-source LLMs, 2024
Yi, J., Ye, R., Chen, Q., Zhu, B. B., Chen, S., Lian, D., Sun, G., Xie, X., and Wu, F · 2024
Closest in time.