Fetching the paper…
Reading the bibliography…
The emergent capabilities of Large Language Models (LLMs) have made it crucial to align their values with those of humans.
Language models are few-shot learners
Brown, T.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J. D.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al. 2020 · 1901
Earlier work this paper cites.
A behavioral model of rational choice
Simon, H. A. 1955 · 1955
Earlier work this paper cites.
Job satisfaction and job performance: A theoretical analysis
Locke, E. A. 1970 · 1970
Earlier work this paper cites.
Social values and rules of fairness: A theoretical perspective
McClintock, C. G.; and Van Avermaet, E. 1982 · 1982
Earlier work this paper cites.
Social diversity and social preferences in mixed-motive reinforcement learning
McKee, K. R.; Gemp, I.; McWilliams, B.; Duéñez-Guzmán, E. A.; Hughes, E.; and Leibo, J. Z. 2020 · 2002
Earlier work this paper cites.
Measuring social value orientation
Murphy, R. O.; Ackermann, K. A.; and Handgraaf, M. J. 2011 · 2011
Earlier work this paper cites.
The moral machine experiment
Awad, E.; Dsouza, S.; Kim, R.; Schulz, J.; Henrich, J.; Shariff, A.; Bonnefon, J.-F.; and Rahwan, I. 2018 · 2018
Earlier work this paper cites.
Reinforcement learning: An introduction
Sutton, R. S.; and Barto, A. G. 2018 · 2018
Earlier work this paper cites.
Social behavior for autonomous vehicles
Schwarting, W.; Pierson, A.; Alonso-Mora, J.; Karaman, S.; and Rus, D. 2019 · 2019
Earlier work this paper cites.
Virtuous vs. utilitarian artificial moral agents
Bauer, W. A. 2020 · 2020
Earlier work this paper cites.
A general language assistant as a laboratory for alignment
Askell, A.; Bai, Y.; Chen, A.; Drain, D.; Ganguli, D.; Henighan, T.; Jones, A.; Joseph, N.; Mann, B.; DasSarma, N.; et al. 2021 · 2021
Earlier work this paper cites.
Value alignment verification
Brown, D. S.; Schneider, J.; Dragan, A.; and Niekum, S. 2021 · 2021
Cited alongside, same era.
Evaluating large language models trained on code
Chen, M.; Tworek, J.; Jun, H.; Yuan, Q.; Pinto, H. P. d. O.; Kaplan, J.; Edwards, H.; Burda, Y.; Joseph, N.; Brockman, G.; et al. 2021 · 2021
Cited alongside, same era.
Constitutional AI: Harmlessness from AI Feedback
Bai, Y.; Kadavath, S.; Kundu, S.; Askell, A.; Kernion, J.; Jones, A.; Chen, A.; Goldie, A.; Mirhoseini, A.; McKinnon, C.; et al. 2022 · 2022
Cited alongside, same era.
A virtue-based framework to support putting AI ethics into practice
Hagendorff, T. 2022 · 2022
Cited alongside, same era.
MPI: Evaluating and Inducing Personality in Pre-trained Language Models
Jiang, G.; Xu, M.; Zhu, S.-C.; Han, W.; Zhang, C.; and Zhu, Y. 2022 · 2022
Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90%* ChatGPT Quality
Chiang, W.-L.; Li, Z.; Lin, Z.; Sheng, Y.; Wu, Z.; Zhang, H.; Zheng, L.; Zhuang, S.; Zhuang, Y.; Gonzalez, J. E.; Stoica, I.; and Xing, E. P. 2023 · 2023
Closest in time.
Koala: A Dialogue Model for Academic Research
Geng, X.; Gudibande, A.; Liu, H.; Wallace, E.; Abbeel, P.; Levine, S.; and Song, D. 2023 · 2023
Closest in time.
AI Principles
of Life Institute, F. 2023 · 2023
Closest in time.
Introducing ChatGPT
OpenAI. 2023b · 2023
Closest in time.
Model index for researchers
OpenAI. 2023c · 2023
Closest in time.
Selenium
SeleniumHQ. 2023 · 2023
Closest in time.
Stanford Alpaca: An Instruction-following LLaMA model
Taori, R.; Gulrajani, I.; Zhang, T.; Dubois, Y.; Li, X.; Guestrin, C.; Liang, P.; and Hashimoto, T. B. 2023 · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L.; Wu, J.; Jiang, X.; Almeida, D.; Wainwright, C.; Mishkin, P.; Zhang, C.; Agarwal, S.; Slama, K.; Ray, A.; et al. 2022 · 2022
Cited alongside, same era.
Self-Instruct: Aligning Language Model with Self Generated Instructions
Wang, Y.; Kordi, Y.; Mishra, S.; Liu, A.; Smith, N. A.; Khashabi, D.; and Hajishirzi, H. 2022 · 2022
Cited alongside, same era.
In situ bidirectional human-robot value alignment
Yuan, L.; Gao, X.; Zheng, Z.; Edmonds, M.; Wu, Y. N.; Rossano, F.; Lu, H.; Zhu, Y.; and Zhu, S.-C. 2022 · 2022
Cited alongside, same era.
Introducing 100K Context Windows
Anthropic. 2023a · 2023
Cited alongside, same era.
Meet Claude
Anthropic. 2023b · 2023
Cited alongside, same era.
Sparks of artificial general intelligence: Early experiments with gpt-4
Bubeck, S.; Chandrasekaran, V.; Eldan, R.; Gehrke, J.; Horvitz, E.; Kamar, E.; Lee, P.; Lee, Y. T.; Li, Y.; Lundberg, S.; et al. 2023 · 2023
Cited alongside, same era.
OpenAI. 2023a
Cited in the paper.
Closest in time.
Chat with Open Large Language Models
The Vicuna Team. 2023 · 2023
Closest in time.
LLaMA: Open and efficient foundation language models
Touvron, H.; Lavril, T.; Izacard, G.; Martinet, X.; Lachaux, M.-A.; Lacroix, T.; Rozière, B.; Goyal, N.; Hambro, E.; Azhar, F.; et al. 2023 · 2023
Closest in time.
Using the Veil of Ignorance to align AI systems with principles of justice
Weidinger, L.; McKee, K. R.; Everett, R.; Huang, S.; Zhu, T. O.; Chadwick, M. J.; Summerfield, C.; and Gabriel, I. 2023 · 2023
Closest in time.