Fetching the paper…
Reading the bibliography…
Increasing interest in ensuring the safety of next-generation Artificial Intelligence (AI) systems calls for novel approaches to embedding morality into autonomous agents.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D. (2020b) · 1901
Earlier work this paper cites.
Fine-tuning language models from human preferences
Ziegler, D. M., Stiennon, N., Wu, J., Brown, T. B., Radford, A., Amodei, D., Christiano, P., and Irving, G. (2019) · 1909
Earlier work this paper cites.
Variabilit(à) e Mutabilit(à): Contributo allo studio delle distribuzioni e delle relazioni statistiche. [Fasc. I.]
Gini, C. (1912) · 1912
Earlier work this paper cites.
I, Robot
Asimov, I. (1950) · 1950
Earlier work this paper cites.
On the rationality postulates underlying the theory of cooperative games
Harsanyi, J. C. (1961) · 1961
Earlier work this paper cites.
The tragedy of the commons. the population problem has no technical solution; it requires a fundamental extension in morality
Hardin, G. (1968) · 1968
Earlier work this paper cites.
A Theory of Justice
Rawls, J. (1971) · 1971
Earlier work this paper cites.
Prisoner’s dilemma — recollections and observations
Rapoport, A. (1974) · 1974
Earlier work this paper cites.
The cognitive-developmental approach to moral education
Kohlberg, L. (1975) · 1975
Earlier work this paper cites.
Killing, letting die, and the trolley problem
Thomson, J. J. (1976) · 1976
Earlier work this paper cites.
The evolution of cooperation
Axelrod, R., and Hamilton, W. D. (1981) · 1981
Earlier work this paper cites.
Prolog: A step towards the ultimate computer language
Ferguson, R. (1981) · 1981
Earlier work this paper cites.
In a Different Voice: Psychological Theory and Women’s Development
Gilligan, C. (1982) · 1982
Earlier work this paper cites.
A classification of social dilemma games
Liebrand, W. B. (1983) · 1983
Earlier work this paper cites.
Intrinsic Motivation and Self-determination in Human Behavior
Deci, E. L., and Ryan, R. M. (1985) · 1985
Earlier work this paper cites.
Morals by Agreement
Gauthier, D. (1987) · 1987
Earlier work this paper cites.
Artificial Morality: Virtuous Robots for Virtual Games
Danielson, P. (1992) · 1992
Earlier work this paper cites.
Q-learning
Watkins, C. J., and Dayan, P. (1992) · 1992
Earlier work this paper cites.
Asimov’s laws of robotics: implications for information technology - part i
Clarke, R. (1993) · 1993
Earlier work this paper cites.
Prisoner’s Dilemma: John von Neumann, Game Theory, and the Puzzle of the Bomb
Poundstone, W. (1993) · 1993
Earlier work this paper cites.
Markov games as a framework for multi-agent reinforcement learning
Littman, M. L. (1994) · 1994
Earlier work this paper cites.
Games with incomplete information
Harsanyi, J. C. (1995) · 1995
Earlier work this paper cites.
Public goods: a survey of experimental research
Ledyard, J. O. (1995) · 1995
Earlier work this paper cites.
Multiagent reinforcement learning in the iterated prisoner’s dilemma
Sandholm, T. W., and Crites, R. H. (1996) · 1996
Earlier work this paper cites.
Jobs, careers, and callings: people’s relations to their work
Wrzesniewski, A., McCauley, C., Rozin, P., and Schwartz, B. (1997) · 1997
Earlier work this paper cites.
Modeling altruism and spitefulness in experiments
Levine, D. K. (1998) · 1998
Earlier work this paper cites.
A meta-analytic review of experiments examining the effects of extrinsic rewards on intrinsic motivation
Deci, E. L., Koestner, R., and Ryan, R. M. (1999) · 1999
Earlier work this paper cites.
A theory of fairness, competition, and cooperation
Fehr, E., and Schmidt, K. M. (1999) · 1999
Earlier work this paper cites.
ERC: A theory of equity, reciprocity, and competition
Bolton, G. E., and Ockenfels, A. (2000) · 2000
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
Ng, A. Y., and Russell, S. J. (2000) · 2000
Earlier work this paper cites.
The stag hunt
Skyrms, B. (2001) · 2001
Earlier work this paper cites.
Giving according to GARP: An experimental test of the consistency of preferences for altruism
Andreoni, J., and Miller, J. (2002) · 2002
Earlier work this paper cites.
Understanding social preferences with simple tests
Charness, G., and Rabin, M. (2002) · 2002
Earlier work this paper cites.
Why social preferences matter – the impact of non-selfish motives on competition, cooperation and incentives
Fehr, E., and Fischbacher, U. (2002) · 2002
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
Abbeel, P., and Ng, A. Y. (2004) · 2004
Earlier work this paper cites.
Intrinsically motivated reinforcement learning
Chentanez, N., Barto, A., and Singh, S. (2004) · 2004
Earlier work this paper cites.
Toward ethical robots via mechanized deontic logic
Arkoudas, K., Bringsjord, S., and Bello, P. (2005) · 2005
Earlier work this paper cites.
The Grammar of Society: The Nature and Dynamics of Social Norms
Bicchieri, C. (2005) · 2005
Earlier work this paper cites.
Natural Justice
Binmore, K. (2005) · 2005
Earlier work this paper cites.
MedEthEx: a prototype medical ethics advisor
Anderson, M., Anderson, S. L., and Armen, C. (2006) · 2006
Earlier work this paper cites.
Toward a general logicist methodology for engineering ethically correct robots
Bringsjord, S., Arkoudas, K., and Bello, P. (2006) · 2006
Earlier work this paper cites.
Particularism and the classification and reclassification of moral cases
Guarini, M. (2006) · 2006
Earlier work this paper cites.
What do laboratory experiments measuring social preferences reveal about the real world?
Levitt, S. D., and List, J. A. (2007) · 2007
Earlier work this paper cites.
What is intrinsic motivation? A typology of computational approaches
Oudeyer, P.-Y., and Kaplan, F. (2007) · 2007
Earlier work this paper cites.
Core knowledge
Spelke, E. S., and Kinzler, K. D. (2007) · 2007
Earlier work this paper cites.
Machine morality: bottom-up and top-down approaches for modelling human moral faculties
Wallach, W., Allen, C., and Smit, I. (2008) · 2008
Earlier work this paper cites.
The Philosophical Baby: What Children’s Minds Tell Us About Truth, Love and the Meaning of Life
Gopnik, A. (2009) · 2009
Earlier work this paper cites.
Liberals and conservatives rely on different sets of moral foundations.
Graham, J., Haidt, J., and Nosek, B. A. (2009) · 2009
Earlier work this paper cites.
Moral Machines: Teaching Robots Right from Wrong
Wallach, W., and Allen, C. (2009) · 2009
Earlier work this paper cites.
Robot minds and human ethics: the need for a comprehensive model of moral decision making
Wallach, W. (2010) · 2010
Cited alongside, same era.
Behavioral Game Theory: Experiments in Strategic Interaction
Camerer, C. (2011) · 2011
Cited alongside, same era.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E. (2012) · 2012
Cited alongside, same era.
Moral Foundations Theory: the pragmatic validity of moral pluralism
Graham, J., Haidt, J., Koleva, S., Motyl, M., Iyer, R., Wojcik, S. P., and Ditto, P. H. (2013) · 2013
Cited alongside, same era.
Identifying social norms using coordination games: why does dictator game sharing vary?
Krupka, E. L., and Weber, R. A. (2013) · 2013
Cited alongside, same era.
The ethics of artificial intelligence
Bostrom, N., and Yudkowsky, E. (2014) · 2014
Mathematical foundations of moral preferences
Capraro, V., and Perc, M. (2021) · 2021
Later among the works it cites.
Reinforcement learning under moral uncertainty
Ecoffet, A., and Lehman, J. (2021) · 2021
Later among the works it cites.
Taking principles seriously: a hybrid approach to value alignment in artificial intelligence
Kim, T. W. N., Hooker, J. N., and Donaldson, T. (2021) · 2021
Later among the works it cites.
The computational philosophy: simulation as a core philosophical method
Mayo-Wilson, C., and Zollman, K. J. S. (2021) · 2021
Later among the works it cites.
Reinforcement learning with human advice: a survey
Najar, A., and Chetouani, M. (2021) · 2021
Later among the works it cites.
Reward is enough
Silver, D., Singh, S., Precup, D., and Sutton, R. S. (2021) · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
‘Guns don’t kill, people kill’; Values in and/or around technologies
Pitt, J. C. (2014) · 2014
Cited alongside, same era.
Towards an ethical robot: internal models, consequences and ethical action selection
Winfield, A. F. T., Blum, C., and Liu, W. (2014) · 2014
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., Petersen, S., Beattie, C., Sadik, A., Antonoglou, I., King, H., Kumaran, D., Wierstra, D., Legg, S., and Hassabis, D. (2015) · 2015
Cited alongside, same era.
Reinforcement learning as a framework for ethical decision making
Abel, D., MacGlashan, J., and Littman, M. L. (2016) · 2016
Cited alongside, same era.
Concrete problems in AI safety
Amodei, D., Olah, C., Steinhardt, J., Christiano, P., Schulman, J., and Mané, D. (2016) · 2016
Cited alongside, same era.
The social dilemma of autonomous vehicles
Bonnefon, J.-F., Shariff, A., and Rahwan, I. (2016) · 2016
Cited alongside, same era.
Implementations in machine ethics: a survey
Tolmeijer, S., Kneer, M., Sarasua, C., Christen, M., and Bernstein, A. (2021) · 2021
Later among the works it cites.
Computational ethics
Awad, E., Levine, S., Anderson, M., Anderson, S. L., Conitzer, V., Crockett, M., Everett, J. A., Evgeniou, T., Gopnik, A., Jamison, J. C., Kim, T. W., Liao, S. M., Meyer, M. N., Mikhail, J., Opoku-Agyemang, K., Borg, J. S., Schroeder, J., Sinnott-Armstrong, W., Slavkovik, M., and Tenenbaum, J. B. (2022) · 2022
Later among the works it cites.
Constitutional AI: harmlessness from AI feedback
Bai, Y., Kadavath, S., Kundu, S., Askell, A., Kernion, J., Jones, A., Chen, A., Goldie, A., Mirhoseini, A., McKinnon, C., et al. (2022) · 2022
Later among the works it cites.
Fine-tuning language models to find agreement among humans with diverse preferences
Bakker, M., Chadwick, M., Sheahan, H., Tessler, M., Campbell-Gillingham, L., Balaguer, J., McAleese, N., Glaese, A., Aslanides, J., Botvinick, M., et al. (2022) · 2022
Later among the works it cites.
AI in Finance: challenges, techniques, and opportunities
Cao, L. (2022) · 2022
Later among the works it cites.
Human-level play in the game of Diplomacy by combining language models with strategic reasoning
FAIR, Bakhtin, A., Brown, N., Dinan, E., Farina, G., Flaherty, C., Fried, D., Goff, A., Gray, J., Hu, H., Jacob, A. P., Komeili, M., Konath, K., Kwon, M., Lerer, A., Lewis, M., Miller, A. H., Mitts, S., Renduchintala, A., Roller, S., Rowe, D., Shi, W., Spisak, J., Wei, A., Wu, D., Zhang, H., and Zijlstra, M. (2022) · 2022
Later among the works it cites.
Improving alignment of dialogue agents via targeted human judgements
Glaese, A., McAleese, N., Trebacz, M., Aslanides, J., Firoiu, V., Ewalds, T., Rauh, M., Weidinger, L., Chadwick, M., Thacker, P., et al. (2022) · 2022
Later among the works it cites.
A practical guide to multi-objective reinforcement learning and planning
Hayes, C. F., Ruadulescu, R., Bargiacchi, E., Kallstrom, J., Macfarlane, M., Reymond, M., Verstraeten, T., Zintgraf, L. M., Dazeley, R., Heintz, F., Howley, E., Irissappane, A. A., Mannion, P., Now’e, A., de Oliveira Ramos, G., Restelli, M., Vamplew, P., and Roijers, D. M. (2022) · 2022
Later among the works it cites.
Recommender systems: Trends and frontiers
Jannach, D., Pu, P., Ricci, F., and Zanker, M. (2022) · 2022
Later among the works it cites.
Human-centred mechanism design with Democratic AI
Koster, R., Balaguer, J., Tacchetti, A., Weinstein, A., Zhu, T., Hauser, O., Williams, D., Campbell-Gillingham, L., Thacker, P., Botvinick, M., and Summerfield, C. (2022) · 2022
Later among the works it cites.
Enforcing ethical goals over reinforcement-learning policies
Neufeld, E. A., Bartocci, E., Ciabattoni, A., and Governatori, G. (2022) · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P. F., Leike, J., and Lowe, R. (2022) · 2022
Later among the works it cites.
The effects of Reward Misspecification: mapping and mitigating misaligned models
Pan, A., Bhatia, K., and Steinhardt, J. (2022) · 2022
Later among the works it cites.
MORAL: Aligning ai with human norms through multi-objective reinforced active learning
Peschl, M., Zgonnikov, A., Oliehoek, F. A., and Siebert, L. C. (2022) · 2022
Later among the works it cites.
AI in Health and Medicine
Rajpurkar, P., Chen, E., Banerjee, O., and Topol, E. J. (2022) · 2022
Later among the works it cites.
Defining and characterizing reward gaming
Skalse, J., Howe, N., Krasheninnikov, D., and Krueger, D. (2022) · 2022
Later among the works it cites.
Playing repeated games with Large Language Models
Akata, E., Schulz, L., Coda-Forno, J., Oh, S. J., Bethge, M., and Schulz, E. (2023) · 2023
Closest in time.
Quantifying harm
Beckers, S., Chockler, H., and Halpern, J. Y. (2023) · 2023
Closest in time.
Open problems and fundamental limitations of reinforcement learning from human feedback
Casper, S., Davies, X., Shi, C., Gilbert, T. K., Scheurer, J., Rando, J., Freedman, R., Korbak, T., Lindner, D., Freire, P., Wang, T., Marks, S., Ségerie, C.-R., Carroll, M., Peng, A., Christoffersen, P. J., Damani, M., Slocum, S., Anwar, U., Siththaranjan, A., Nadeau, M., Michaud, E. J., Pfau, J., Krasheninnikov, D., Chen, X., di Langosco, L. L., Hase, P., Biyik, E., Dragan, A. D., Krueger, D., Sadigh, D., and Hadfield-Menell, D. (2023) · 2023
Closest in time.
Strategic reasoning with language models
Gandhi, K., Sadigh, D., and Goodman, N. D. (2023) · 2023
Closest in time.
Mechanistically analyzing the effects of fine-tuning on procedurally defined tasks
Jain, S., Kirk, R., Singh, E., Robert, L., Dick, P., Tanaka, H., Grefenstette, E., Rocktäschel, T., and Krueger, D. S. (2023) · 2023
Closest in time.
A theory of injunctive norms
Kimbrough, E. O., and Vostroknutov, A. (2023) · 2023
Closest in time.
The debate over understanding in ai’s large language models
Mitchell, M., and Krakauer, D. C. (2023) · 2023
Closest in time.
Do the rewards justify the means? Measuring trade-offs between rewards and ethical behavior in the MACHIAVELLI benchmark
Pan, A., Chan, J. S., Zou, A., Li, N., Basart, S., Woodside, T., Zhang, H., Emmons, S., and Hendrycks, D. (2023) · 2023
Closest in time.
Generative agents: interactive simulacra of human behavior
Park, J. S., O’Brien, J. C., Cai, C. J., Morris, M. R., Liang, P., and Bernstein, M. S. (2023) · 2023
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C. D., and Finn, C. (2023) · 2023
Closest in time.
The puzzle of evaluating moral cognition in artificial agents
Reinecke, M. G., Mao, Y., Kunesch, M., Duéñez-Guzmán, E. A., Haas, J., and Leibo, J. Z. (2023) · 2023
Closest in time.
Offline RL for Natural Language Generation with Implicit Language Q Learning
Snell, C., Kostrikov, I., Su, Y., Yang, M., and Levine, S. (2023) · 2023
Closest in time.
Modeling moral choices in social dilemmas with multi-agent reinforcement learning
Tennant, E., Hailes, S., and Musolesi, M. (2023) · 2023
Closest in time.
AI Safety Summit 2023: The Bletchley Declaration
UK Government (2023) · 2023
Closest in time.
Vezhnevets, A. S., Agapiou, J. P., Aharon, A., Ziv, R., Matyas, J., Duéñez-Guzmán, E. A., Cunningham, W. A., Osindero, S., Karmon, D., and Leibo, J. Z. (2023) · 2023
Closest in time.
Using the veil of ignorance to align AI systems with principles of justice
Weidinger, L., McKee, K. R., Everett, R., Huang, S., Zhu, T. O., Chadwick, M. J., Summerfield, C., and Gabriel, I. (2023) · 2023
Closest in time.
Collective Constitutional AI
Anthropic (2024) · 2024
Closest in time.
Moral AI: And How We Get There
Borg, J. S., Sinnott-Armstrong, W., and Conitzer, V. (2024) · 2024
Closest in time.
AI Alignment with Changing and Influenceable Reward Functions
Carroll, M., Foote, D., Siththaranjan, A., Russell, S., and Dragan, A. (2024) · 2024
Closest in time.
Can large language models serve as rational players in game theory? A systematic analysis
Fan, C., Chen, J., Jin, Y., and He, H. (2024) · 2024
Closest in time.
Do moral values change with the seasons?
Hohm, I., O’Shea, B. A., and Schaller, M. (2024) · 2024
Closest in time.
AI Alignment: A Comprehensive Survey
Ji, J., Qiu, T., Chen, B., Zhang, B., Lou, H., Wang, K., Duan, Y., He, Z., Zhou, J., Zhang, Z., Zeng, F., Ng, K. Y., Dai, J., Pan, X., O’Gara, A., Lei, Y., Xu, H., Tse, B., Fu, J., McAleer, S., Yang, Y., Wang, Y., Zhu, S.-C., Guo, Y., and Gao, W. (2024) · 2024
Closest in time.
Position: Levels of AGI for operationalizing progress on the path to AGI
Morris, M. R., Sohl-Dickstein, J., Fiedel, N., Warkentin, T., Dafoe, A., Faust, A., Farabet, C., and Legg, S. (2024) · 2024
Closest in time.
Model Specs
OpenAI (2024) · 2024
Closest in time.
Position: A roadmap to pluralistic alignment
Sorensen, T., Moore, J., Fisher, J., Gordon, M. L., Mireshghallah, N., Rytting, C. M., Ye, A., Jiang, L., Lu, X., Dziri, N., Althoff, T., and Choi, Y. (2024) · 2024
Closest in time.
Artificial intelligence and agency: Tie-breaking in AI decision-making
Swanepoel, D., and Corks, D. (2024) · 2024
Closest in time.
LLM as a Mastermind: A Survey of Strategic Reasoning with Large Language Models
Zhang, Y., Mao, S., Ge, T., Wang, X., de Wynter, A., Xia, Y., Wu, W., Song, T., Lan, M., and Wei, F. (2024) · 2024
Closest in time.