Fetching the paper…
Reading the bibliography…
The dominant practice of AI alignment assumes (1) that preferences are an adequate representation of human values, (2) that human rationality can be understood in terms of maximizing the satisfaction of preferences, and (3) that AI systems should be aligned with the preferences of one or more humans to ensure that they behave safely and in accordance with our values.
Eckersley, P. (2018) · 1901
Earlier work this paper cites.
Risks from learned optimization in advanced machine learning systems
Hubinger, E., van Merwijk, C., Mikulik, V., Skalse, J., and Garrabrant, S. (2019) · 1906
Earlier work this paper cites.
Corrigibility with utility preservation
Holtman, K. (2019) · 1908
Earlier work this paper cites.
A Note on the Pure Theory of Consumer’s Behaviour
Samuelson, P. A. (1938) · 1938
Earlier work this paper cites.
Theory of Games and Economic Behavior
von Neumann, J. and Morgenstern, O. (1944) · 1944
Earlier work this paper cites.
The use of knowledge in society
Hayek, F. (1945) · 1945
Earlier work this paper cites.
Deontic logic
von Wright, G. H. (1951) · 1951
Earlier work this paper cites.
Ethical absolutism and the ideal observer
Firth, R. (1952) · 1952
Earlier work this paper cites.
Cardinal utility in welfare economics and in the theory of risk-taking
Harsanyi, J. C. (1953) · 1953
Earlier work this paper cites.
Cardinal utility
Strotz, R. H. (1953) · 1953
Earlier work this paper cites.
The definition of an “ideal observer” theory in ethics
Brandt, R. B. (1955) · 1955
Earlier work this paper cites.
Cardinal welfare, individualistic ethics, and interpersonal comparisons of utility
Harsanyi, J. C. (1955) · 1955
Earlier work this paper cites.
A behavioral model of rational choice
Simon, H. A. (1957) · 1957
Earlier work this paper cites.
A Simultaneous Axiomatization of Utility and Subjective Probability
Bolker, E. D. (1967) · 1967
Earlier work this paper cites.
Prior probabilities
Jaynes, E. T. (1968) · 1968
Earlier work this paper cites.
STRIPS: A new approach to the application of theorem proving to problem solving
Fikes, R. E. and Nilsson, N. J. (1971) · 1971
Earlier work this paper cites.
A Theory of Justice: Original Edition
Rawls, J. (1971) · 1971
Earlier work this paper cites.
History and Class Consciousness: Studies in Marxist Dialectics
Lukacs, G. and Livingstone, R. (1972) · 1972
Earlier work this paper cites.
The Foundations of Statistics
Savage, L. J. (1972) · 1972
Earlier work this paper cites.
The logic of preference reconsidered
von Wright, G. H. (1972) · 1972
Earlier work this paper cites.
The informational size of message spaces
Mount, K. and Reiter, S. (1974) · 1974
Earlier work this paper cites.
Can the maximin principle serve as a basis for morality? A critique of John Rawls’s theory
Harsanyi, J. C. (1975) · 1975
Earlier work this paper cites.
Other solutions to Nash’s bargaining problem
Kalai, E. and Smorodinsky, M. (1975) · 1975
Earlier work this paper cites.
Economy and Society: An Outline of Interpretive Sociology
Weber, M. (1978) · 1978
Earlier work this paper cites.
Prospect theory: An analysis of decision under risk
Kahneman, D. and Tversky, A. (1979) · 1979
Earlier work this paper cites.
Individual Choice Behavior: A Theoretical Analysis
Luce, R. D. (1979) · 1979
Earlier work this paper cites.
Rational decision making in business organizations
Simon, H. A. (1979) · 1979
Earlier work this paper cites.
Moral Thinking: Its Levels, Method, and Point
Hare, R. M. (1981) · 1981
Earlier work this paper cites.
The competitive allocation process is informationally efficient uniquely
Jordan, J. S. (1982) · 1982
Earlier work this paper cites.
Morals by Agreement
Gauthier, D. (1986) · 1986
Earlier work this paper cites.
Intention, Plans, and Practical Reason
Bratman, M. (1987) · 1987
Earlier work this paper cites.
The Complexity of Markov Decision Processes
Papadimitriou, C. H. and Tsitsiklis, J. N. (1987) · 1987
Earlier work this paper cites.
Plans and resource-bounded practical reasoning
Bratman, M. E., Israel, D. J., and Pollack, M. E. (1988) · 1988
Earlier work this paper cites.
Personal identity and the unity of agency: A Kantian response to Parfit
Korsgaard, C. M. (1989) · 1989
Earlier work this paper cites.
What was socialism, and why did it fall?
Verdery, K. (2005) · 1989
Earlier work this paper cites.
Preference-based deontic logic
Hansson, S. O. (1990) · 1990
Earlier work this paper cites.
Economic Calculation in the Socialist Commonwealth
von Mises, L. (1990) · 1990
Earlier work this paper cites.
The Logic of Decision
Jeffrey, R. C. (1991) · 1991
Earlier work this paper cites.
Probabilistic logic programming
Ng, R. and Subrahmanian, V. S. (1992) · 1992
Earlier work this paper cites.
Advances in prospect theory: Cumulative representation of uncertainty
Tversky, A. and Kahneman, D. (1992) · 1992
Earlier work this paper cites.
Alienation, Consequentialism, and the Demands of Morality
Railton, P. (1993) · 1993
Earlier work this paper cites.
Political Liberalism
Rawls, J. (1993) · 1993
Earlier work this paper cites.
Game Theory and the Social Contract
Binmore, K. G. (1994) · 1994
Earlier work this paper cites.
Act utilitarianism and decision procedures
Frazier, R. L. (1994) · 1994
Earlier work this paper cites.
Advances in random utility models
Horowitz, J. L., Bolduc, D., Divakar, S., Geweke, J., Gönül, F., Hajivassiliou, V., Koppelman, F. S., Keane, M., Matzkin, R., Rossi, P., et al. (1994) · 1994
Earlier work this paper cites.
Provably bounded-optimal agents
Russell, S. J. and Subramanian, D. (1994) · 1994
Earlier work this paper cites.
Value in Ethics and Economics
Anderson, E. (1995) · 1995
Earlier work this paper cites.
On the acceptability of arguments and its fundamental role in nonmonotonic reasoning, logic programming and n-person games
Dung, P. M. (1995) · 1995
Earlier work this paper cites.
Dealing with the complexity of economic calculations
Rust, J. P. (1996) · 1996
Earlier work this paper cites.
Incommensurability, Incomparability, and Practical Reason
Chang, R., editor (1997) · 1997
Earlier work this paper cites.
A case for happiness, cardinalism, and interpersonal comparability
Ng, Y.-K. (1997) · 1997
Earlier work this paper cites.
Decision procedures, standards of rightness and impartiality
Stark, C. A. (1997) · 1997
Earlier work this paper cites.
On the acceptability of arguments in preference-based argumentation
Amgoud, L. and Cayrol, C. (1998) · 1998
Earlier work this paper cites.
Commensuration as a social process
Espeland, W. N. and Stevens, M. L. (1998) · 1998
Earlier work this paper cites.
Seeing Like A State: How Certain Schemes to Improve the Human Condition Have Failed
Scott, J. C. (1998) · 1998
Earlier work this paper cites.
Ten years of the rational analysis of cognition
Chater, N. and Oaksford, M. (1999) · 1999
Earlier work this paper cites.
Egalitarian justice and interpersonal comparison
Clayton, M. and Williams, A. (1999) · 1999
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
Ng, A. Y., Harada, D., and Russell, S. J. (1999) · 1999
Earlier work this paper cites.
Engaging Reason: On the Theory of Value and Action
Raz, J. (1999) · 1999
Earlier work this paper cites.
Commodities and capabilities
Sen, A. et al. (1999) · 1999
Earlier work this paper cites.
Rational choice theory in law and economics
Ulen, T. S. (1999) · 1999
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
Ng, A. Y. and Russell, S. J. (2000) · 2000
Earlier work this paper cites.
What We Owe To Each Other
Scanlon, T. (2000) · 2000
Earlier work this paper cites.
Symposium on Amartya Sen’s philosophy: Adaptive preferences and women’s options
Nussbaum, M. C. (2001) · 2001
Earlier work this paper cites.
Beyond rational choice theory
Boudon, R. (2003) · 2003
Earlier work this paper cites.
Probabilistic logic learning
De Raedt, L. and Kersting, K. (2003) · 2003
Earlier work this paper cites.
Predicting and indulging changing preferences
Loewenstein, G. and Angner, E. (2003) · 2003
Earlier work this paper cites.
Argumentation-based negotiation
Rahwan, I., Ramchurn, S. D., Jennings, N. R., McBurney, P., Parsons, S., and Sonenberg, L. (2003) · 2003
Earlier work this paper cites.
Implications of rational inattention
Sims, C. A. (2003) · 2003
Earlier work this paper cites.
Modeling interdependent consumer preferences
Yang, S. and Allenby, G. M. (2003) · 2003
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
Abbeel, P. and Ng, A. Y. (2004) · 2004
Earlier work this paper cites.
CP-nets: A Tool for Representing and Reasoning with Conditional Ceteris Paribus Preference Statements
Boutilier, C., Brafman, R. I., Domshlak, C., Hoos, H. H., and Poole, D. (2004) · 2004
Earlier work this paper cites.
Can desires provide reasons for action
Chang, R. (2004) · 2004
Earlier work this paper cites.
Three Faces of Desire
Schroeder, T. (2004) · 2004
Earlier work this paper cites.
Coherent extrapolated volition
Yudkowsky, E. (2004) · 2004
Earlier work this paper cites.
The Grammar of Society: The Nature and Dynamics of Social Norms
Bicchieri, C. (2005) · 2005
Earlier work this paper cites.
Plan constraints and preferences in PDDL3
Gerevini, A. and Long, D. (2005) · 2005
Earlier work this paper cites.
Living, and thinking about it: Two perspectives on life
Kahneman, D. and Riis, J. (2005) · 2005
Earlier work this paper cites.
Prioritarian welfare functions: An elaboration and justification
Lumer, C. et al. (2005) · 2005
Earlier work this paper cites.
Interdependent preferences and reciprocity
Sobel, J. (2005) · 2005
Earlier work this paper cites.
Ideology and Ideological State Apparatuses
Althusser, L. et al. (2006) · 2006
Earlier work this paper cites.
AI research considerations for human existential safety (ARCHES)
Critch, A. and Krueger, D. (2020) · 2006
Earlier work this paper cites.
The Construction of Preference
Lichtenstein, S. and Slovic, P. (2006) · 2006
Earlier work this paper cites.
Cantor’s diagonal argument: An extension to the socialist calculation debate
Murphy, R. (2006) · 2006
Earlier work this paper cites.
Multi-Principal Assistance Games
Fickinger, A., Zhuang, S., Hadfield-Menell, D., and Russell, S. (2020) · 2007
Earlier work this paper cites.
Epistemic Injustice: Power and the Ethics of Knowing
Fricker, M. (2007) · 2007
Earlier work this paper cites.
Safety in Markets: An Impossibility Theorem for Dutch Books
Laibson, D. and Yariv, L. (2007) · 2007
Earlier work this paper cites.
Discredited data
Okidegbe, N. (2021) · 2007
Earlier work this paper cites.
Why heuristics work
Gigerenzer, G. (2008) · 2008
Earlier work this paper cites.
The nature of self-improving artificial intelligence
Omohundro, S. M. (2007) · 2008
Earlier work this paper cites.
The basic AI drives
Omohundro, S. M. (2008) · 2008
Earlier work this paper cites.
The Tractable Cognition Thesis
van Rooij, I. (2008) · 2008
Earlier work this paper cites.
Evolution of Non-Expected Utility Preferences
von Widekind, S. (2008) · 2008
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
Ziebart, B. D., Maas, A. L., Bagnell, J. A., Dey, A. K., et al. (2008) · 2008
Earlier work this paper cites.
Action understanding as inverse planning
Baker, C. L., Saxe, R., and Tenenbaum, J. B. (2009) · 2009
Earlier work this paper cites.
Voluntarist Reasons and the Sources of Normativity
Chang, R. (2009) · 2009
Earlier work this paper cites.
Reasoning about preferences in argumentation frameworks
Modgil, S. (2009) · 2009
Earlier work this paper cites.
Where do rewards come from
Singh, S., Lewis, R. L., and Barto, A. G. (2009) · 2009
Earlier work this paper cites.
Help or hinder: Bayesian models of social goal inference
Ullman, T., Baker, C., Macindoe, O., Evans, O., Goodman, N., and Tenenbaum, J. (2009) · 2009
Earlier work this paper cites.
Social norms as choreography
Gintis, H. (2010) · 2010
Earlier work this paper cites.
A contextual-bandit approach to personalized news article recommendation
Li, L., Chu, W., Langford, J., and Schapire, R. E. (2010) · 2010
Earlier work this paper cites.
Modeling interaction via the principle of maximum causal entropy
Ziebart, B. D., Bagnell, J. A., and Dey, A. K. (2010) · 2010
Earlier work this paper cites.
Preference-Satisfaction
Baber, H. E. (2011) · 2011
Earlier work this paper cites.
Learning what to value
Dewey, D. (2011) · 2011
Earlier work this paper cites.
Augmenting reinforcement learning with human feedback
Knox, W. B. and Stone, P. (2011) · 2011
Earlier work this paper cites.
Reasoning about Preference Dynamics
Liu, F. (2011) · 2011
Earlier work this paper cites.
Why do humans reason? Arguments for an argumentative theory
Mercier, H. and Sperber, D. (2011) · 2011
Earlier work this paper cites.
On What Matters
Parfit, D. (2011) · 2011
Earlier work this paper cites.
Dutch book arguments
Vineberg, S. (2011) · 2011
Earlier work this paper cites.
Values and preferences: defining preference construction
Warren, C., McGraw, A. P., and Van Boven, L. (2011) · 2011
Earlier work this paper cites.
Open problems in cooperative AI
Dafoe, A., Hughes, E., Bachrach, Y., Collins, T., McKee, K. R., Leibo, J. Z., Larson, K., and Graepel, T. (2020) · 2012
Earlier work this paper cites.
Eliciting welfare preferences from behavioural data sets
Rubinstein, A. and Salant, Y. (2012) · 2012
Earlier work this paper cites.
Generalized random utility models with multiple types
Azari Soufiani, H., Diao, H., Lai, Z., and Parkes, D. C. (2013) · 2013
Earlier work this paper cites.
Statistical Decision Theory: Foundations, Concepts, and Methods
Berger, J. (2013) · 2013
Earlier work this paper cites.
Updates and Uncertainty in CP-Nets
Cornelio, C., Goldsmith, J., Mattei, N., Rossi, F., and Venable, K. B. (2013) · 2013
Cited alongside, same era.
Policy Shaping: Integrating Human Feedback with Reinforcement Learning
Griffith, S., Subramanian, K., Scholz, J., Isbell, C. L., and Thomaz, A. L. (2013) · 2013
Cited alongside, same era.
Thermodynamics as a theory of decision-making with information-processing costs
Ortega, P. A. and Braun, D. A. (2013) · 2013
Cited alongside, same era.
Programming by feedback
Akrour, R., Schoenauer, M., Sebag, M., and Souplet, J.-C. (2014) · 2014
Cited alongside, same era.
Should subjective probabilities be sharp?
Bradley, S. and Steele, K. (2014) · 2014
Cited alongside, same era.
Computational rationality: Linking mechanism and behavior through bounded utility maximization
Lewis, R. L., Howes, A., and Singh, S. (2014) · 2014
Cited alongside, same era.
Geometric rationality
Garrabrant, S. (2022) · 2022
Later among the works it cites.
Jury learning: Integrating dissenting voices into machine learning models
Gordon, M. L., Lam, M. S., Park, J. S., Patel, K., Hancock, J., Hashimoto, T., and Bernstein, M. S. (2022) · 2022
Later among the works it cites.
Money-pump arguments
Gustafsson, J. E. (2022) · 2022
Later among the works it cites.
Cognitive science as a source of forward and inverse models of human decisions for robotics and control
Ho, M. K. and Griffiths, T. L. (2022) · 2022
Later among the works it cites.
Reward machines: Exploiting reward function structure in reinforcement learning
Icarte, R. T., Klassen, T. Q., Valenzano, R., and McIlraith, S. A. (2022) · 2022
Later among the works it cites.
When to make exceptions: Exploring language models as accounts of human moral judgment
Jin, Z., Levine, S., Gonzalez Adauto, F., Kamal, O., Sap, M., Sachan, M., Mihalcea, R., Tenenbaum, J., and Schölkopf, B. (2022) · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Decision-making under risk: Integrating perspectives from biology, economics, and psychology
Mishra, S. (2014) · 2014
Cited alongside, same era.
Unwritten rules: Virtual bargaining underpins social interaction, culture, and society
Misyak, J. B., Melkonyan, T., Zeitoun, H., and Chater, N. (2014) · 2014
Cited alongside, same era.
Transformative Experience
Paul, L. A. (2014) · 2014
Cited alongside, same era.
Staying Alive: Personal Identity, Practical Concerns, and the Unity of a Life
Schechtman, M. (2014) · 2014
Cited alongside, same era.
Computational rationality: A converging paradigm for intelligence in brains, minds, and machines
Gershman, S. J., Horvitz, E. J., and Tenenbaum, J. B. (2015) · 2015
Cited alongside, same era.
Algorithmic rationality: Game theory with costly computation
Halpern, J. Y. and Pass, R. (2015) · 2015
Cited alongside, same era.
Later among the works it cites.
Aligned with whom? Direct and social goals for AI systems
Korinek, A. and Balwit, A. (2022) · 2022
Later among the works it cites.
The Boltzmann policy distribution: Accounting for systematic suboptimality in human models
Laidlaw, C. and Dragan, A. (2022) · 2022
Later among the works it cites.
End-user audits: A system empowering communities to lead large-scale investigations of harmful algorithmic behavior
Lam, M. S., Gordon, M. L., Metaxa, D., Hancock, J. T., Landay, J. A., and Bernstein, M. S. (2022) · 2022
Later among the works it cites.
Inferring rewards from language in context
Lin, J., Fried, D., Klein, D., and Dragan, A. (2022) · 2022
Later among the works it cites.
Normative Reasons: Between Reasoning and Explanation
Logins, A. (2022) · 2022
Later among the works it cites.
The alignment problem from a deep learning perspective
Ngo, R., Chan, L., and Mindermann, S. (2022) · 2022
Later among the works it cites.
Computational Rationality as a Theory of Interaction
Oulasvirta, A., Jokinen, J. P. P., and Howes, A. (2022) · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P., Leike, J., and Lowe, R. (2022) · 2022
Later among the works it cites.
TALM: Tool augmented language models
Parisi, A., Zhao, Y., and Fiedel, N. (2022) · 2022
Later among the works it cites.
Meaning without reference in large language models
Piantadosi, S. T. and Hill, F. (2022) · 2022
Later among the works it cites.
Self-critiquing models for assisting human evaluators
Saunders, W., Yeh, C., Wu, J., Bills, S., Ouyang, L., Ward, J., and Leike, J. (2022) · 2022
Later among the works it cites.
How AI Fails Us
Siddarth, D., Acemoglu, D., Allen, D., Crawford, K., Evans, J., Jordan, M., and Weyl, E. G. (2022) · 2022
Later among the works it cites.
Epistemic injustice and data science technologies
Symons, J. and Alvarado, R. (2022) · 2022
Later among the works it cites.
What does it mean to give someone what they want? the nature of preferences in recommender systems
Thorburn, L., Stray, J., and Bengani, P. (2022) · 2022
Later among the works it cites.
What Should AI Owe To Us? Accountable and Aligned AI Systems via Contractualist AI Alignment
Zhi-Xuan, T. (2022) · 2022
Later among the works it cites.
A hierarchical Bayesian approach to inverse reinforcement learning with symbolic reward machines
Zhou, W. and Li, W. (2022) · 2022
Later among the works it cites.
The Value Change Problem (sequence)
Ammann, N. (2023) · 2023
Later among the works it cites.
Will AI avoid exploitation? Artificial General Intelligence and Expected Utility Theory
Bales, A. (2023) · 2023
Later among the works it cites.
AI scientists: Safe and useful AI?
Bengio, Y. (2023) · 2023
Later among the works it cites.
Thinking about thinking as rational computation
Berke, M., Tenenbaum, A., Sterling, B., and Jara-Ettinger, J. (2023) · 2023
Later among the works it cites.
Making intelligence: Ethical values in IQ and ML benchmarks
Blili-Hamelin, B. and Hancox-Li, L. (2023) · 2023
Later among the works it cites.
The perils of trial-and-error reward design: Misdesign through overfitting and invalid task specifications
Booth, S., Knox, W. B., Shah, J., Niekum, S., Stone, P., and Allievi, A. (2023) · 2023
Later among the works it cites.
Settling the Reward Hypothesis
Bowling, M., Martin, J. D., Abel, D., and Dabney, W. (2023) · 2023
Later among the works it cites.
Characterizing manipulation from AI systems
Carroll, M., Chan, A., Ashton, H., and Krueger, D. (2023) · 2023
Later among the works it cites.
Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Casper, S., Davies, X., Shi, C., Gilbert, T. K., Scheurer, J., Rando, J., Freedman, R., Korbak, T., Lindner, D., Freire, P., Wang, T., Marks, S., Segerie, C.-R., Carroll, M., Peng, A., Christoffersen, P., Damani, M., Slocum, S., Anwar, U., Siththaranjan, A., Nadeau, M., Michaud, E. J., Pfau, J., Krasheninnikov, D., Chen, X., Langosco, L., Hase, P., Bıyık, E., Dragan, A., Krueger, D., Sadigh, D., and Hadfield-Menell, D. (2023) · 2023
Later among the works it cites.
How could we make a social robot? A virtual bargaining approach
Chater, N. (2023) · 2023
Later among the works it cites.
Get it in writing: Formal contracts mitigate social dilemmas in multi-agent RL
Christoffersen, P. J., Haupt, A. A., and Hadfield-Menell, D. (2023) · 2023
Later among the works it cites.
Is there an epistemic advantage to being oppressed?
Dror, L. (2023) · 2023
Later among the works it cites.
Faith and fate: Limits of transformers on compositionality
Dziri, N., Lu, X., Sclar, M., Li, X. L., Jian, L., Lin, B. Y., West, P., Bhagavatula, C., Bras, R. L., Hwang, J. D., et al. (2023) · 2023
Later among the works it cites.
AI-powered Bing chat gains three distinct personalities
Edwards, B. (2023) · 2023
Later among the works it cites.
Scaling laws for reward model overoptimization
Gao, L., Schulman, J., and Hilton, J. (2023) · 2023
Later among the works it cites.
The effect of modeling human rationality level on learning rewards from multiple feedback types
Ghosal, G. R., Zurek, M., Brown, D. S., and Dragan, A. D. (2023) · 2023
Later among the works it cites.
Natural selection favors AIs over humans
Hendrycks, D. (2023) · 2023
Later among the works it cites.
Dirty data labeled dirt cheap: Epistemic injustice in machine learning systems
Hull, G. (2023) · 2023
Later among the works it cites.
ChatGPT by OpenAI: The end of litigation lawyers?
Iu, K. Y. and Wong, V. M.-Y. (2023) · 2023
Later among the works it cites.
Language agents as digital representatives in collective decision-making
Jarrett, D., Pislar, M., Bakker, M. A., Tessler, M. H., Koster, R., Balaguer, J., Elie, R., Summerfield, C., and Tacchetti, A. (2023) · 2023
Later among the works it cites.
In Conversation with Artificial Intelligence: Aligning language models with human values
Kasirzadeh, A. and Gabriel, I. (2023) · 2023
Later among the works it cites.
FANToM: A benchmark for stress-testing machine theory of mind in interactions
Kim, H., Sclar, M., Zhou, X., Bras, R., Kim, G., Choi, Y., and Sap, M. (2023) · 2023
Later among the works it cites.
Power-seeking can be probable and predictive for trained agents
Krakovna, V. and Kramar, J. (2023) · 2023
Later among the works it cites.
Neuro-symbolic models of human moral judgment: LLMs as automatic feature extractors
Kwon, J., Levine, S., and Tenenbaum, J. B. (2023a) · 2023
Later among the works it cites.
The Alignment Ceiling: Objective Mismatch in Reinforcement Learning from Human Feedback
Lambert, N. and Calandra, R. (2023) · 2023
Later among the works it cites.
Entangled Preferences: The History and Risks of Reinforcement Learning and Human Feedback
Lambert, N., Gilbert, T. K., and Zick, T. (2023) · 2023
Later among the works it cites.
AI safety on whose terms?
Lazar, S. and Nelson, A. (2023) · 2023
Later among the works it cites.
Value as semantics: Representations of human moral and hedonic value in large language models
Leshinskaya, A., San Franscisco, C., and Chakroff, A. (2023) · 2023
Later among the works it cites.
Resource-rational contractualism: A triple theory of moral cognition
Levine, S., Chater, N., Tenenbaum, J., and Cushman, F. (2023) · 2023
Later among the works it cites.
A.I. is coming for lawyers, again
Lohr, S. (2023) · 2023
Later among the works it cites.
A discerning several thousand judgments: GPT-3 rates the article+ adjective+ numeral+ noun construction
Mahowald, K. (2023) · 2023
Later among the works it cites.
AI alignment and social choice: Fundamental limitations and policy implications
Mishra, A. (2023) · 2023
Later among the works it cites.
A goal-centric outlook on learning
Molinaro, G. and Collins, A. G. (2023) · 2023
Later among the works it cites.
Revealed incomplete preferences
Nielsen, K. and Rigotti, L. (2023) · 2023
Later among the works it cites.
Reimagining democracy for AI
Ovadya, A. (2023) · 2023
Later among the works it cites.
Invulnerable incomplete preferences: A formal statement
Petersen, S. (2023) · 2023
Later among the works it cites.
The best game in town: The reemergence of the language-of-thought hypothesis across the cognitive sciences
Quilty-Dunn, J., Porot, N., and Mandelbaum, E. (2023) · 2023
Later among the works it cites.
Whitepaper
Siddarth, D. and Huang, S. (2023) · 2023
Later among the works it cites.
Invariance in policy optimisation and partial identifiability in reward learning
Skalse, J. M. V., Farrugia-Roberts, M., Russell, S., Abate, A., and Gleave, A. (2023) · 2023
Later among the works it cites.
GPT-4 doesn’t know it’s wrong: An analysis of iterative prompting for reasoning problems
Stechly, K., Marquez, M., and Kambhampati, S. (2023) · 2023
Later among the works it cites.
There are no coherence theorems
Thornley, E. (2023) · 2023
Later among the works it cites.
Can large language models really improve by self-critiquing their own plans?
Valmeekam, K., Marquez, M., and Kambhampati, S. (2023a) · 2023
Later among the works it cites.
Using the Veil of Ignorance to align AI systems with principles of justice
Weidinger, L., McKee, K. R., Everett, R., Huang, S., Zhu, T. O., Chadwick, M. J., Summerfield, C., and Gabriel, I. (2023) · 2023
Later among the works it cites.
Why not subagents?
Wentworth, J. (2023) · 2023
Later among the works it cites.
Deceptive alignment is < < 1% likely by default
Wheaton, D. (2023) · 2023
Later among the works it cites.
Wong, L., Grand, G., Lew, A. K., Goodman, N. D., Mansinghka, V. K., Andreas, J., and Tenenbaum, J. B. (2023) · 2023
Later among the works it cites.
From instructions to intrinsic human values–a survey of alignment goals for big models
Yao, J., Yi, X., Wang, X., Wang, J., and Xie, X. (2023) · 2023
Later among the works it cites.
Principled reinforcement learning with human feedback from pairwise or k-wise comparisons
Zhu, B., Jordan, M., and Jiao, J. (2023) · 2023
Later among the works it cites.
Deductive closure training of language models for coherence, accuracy, and updatability
Akyürek, A. F., Akyürek, E., Choshen, L., Wijaya, D., and Andreas, J. (2024) · 2024
Closest in time.
Roles guide rapid inferences about agent knowledge and behavior
Baker, A., Dunham, Y., and Jara-Ettinger, J. (2024) · 2024
Closest in time.
Unsocial Intelligence: An Investigation of the Assumptions of AGI Discourse
Blili-Hamelin, B., Hancox-Li, L., and Smart, A. (2024) · 2024
Closest in time.
Aligning robot and human representations
Bobu, A., Peng, A., Agrawal, P., Shah, J., and Dragan, A. D. (2024) · 2024
Closest in time.
AI alignment with changing and influenceable reward functions
Carroll, M., Foote, D., Siththaranjan, A., Russell, S., and Dragan, A. (2024) · 2024
Closest in time.
Can formal argumentative reasoning enhance LLMs’ performances?
Castagna, F., Sassoon, I., and Parsons, S. (2024) · 2024
Closest in time.
Position: Social choice should guide AI alignment in dealing with diverse human feedback
Conitzer, V., Freedman, R., Heitzig, J., Holliday, W. H., Jacobs, B. M., Lambert, N., Mossé, M., Pacuit, E., Russell, S., Schoelkopf, H., Tewolde, E., and Zwicker, W. S. (2024) · 2024
Closest in time.
Revisiting the computation problem
Cwik, P. and Engelhardt, L. (2024) · 2024
Closest in time.
Safeguarded AI: Constructing guaranteed safety
Dalrymple, D. D. (2024) · 2024
Closest in time.
Towards guaranteed safe AI: A framework for ensuring robust and reliable AI systems
Dalrymple, D. D., Skalse, J., Bengio, Y., Russell, S., Tegmark, M., Seshia, S., Omohundro, S., Szegedy, C., Goldhaber, B., Ammann, N., Abate, A., Halpern, J., Barrett, C., Zhao, D., Zhi-Xuan, T., Wing, J., and Tenenbaum, J. (2024) · 2024
Closest in time.
Goals as reward-producing programs
Davidson, G., Todd, G., Togelius, J., Gureckis, T. M., and Lake, B. M. (2024) · 2024
Closest in time.
A density estimation perspective on learning from pairwise human preferences
Dumoulin, V., Johnson, D. D., Castro, P. S., Larochelle, H., and Dauphin, Y. (2024) · 2024
Closest in time.
A matter of principle? AI alignment as the fair treatment of claims
Gabriel, I. and Keeling, G. (2024) · 2024
Closest in time.
Compositional preference models for aligning LMs
Go, D., Korbak, T., Kruszewski, G., Rozen, J., and Dymetman, M. (2024) · 2024
Closest in time.
Contrastive preference learning: Learning from human feedback without reinforcement learning
Hejna, J., Rafailov, R., Sikchi, H., Finn, C., Niekum, S., Knox, W. B., and Sadigh, D. (2024) · 2024
Closest in time.
Collective Constitutional AI: Aligning a language model with public input
Huang, S., Siddarth, D., Lovitt, L., Liao, T. I., Durmus, E., Tamkin, A., and Ganguli, D. (2024) · 2024
Closest in time.
Modeling boundedly rational agents with latent inference budgets
Jacob, A. P., Gupta, A., and Andreas, J. (2024) · 2024
Closest in time.
The institutional stance
Jara-Ettinger, J. and Dunham, Y. (2024) · 2024
Closest in time.
The benefits, risks and bounds of personalizing the alignment of large language models to individuals
Kirk, H. R., Vidgen, B., Röttger, P., and Hale, S. A. (2024) · 2024
Closest in time.
What are human values, and how do we align AI to them?
Klingefjord, O., Lowe, R., and Edelman, J. (2024) · 2024
Closest in time.
Legitimacy, authority, and democratic duties of explanation
Lazar, S. (2024) · 2024
Closest in time.
When rules are over-ruled: Virtual bargaining as a contractualist method of moral judgment
Levine, S., Kleiman-Weiner, M., Chater, N., Cushman, F., and Tenenbaum, J. B. (2024) · 2024
Closest in time.
Beneficent intelligence: a capability approach to modeling benefit, assistance, and associated moral failures through ai systems
London, A. J. and Heidari, H. (2024) · 2024
Closest in time.
Dissociating language and thought in large language models
Mahowald, K., Ivanova, A. A., Blank, I. A., Kanwisher, N., Tenenbaum, J. B., and Fedorenko, E. (2024) · 2024
Closest in time.
Can you learn semantics through next-word prediction? The case of entailment
Merrill, W., Wu, Z., Naka, N., Kim, Y., and Linzen, T. (2024) · 2024
Closest in time.
Evaluating cognitive maps and planning in large language models with CogEval
Momennejad, I., Hasanbeig, H., Vieira Frujeri, F., Sharma, H., Jojic, N., Palangi, H., Ness, R., and Larson, J. (2024) · 2024
Closest in time.
Confronting reward model overoptimization with constrained RLHF
Moskovitz, T., Singh, A. K., Strouse, D., Sandholm, T., Salakhutdinov, R., Dragan, A. D., and McAleer, S. (2024) · 2024
Closest in time.
Learning and sustaining shared normative systems via Bayesian rule induction in Markov Games
Oldenburg, N. and Zhi-Xuan, T. (2024) · 2024
Closest in time.
Improving context-aware preference modeling for language models
Pitis, S., Xiao, Z., Roux, N. L., and Sordoni, A. (2024) · 2024
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R., Sharma, A., Mitchell, E., Manning, C. D., Ermon, S., and Finn, C. (2024) · 2024
Closest in time.
Compositional capabilities of autoregressive transformers: A study on synthetic, interpretable tasks
Ramesh, R., Lubana, E. S., Khona, M., Dick, R. P., and Tanaka, H. (2024) · 2024
Closest in time.
Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF
Siththaranjan, A., Laidlaw, C., and Hadfield-Menell, D. (2024) · 2024
Closest in time.
A roadmap to pluralistic alignment
Sorensen, T., Moore, J., Fisher, J., Gordon, M., Mireshghallah, N., Rytting, C. M., Ye, A., Jiang, L., Lu, X., Dziri, N., et al. (2024) · 2024
Closest in time.
Legitimate power, illegitimate automation: The problem of ignoring legitimacy in automated decision systems
Stone, J. and Mittelstadt, B. (2024) · 2024
Closest in time.
Cognitive architectures for language agents
Sumers, T., Yao, S., Narasimhan, K., and Griffiths, T. (2024) · 2024
Closest in time.
Participation in the age of foundation models
Suresh, H., Tseng, E., Young, M., Gray, M., Pierson, E., and Levy, K. (2024) · 2024
Closest in time.
AI can help humans find common ground in democratic deliberation
Tessler, M. H., Bakker, M. A., Jarrett, D., Sheahan, H., Chadwick, M. J., Koster, R., Evans, G., Campbell-Gillingham, L., Collins, T., Parkes, D. C., et al. (2024) · 2024
Closest in time.
The shutdown problem: an AI engineering puzzle for decision theorists
Thornley, E. (2024) · 2024
Closest in time.
Towards shutdownable agents via stochastic choice
Thornley, E., Roman, A., Ziakas, C., Ho, L., and Thomson, L. (2024) · 2024
Closest in time.
AI firms mustn’t govern themselves, say ex-members of OpenAI’s board
Toner, H. and McCauley, T. (2024) · 2024
Closest in time.
Reclaiming ai as a theoretical tool for cognitive science
Van Rooij, I., Guest, O., Adolfi, F., de Haan, R., Kolokolova, A., and Rich, P. (2024) · 2024
Closest in time.
Fine-grained human feedback gives better rewards for language model training
Wu, Z., Hu, Y., Shi, W., Dziri, N., Suhr, A., Ammanabrolu, P., Smith, N. A., Ostendorf, M., and Hajishirzi, H. (2024) · 2024
Closest in time.
The Perfect Blend: Redefining RLHF with Mixture of Judges
Xu, T., Helenowski, E., Sankararaman, K. A., Jin, D., Peng, K., Han, E., Nie, S., Zhu, C., Zhang, H., Zhou, W., et al. (2024) · 2024
Closest in time.
Pragmatic instruction following and goal assistance via cooperative language-guided inverse planning
Zhi-Xuan, T., Ying, L., Mansinghka, V., and Tenenbaum, J. B. (2024b) · 2094
Closest in time.