Fetching the paper…
Reading the bibliography…
Existing AI alignment approaches assume that preferences are static, which is unrealistic: our preferences change, and may even be influenced by our interactions with AI systems themselves.
The Assistive Multi-Armed Bandit
Chan, L., Hadfield-Menell, D., Srinivasa, S., and Dragan, A · 1901
Earlier work this paper cites.
Preferences Implicit in the State of the World
Shah, R., Krasheninnikov, D., Alexander, J., Abbeel, P., and Dragan, A · 1902
Earlier work this paper cites.
On the Feasibility of Learning, Rather than Assuming, Human Biases for Reward Inference
Shah, R., Gundotra, N., Abbeel, P., and Dragan, A. D · 1906
Earlier work this paper cites.
Everitt, T., Hutter, M., Kumar, R., and Krakovna, V · 1908
Earlier work this paper cites.
Discount Rates Inferred from Decisions: An Experimental Study
Benzion, U., Rapoport, A., and Yagil, J · 1909
Earlier work this paper cites.
Hot-cold empathy gaps and medical decision making
Loewenstein, G · 1930
Earlier work this paper cites.
A Note on Measurement of Utility
Samuelson, P. A · 1937
Earlier work this paper cites.
Specious reward: A behavioral theory of impulsiveness and impulse control
Ainslie, G · 1939
Earlier work this paper cites.
The Pure Theory of Capital
Hayek, F. A., White, L. H., and White, L. H · 1941
Earlier work this paper cites.
Rank Analysis of Incomplete Block Designs: I. The Method of Paired Comparisons
Bradley, R. A. and Terry, M. E · 1952
Earlier work this paper cites.
Ethical Absolutism and the Ideal Observer
Firth, R · 1952
Earlier work this paper cites.
Welfare Economics of Variable Tastes
Harsanyi, J. C · 1953
Earlier work this paper cites.
Myopia and Inconsistency in Dynamic Utility Maximization
Strotz, R. H · 1955
Earlier work this paper cites.
Ethical Theory: The Problems of Normative and Critical Ethics
Brandt, R. B · 1959
Earlier work this paper cites.
Chapter 17. Changing Utility Functions
Peston, M. H · 1967
Earlier work this paper cites.
Consistent Planning
Pollak, R. A · 1968
Earlier work this paper cites.
Notes on endogenous change of tastes
von Weizsäcker, C. C · 1971
Earlier work this paper cites.
De Gustibus Non Est Disputandum
Stigler, G. J. and Becker, G. S · 1977
Earlier work this paper cites.
Endogenous Tastes in Demand and Welfare Analysis
Pollak, R. A · 1978
Earlier work this paper cites.
Egonomics, or the Art of Self-Management
Schelling, T. C · 1978
Earlier work this paper cites.
Ulysses and the Sirens: Studies in Rationality and Irrationality
Elster, J. (ed.) · 1979
Earlier work this paper cites.
Education’s Lasting Influence on Values
Hyman, H. H. and Wright, C. R · 1979
Earlier work this paper cites.
Personal Identity and Rationality
Parfit, D · 1982
Earlier work this paper cites.
Sour grapes: studies in the subversion of rationality
Elster, J · 1983
Earlier work this paper cites.
Reasons and persons
Parfit, D · 1984
Earlier work this paper cites.
Self-Command in Practice, in Policy, and in a Theory of Rational Choice
Schelling, T. C · 1984
Earlier work this paper cites.
Weakness of Will and the Free-Rider Problem
Elster, J · 1985
Earlier work this paper cites.
Enforcing Rules on Oneself
Schelling, T. C · 1985
Earlier work this paper cites.
Well-Being: Its Meaning, Measurement, and Moral Importance
Griffin, J · 1986
Earlier work this paper cites.
Intention, Plans, and Practical Reason
Bratman, M · 1987
Earlier work this paper cites.
Well-Being And Time
Velleman, J. D · 1991
Earlier work this paper cites.
Anomalies in Intertemporal Choice: Evidence and an Interpretation
Loewenstein, G. and Prelec, D · 1992
Earlier work this paper cites.
Learning agents for uncertain environments (extended abstract)
Russell, S · 1998
Earlier work this paper cites.
Stochastic dynamic programming with factored representations
Boutilier, C., Dearden, R., and Goldszmidt, M · 2000
Earlier work this paper cites.
Algorithms for Inverse Reinforcement Learning
Ng, A. Y. and Russell, S. J · 2000
Earlier work this paper cites.
Artificial Intelligence, Values and Alignment
Gabriel, I · 2001
Earlier work this paper cites.
Preference pollution: how markets create the desires we dislike
George, D · 2001
Earlier work this paper cites.
Time Discounting and Time Preference: A Critical Review
Frederick, S., Loewenstein, G., and O’Donoghue, T · 2002
Earlier work this paper cites.
Reward-rational (implicit) choice: A unifying formalism for reward learning, December 2020
Jeon, H. J., Milli, S., and Dragan, A. D · 2002
Earlier work this paper cites.
Perdomo, J. C., Zrnic, T., Mendler-Dünner, C., and Hardt, M · 2002
Earlier work this paper cites.
Predicting and indulging changing preferences
Loewenstein, G. and Angner, E · 2003
Earlier work this paper cites.
Time and decision: Economic and psychological perspectives on intertemporal choice
Loewenstein, G., Read, D., and Baumeister, R. (eds.) · 2003
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
Abbeel, P. and Ng, A. Y · 2004
Earlier work this paper cites.
Coherent Extrapolated Volition
Yudkowsky, E · 2004
Earlier work this paper cites.
The Influence of Advertising on Consumer Brand Preference
Ayanwale, A. B., Alimi, T., and Ayanbimipe, M. A · 2005
Earlier work this paper cites.
Therapist influence on client language during motivational interviewing sessions
Moyers, T. B. and Martin, T · 2005
Earlier work this paper cites.
Prudence for Changing Selves
Bykvist, K · 2006
Earlier work this paper cites.
AI Research Considerations for Human Existential Safety (ARCHES), May 2020
Critch, A. and Krueger, D · 2006
Earlier work this paper cites.
Big Decisions: Opting, Converting, Drifting
Ullmann-Margalit, E · 2006
Earlier work this paper cites.
Algorithmic Game Theory
Nisan, N., Roughgarden, T., Tardos, E., and Vazirani, V. V. (eds.) · 2007
Earlier work this paper cites.
Narrative identity and eudaimonic well-being
Bauer, J. J., McAdams, D. P., and Pals, J. L · 2008
Earlier work this paper cites.
Individual laboratory-measured discount rates predict field behavior
Chabris, C. F., Laibson, D., Morris, C. L., Schuldt, J. P., and Taubinsky, D · 2008
Earlier work this paper cites.
Multiobjective Decision Making: Theory and Methodology
Chankong, V. and Haimes, Y. Y · 2008
Earlier work this paper cites.
Pro‐environmental products: marketing influence on consumer purchase decision
Pickett‐Baker, J. and Ozaki, R · 2008
Earlier work this paper cites.
Nudge: Improving decisions about health, wealth, and happiness
Thaler, R. H. and Sunstein, C. R · 2008
Earlier work this paper cites.
In Defense of Adaptive Preferences
Bruckner, D. W · 2009
Earlier work this paper cites.
The EMPATHIC Framework for Task Learning from Implicit Human Feedback, December 2020
Cui, Y., Zhang, Q., Allievi, A., Stone, P., Niekum, S., and Knox, W. B · 2009
Earlier work this paper cites.
Preference Change
Grüne-Yanoff, T. and Hansson, S. O. (eds.) · 2009
Earlier work this paper cites.
Debate: To Nudge or Not to Nudge*
Hausman, D. and Welch, B · 2009
Cited alongside, same era.
Hidden Incentives for Auto-Induced Distributional Shift
Krueger, D., Maharaj, T., and Leike, J · 2009
Cited alongside, same era.
Robust Policy Computation in Reward-Uncertain MDPs Using Nondominated Policies
Regan, K. and Boutilier, C · 2010
Cited alongside, same era.
Modeling Interaction via the Principle of Maximum Causal Entropy
Ziebart, B. D., Bagnell, J. A., and Dey, A. K · 2010
Cited alongside, same era.
Dwork, C., Hardt, M., Pitassi, T., Reingold, O., and Zemel, R · 2011
Cited alongside, same era.
The precipice: existential risk and the future of humanity
Ord, T · 2021
Later among the works it cites.
Where To Next? A Dynamic Model of User Preferences
Sanna Passino, F., Maystre, L., Moor, D., Anderson, A., and Lalmas, M · 2021
Later among the works it cites.
What are you optimizing for? Aligning Recommender Systems with Human Values
Stray, J., Adler, S., and Hadfield-Menell, D · 2021
Later among the works it cites.
Constitutional AI: Harmlessness from AI Feedback, December 2022
Bai, Y., Kadavath, S., Kundu, S., Askell, A., Kernion, J., Jones, A., Chen, A., Goldie, A., Mirhoseini, A., McKinnon, C., Chen, C., Olsson, C., Olah, C., Hernandez, D., Drain, D., Ganguli, D., Li, D., Tran-Johnson, E., Perez, E., Kerr, J., Mueller, J., Ladish, J., Landau, J., Ndousse, K., Lukosuite, K., Lovitt, L., Sellitto, M., Elhage, N., Schiefer, N., Mercado, N., DasSarma, N., Lasenby, R., Larson, R., Ringer, S., Johnston, S., Kravec, S., Showk, S. E., Fort, S., Lanham, T., Telleen-Lawton, T., Conerly, T., Henighan, T., Hume, T., Bowman, S. R., Hatfield-Dodds, Z., Mann, B., Amodei, D., Joseph, N., McCandlish, S., Brown, T., and Kaplan, J · 2022
Later among the works it cites.
Estimating and Penalizing Induced Preference Shifts in Recommender Systems, July 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kulkarni, K. and Neth, S · 2011
Cited alongside, same era.
Computational Social Choice
Brandt, F., Conitzer, V., and Endriss, U · 2012
Cited alongside, same era.
Where do preferences come from?
Dietrich, F. and List, C · 2013
Cited alongside, same era.
Nudge and the Manipulation of Choice: A Framework for the Responsible Use of the Nudge Approach to Behaviour Change in Public Policy
Hansen, P. G. and Jespersen, A. M · 2013
Cited alongside, same era.
Training a Robot via Human Feedback: A Case Study
Knox, W. B., Stone, P., and Breazeal, C · 2013
Cited alongside, same era.
Dynamic Social Choice: Foundations and Algorithms
Parkes, D. C. and Procaccia, A. D · 2013
Cited alongside, same era.
A Survey of Multi-Objective Sequential Decision-Making
Roijers, D. M., Vamplew, P., Whiteson, S., and Dazeley, R · 2013
Cited alongside, same era.
Carroll, M., Dragan, A., Russell, S., and Hadfield-Menell, D · 2022
Later among the works it cites.
Towards Psychologically-Grounded Dynamic Preference Models
Curmei, M., Haupt, A. A., Recht, B., and Hadfield-Menell, D · 2022
Later among the works it cites.
Preference Dynamics Under Personalized Recommendations
Dean, S. and Morgenstern, J · 2022
Later among the works it cites.
Path-Specific Objectives for Safer Agent Incentives
Farquhar, S., Carey, R., and Everitt, T · 2022
Later among the works it cites.
Franklin, M., Ashton, H., Gorman, R., and Armstrong, S · 2022
Later among the works it cites.
Hardt, M., Jagadeesan, M., and Mendler-Dünner, C · 2022
Later among the works it cites.
On the Sensitivity of Reward Inference to Misspecified Human Models, December 2022
Hong, J., Bhatia, K., and Dragan, A · 2022
Later among the works it cites.
Influencing Long-Term Behavior in Multiagent Reinforcement Learning, October 2022
Kim, D.-K., Riemer, M., Liu, M., Foerster, J. N., Everett, M., Sun, C., Tesauro, G., and How, J. P · 2022
Later among the works it cites.
The Challenge of Understanding What Users Want: Inconsistent Preferences and Engagement Optimization
Kleinberg, J., Mullainathan, S., and Raghavan, M · 2022
Later among the works it cites.
AI Safety and Preference Change, September 2022
Kolodny, N · 2022
Later among the works it cites.
Lindner, D. and El-Assady, M · 2022
Later among the works it cites.
Social Choice Theory
List, C · 2022
Later among the works it cites.
What we owe the future
MacAskill, W · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P., Leike, J., and Lowe, R · 2022
Later among the works it cites.
Choosing for Changing Selves
Paul, L. A · 2022
Later among the works it cites.
Nudging for Changing Selves
Pettigrew, R · 2022
Later among the works it cites.
How Platform Recommenders Work, November 2022
Thorburn, L · 2022
Later among the works it cites.
What Will “Amplification” Mean in Court?, 2022
Thorburn, L., Stray, J., and Bengani, P · 2022
Later among the works it cites.
Artificial Intelligence–Based Chatbots for Promoting Health Behavioral Changes: Systematic Review
Aggarwal, A., Tam, C. C., Wu, D., Li, X., and Qiao, S · 2023
Later among the works it cites.
Designing Fiduciary Artificial Intelligence, July 2023
Benthall, S. and Shekman, D · 2023
Later among the works it cites.
SHAPE: A Framework for Evaluating the Ethicality of Influence
Bezou-Vrakatseli, E., Brückner, B., and Thorburn, L · 2023
Later among the works it cites.
Artificial Influence: An Analysis Of AI-Driven Persuasion, March 2023
Burtell, M. and Woodside, T · 2023
Later among the works it cites.
Reinforcing User Retention in a Billion Scale Short Video Recommender System, February 2023
Cai, Q., Liu, S., Wang, X., Zuo, T., Xie, W., Yang, B., Zheng, D., Jiang, P., and Gai, K · 2023
Later among the works it cites.
Characterizing Manipulation from AI Systems, March 2023
Carroll, M., Chan, A., Ashton, H., and Krueger, D · 2023
Later among the works it cites.
Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback, July 2023
Casper, S., Davies, X., Shi, C., Gilbert, T. K., Scheurer, J., Rando, J., Freedman, R., Korbak, T., Lindner, D., Freire, P., Wang, T., Marks, S., Segerie, C.-R., Carroll, M., Peng, A., Christoffersen, P., Damani, M., Slocum, S., Anwar, U., Siththaranjan, A., Nadeau, M., Michaud, E. J., Pfau, J., Krasheninnikov, D., Chen, X., Langosco, L., Hase, P., Bıyık, E., Dragan, A., Krueger, D., Sadigh, D., and Hadfield-Menell, D · 2023
Later among the works it cites.
Incentives from a causal perspective
Everitt, T., Fox, J., Carey, R., MacDermott, M., Benthall, S., and Richens, J · 2023
Later among the works it cites.
Unveiling the Neely Ethics & Technology Indices, June 2023
Fast, N., Schroeder, J., Iyer, R., and Motyl, M · 2023
Later among the works it cites.
Reasoning about Causality in Games, January 2023
Hammond, L., Fox, J., Everitt, T., Carey, R., Abate, A., and Wooldridge, M · 2023
Later among the works it cites.
An Overview of Catastrophic AI Risks, October 2023
Hendrycks, D., Mazeika, M., and Woodside, T · 2023
Later among the works it cites.
Learning to Influence Human Behavior with Offline Reinforcement Learning, June 2023
Hong, J., Dragan, A., and Levine, S · 2023
Later among the works it cites.
Rewarding Chatbots for Real-World Engagement with Millions of Users, March 2023
Irvine, R., Boubert, D., Raina, V., Liusie, A., Mudupalli, V., Korshuk, A., Liu, Z., Cremer, F., Assassi, V., Beauchamp, C.-C., Lu, X., Rialan, T., and Beauchamp, W · 2023
Later among the works it cites.
User tampering in reinforcement learning recommender systems
Kasirzadeh, A. and Evans, C · 2023
Later among the works it cites.
Kirk, H. R., Vidgen, B., Röttger, P., and Hale, S. A · 2023
Later among the works it cites.
Algorithmic Displacement of Social Trust, 2023
Laufer, B. and Nissenbaum, H · 2023
Later among the works it cites.
Eliciting Human Preferences with Language Models, October 2023
Li, B. Z., Tamkin, A., Goodman, N., and Andreas, J · 2023
Later among the works it cites.
On The Fragility of Learned Reward Functions, January 2023
McKinney, L., Duan, Y., Krueger, D., and Gleave, A · 2023
Later among the works it cites.
Milli, S., Carroll, M., Wang, Y., Pandey, S., Zhao, S., and Dragan, A. D · 2023
Later among the works it cites.
AI Alignment and Social Choice: Fundamental Limitations and Policy Implications, October 2023
Mishra, A · 2023
Later among the works it cites.
The Amplification Paradox in Recommender Systems, February 2023
Ribeiro, M. H., Veselovsky, V., and West, R · 2023
Later among the works it cites.
Towards Understanding Sycophancy in Language Models, October 2023
Sharma, M., Tong, M., Korbak, T., Duvenaud, D., Askell, A., Bowman, S. R., Cheng, N., Durmus, E., Hatfield-Dodds, Z., Johnston, S. R., Kravec, S., Maxwell, T., McCandlish, S., Ndousse, K., Rausch, O., Schiefer, N., Yan, D., Zhang, M., and Perez, E · 2023
Later among the works it cites.
Siththaranjan, A., Laidlaw, C., and Hadfield-Menell, D · 2023
Later among the works it cites.
Making Amplification Measurable, May 2023
Thorburn, L · 2023
Later among the works it cites.
Causal Confusion and Reward Misidentification in Preference-Based Reward Learning, March 2023
Tien, J., He, J. Z.-Y., Erickson, Z., Dragan, A. D., and Brown, D. S · 2023
Later among the works it cites.
Honesty Is the Best Policy: Defining and Mitigating AI Deception
Ward, F. R., Everitt, T., Toni, F., and Belardinelli, F · 2023
Later among the works it cites.
Ahmadian, A., Cremer, C., Gallé, M., Fadaee, M., Kreutzer, J., Pietquin, O., Üstün, A., and Hooker, S · 2024
Closest in time.
The Problem of Legitimate Value Change: Value Malleability and AI Alignment
Ammann, N · 2024
Closest in time.
Social Choice for AI Alignment: Dealing with Diverse Human Feedback, April 2024
Conitzer, V., Freedman, R., Heitzig, J., Holliday, W. H., Jacobs, B. M., Lambert, N., Mossé, M., Pacuit, E., Russell, S., Schoelkopf, H., Tewolde, E., and Zwicker, W. S · 2024
Closest in time.
What We Know About Using Non-Engagement Signals in Content Ranking, February 2024
Cunningham, T., Pandey, S., Sigerson, L., Stray, J., Allen, J., Barrilleaux, B., Iyer, R., Milli, S., Kothari, M., and Rezaei, B · 2024
Closest in time.
The Ethics of Advanced AI Assistants, April 2024
Gabriel, I., Manzini, A., Keeling, G., Hendricks, L. A., Rieser, V., Iqbal, H., Tomašev, N., Ktena, I., Kenton, Z., Rodriguez, M., El-Sayed, S., Brown, S., Akbulut, C., Trask, A., Hughes, E., Bergman, A. S., Shelby, R., Marchal, N., Griffin, C., Mateos-Garcia, J., Weidinger, L., Street, W., Lange, B., Ingerman, A., Lentz, A., Enger, R., Barakat, A., Krakovna, V., Siy, J. O., Kurth-Nelson, Z., McCroskery, A., Bolina, V., Law, H., Shanahan, M., Alberts, L., Balle, B., de Haas, S., Ibitoye, Y., Dafoe, A., Goldberg, B., Krier, S., Reese, A., Witherspoon, S., Hawkins, W., Rauh, M., Wallace, D., Franklin, M., Goldstein, J. A., Lehman, J., Klenk, M., Vallor, S., Biles, C., Morris, M. R., King, H., Arcas, B. A. y., Isaac, W., and Manyika, J · 2024
Closest in time.
Lang, L., Foote, D., Russell, S., Dragan, A., Jenner, E., and Emmons, S · 2024
Closest in time.
Preference Change
Strohmaier, D. and Messerli, M · 2024
Closest in time.
On The Expressivity of Objective-Specification Formalisms in Reinforcement Learning, February 2024
Subramani, R., Williams, M., Heitmann, M., Holm, H., Griffin, C., and Skalse, J · 2024
Closest in time.
The Reasons that Agents Act: Intention and Instrumental Goals, February 2024
Ward, F. R., MacDermott, M., Belardinelli, F., Toni, F., and Everitt, T · 2024
Closest in time.
Beyond Preferences in AI Alignment
Zhi-Xuan, T., Carroll, M., Franklin, M., and Ashton, H · 2024
Closest in time.